Big Data and Machine Learning for Applied Economics
Graduate course on statistical learning for applied economics and finance: the predict-vs-explain distinction, regression and classification as prediction machines, cross-validation and regularization, trees, random forests and boosting, model interpretation (variable importance, partial dependence, SHAP), and representation with PCA, k-means and neural networks — following ISL and applied in R on one running case: predicting the vote of every Colombian polling station.
Instructor: Eduard F. Martínez-González
Institution: Universidad ICESI — Department of Economics
Course code: 60-121 · NRC 10-993
Program: Master's in Economics (ME)
Term: September 15 – November 24, 2026
Credits: 3 · Weekly hours: 3
Location: Room 406-E (Block E)
Time: Tuesdays, 8:00–11:00 a.m.
Original title (in Spanish): Big Data y Machine Learning para Economía Aplicada (60-121, NRC 10-993). Lectures and materials are in Spanish. Prerequisites: Econometrics I & II (or equivalent); basic R is recommended, not required.
All course materials — syllabus, lecture slides, R applications, problem sets and the paper library — are hosted on GitHub:
github.com/eduard-martinez/bdml-applied-economics Full syllabus (PDF, in Spanish) Final project guidelines · soon Course text — ISL, 2nd ed. (free online)
Course description
In causal inference the goal is to identify a parameter correctly: the effect β of an intervention on an outcome. In this course the goal changes place — the interest is Y: building models that predict the variable of interest well on data the model has not seen, evaluating that predictive performance honestly, and opening the model to understand what the prediction depends on. That distinction between explaining and predicting organizes the whole semester and is drawn from the very first session.
The course introduces Master’s students to the statistical-learning framework and to the machine-learning tools most used today in applied economics and finance: regression as a prediction machine, cross-validation and the honest pipeline, regularization (Ridge, Lasso, Elastic Net), classification and the threshold as a decision with costs, trees, random forests and boosting, interpretation tools (permutation importance, partial dependence, SHAP) and representation with PCA, k-means and neural networks. The progression follows An Introduction to Statistical Learning (James, Witten, Hastie & Tibshirani, 2nd ed.), the course textbook.
The treatment combines sufficient formality with a lot of intuition. For each method the course presents its objective function and hyperparameters, explains what problem it solves and how it works, implements it on real data in R, evaluates its out-of-sample performance against a baseline model, and interprets the results. Long derivations live in each session’s Para profundizar appendix.
One running case. From session 2 to session 8 the class works on a single problem: predicting the vote of each of Colombia’s 12,001 polling stations in the 2022 presidential runoff, using the census characteristics of the station’s surroundings and its location. Every week the same table of leaders grows by one model, evaluated on the same untouched test set. The problem sets transfer the tools to a second dataset — housing prices in Cali — and in the final project each team brings its own data.
Learning outcomes
By the end of the course, students will be able to:
- Distinguish a prediction problem from a causal-inference problem, and recognize when out-of-sample predictive performance is the quantity of interest.
- Build the complete workflow in R: splitting, preprocessing inside the pipeline, validation, tuning, evaluation against a baseline and interpretation.
- Evaluate regression and classification models with the appropriate metrics (RMSE, MAE, out-of-sample R², AUC, calibration, costs of each error), and diagnose where the errors concentrate.
- Apply and compare regularization, k-NN, trees, random forests, boosting and neural networks on economic data, knowing when each family is the right choice.
- Open a flexible model with permutation importance, partial dependence and SHAP — and know what it says and what it does not say.
- Read critically and present economic research that uses ML.
How each class works
Each session is organized in three parts:
- Theory and concepts. For every method, the same sequence: the problem it solves, the intuition, the essential formulation, how it works and how it is tuned, the application on real data, the evaluation and the interpretation. Long derivations go to the session’s appendix.
- Applied work in R. The last stretch of each class is hands-on work with prepared code on the running case (week 1 is the conceptual exception, with a small simulated example). Each script reports out-of-sample performance against the previous week’s leader; they are written in base R plus a handful of packages (
rio,dplyr,glmnet,rpart,randomForest,xgboost,nnet), with no hidden machinery. - The papers. Each session has a reference applied article that shows the method at work in published research. In session 2 each team picks the paper it will present — two options per field of economics — and presents it in session 8, in 15 minutes: the question and the decision behind it, the data, the baseline, the out-of-sample evaluation and where the model fails, and the main result.
Getting up to speed in R
The R leveling is mandatory, self-paced and happens outside class hours:
- Introduction to R — R, RStudio, objects and data frames.
- Data wrangling with
dplyr - Reading and writing data
- Good practices in data management
- Good modeling practices with
tidymodels— material posted on Intu.
Equivalent reference material is offered for students who choose the Python track.
Schedule
Each week shows its class materials as buttons — the lecture slides, the R application script and a zip with the script plus its data — with the readings and the week’s milestones on the small lines below.
Key dates. Problem set 1: published Oct 6, due Oct 20 at 8:00 a.m. · Project pitch: Nov 3 · Problem set 2: published Oct 20, due Nov 10 at 11:59 p.m. · Final project document (max. 8 pages) and presentations: Nov 24.
Module 1 — Foundations and evaluation
Week 1 · Sep 15 — Predict, don’t explain: the statistical-learning framework. From β to Y; the generalization risk and why the conditional mean is the best predictor; the two golden rules (training error is optimistic; the test set is opened once); bias, variance and the U-curve.
Lecture slides R application · one dataset, three models Download · script + data (.zip) Readings: ISL ch. 1–2 (section 2.2) · Mullainathan & Spiess (2017), Machine Learning: An Applied Econometric Approach
Week 2 · Sep 22 — Linear regression as a prediction machine, and how to evaluate a prediction. The same OLS with a different criterion: what survives, what stops mattering and what appears; k-NN as the first contrast; metrics, the baseline in the first row of the table, and where to look for the errors. The running case starts.
Lecture slides R application · the vote of a polling station Download · script + data (.zip) Readings: ISL ch. 3 · ch. 7 (steps, splines and GAMs) Milestone: teams choose their session-8 paper (two options per field of economics)
Week 3 · Sep 29 — Cross-validation and regularization: the honest pipeline. Two tools that need each other. The judge: cross-validation estimates out-of-sample error inside the training set, so models and hyperparameters are chosen without touching the test. The dial: regularization puts a price on the size of the coefficients — Ridge shrinks, the Lasso switches variables off — and its intensity λ is set by the judge. With 1,291 columns OLS collapses and the Lasso keeps 202.
Lecture slides R application · cross-validation and the Lasso Download · script + data (.zip) Readings: ISL 5.1 (cross-validation) · 6.2 and 6.4 (Ridge, Lasso, high dimension)
Week 4 · Oct 6 — Classification: from probability to decision. Same election, same stations, same data; the y changes: did Petro win the station? Classifying is estimating a probability and then deciding — the linear probability model breaks out of [0,1] and the logit does not; the model orders the stations (AUC), calibration says whether its probabilities are credible, and the cost of each error fixes the threshold.
Lecture slides R application · does Petro win the station? Download · script + data (.zip) Readings: ISL 4.1–4.3 and 4.4.2 (confusion matrix and ROC) · 4.7.6 (the k-NN classifier lab) · Fawcett (2006) Milestone: Problem set 1 published — PDF · vivienda_cali.rds · diccionario_vivienda.csv (due Oct 20, 8:00 a.m.)
Module 2 — Trees, ensembles and interpretation
Week 5 · Oct 13 — Trees, forests and boosting: split, average and correct. Until now we wrote the flexibility by hand — which variable crossed with which, and how. With 37 variables there are 666 pairwise crossings: nobody writes those. A tree splits the stations into similar groups and predicts the average of each, so a split inside a split is an interaction nobody had to write; the forest averages hundreds of trees and boosting adds small trees that correct the previous one (rpart, randomForest, xgboost).
Lecture slides R application · tree, forest and boosting Download · script + data (.zip) Readings: ISL ch. 8 (8.1 the tree, 8.2 the ensembles) · Breiman (2001) Thread paper: Gelvez, Cardiles, Martínez-González & Muñoz (2026), How Predictable Is an Election? — PDF
Week 6 · Oct 20 — Opening the black box: what the machine learned. No new model: today we walk into last week’s, 738 trees and 38,330 leaves that nobody can read. Four questions, four tools — which variables it uses (importance), with what shape (partial dependence), why it predicts what it predicts at a given station (Shapley and SHAP) and who it fails (error by subgroup) — and a fifth, the important one: what none of them tells us. Each tool moves a variable and watches the prediction: it asks the model, not the world.
Lecture slides R application · what the leader learned Download · script + data (.zip) Readings: Molnar (2025), the permutation-importance, PDP and SHAP chapters (read the disadvantages twice) Milestones: Problem set 1 due (8:00 a.m.) · Problem set 2 published (sessions 5–7; due Nov 10, 11:59 p.m.)
Module 3 — Representation, synthesis and project
Week 7 · Oct 27 — Representing: compressing with meaning. How many different things does the census actually know about a station — and who decides how to summarize them, the variance of the x’s or the y? Without y: principal components compress the 37 columns and k-means groups the stations into types of territory. With y: a neural network builds its own variables and tunes them together with the prediction, and sits the usual exam against boosting. What varies most is not necessarily what predicts.
Lecture slides R application · PCA, k-means and a network Download · script + data (.zip) Readings: ISL 12.2 (principal components) and 12.4.1 (k-means) · 10.1–10.3 and 10.6–10.7 (neural networks) Milestone: the session opens with the Problem set 1 debrief and ranking
Week 8 · Nov 3 — The papers and the synthesis: student presentations, the course map and the project pitches. No new method: today all of them are used. The student paper presentations (15 minutes each), the map that puts papers, methods and workflow together, and the one-page project pitches (five minutes each).
Lecture slides Readings: Athey & Imbens (2019), sections 1–3 (the causal bridge) Session papers: eight fields of economics with two options each — political economy, crime and justice, health, development and poverty, labour and education, finance and credit, macro and forecasting, IO and demand. Each team picked its field and paper in session 2; who presents what is announced in class. Paper library Milestone: project pitch — the idea in one page
Week 9 · Nov 24 — The final project: presentations and course closing. The Problem set 2 ranking and the course in one slide; then each team presents its predictive research proposal and hands in the document.
Lecture slides Milestones: final project presentations · project document submitted (max. 8 pages)
Problem sets
Both problem sets are applied only — no conceptual questions — and work on a second dataset: housing prices in Cali, so the tools travel to a problem other than the class case. Each one closes with a ranking of the teams by test error.
- Problem set 1 — How much is a home in Cali worth? (sessions 2–4). In pairs, due Oct 20 at 8:00 a.m., before session 5. Real for-sale listings in Cali (2019–2020): a baseline and a regression, cross-validation and regularization, the binary version (is it social housing?), and a final model chosen with criteria. The test listings have their price hidden, so every intermediate decision is made with validation inside the training set and the test is used once, to predict.
Problem set 1 (PDF) vivienda_cali.rds diccionario_vivienda.csv
- Problem set 2 — sessions 5–7. Published Oct 20, due Nov 10 at 11:59 p.m. PDF and data: available soon.
Each problem set is submitted on the virtual campus as a .zip with the report in PDF, the code (00_run.R and its auxiliary scripts) and the predicciones.csv that the code produces: the teacher scores those predictions against the hidden prices, and that test error ranks the teams. The ranking is discussed in session 7.
Evaluation
| Component | Weight |
|---|---|
| Paper presentation (15 min, session 8) | 10% |
| Problem set 1 — sessions 2–4 | 15% |
| Problem set 2 — sessions 5–7 | 15% |
| Final project | 60% |
Final project. A predictive research proposal: each team defines an economic question in which prediction is the quantity of interest, identifies its own data, and designs the study that survives the seven questions of the course — the question and the decision behind it, the unit and the outcome, the data and what is known at prediction time, the metric and the baseline, the validation strategy, the risks (leakage, extrapolation, subgroups where it fails), and the reading of the results. The idea is pitched in one page on Nov 3 (session 8); the document — at most eight pages — and the presentation are due on Nov 24 (session 9).
AI policy: AI tools are allowed in every component of the course, as long as their use is explicitly declared — what was used and for what.
Reading library
The repo’s literature/papers/ folder collects the applied articles of the course, organized by field of economics, so every student has them at hand. Session 8 is built on this library: each field offers two options and each team picked one in session 2 to present.
- Political economy — 4 papers.
- Crime and justice — 5 papers.
- Development and poverty — 4 papers.
- Labour, education and human capital — 2 papers.
- Housing and urban economics — 5 papers.
- Macroeconomics and forecasting — 6 papers.
- Finance and credit — 7 papers.
- IO, demand and firms — 4 papers.
- The prediction framework for economists — 2 papers.
The course’s own thread paper — Gelvez, Cardiles, Martínez-González & Muñoz (2026), How Predictable Is an Election? A Machine-Learning Approach to Electoral Behavior in Colombia — is in the political-economy folder; it is the benchmark the class case is measured against.
Core bibliography
- James, G., Witten, D., Hastie, T., & Tibshirani, R. (2021). An Introduction to Statistical Learning with Applications in R (2nd ed.). Springer. [ISL] — free online; a course copy is in the repo’s
literature/books/folder. - Molnar, C. (2025). Interpretable Machine Learning: A Guide for Making Black Box Models Explainable (3rd ed.) — free online.
- Hastie, T., Tibshirani, R., & Friedman, J. (2009). The Elements of Statistical Learning (2nd ed.). Springer. [ESL]
- Goodfellow, I., Bengio, Y., & Courville, A. (2016). Deep Learning. MIT Press.