Stambaugh 1999 Replication#
Last updated: Aug 20, 2026, 1:04:28β―PM
Project Notes
Table of Contents#
Notebooks π
Pipeline Charts π
Pipeline Specs#
Pipeline Name |
Stambaugh 1999 Replication |
|---|---|
Pipeline ID |
|
Maintainer |
Ashish Maheshwari & Omar Anabtawi |
Contributors |
Ashish Maheshwari & Omar Anabtawi |
Repository |
|
Pipeline Web Page |
|
Date of Last Code Update |
2026-08-20 13:04:16 |
OS Compatibility |
Windows, Linux, macOS |
Linked Dataframes |
Build Commands:
pip install -r requirements.txt
doit
Does the dividendβprice ratio predict stock returns? For decades the standard test regressed next monthβs excess return on this monthβs dividend yield and usually found a positive, βsignificantβ slope. Stambaugh (1999, Journal of Financial Economics 54) showed the test is broken in a quantifiable way: the dividend yield is highly persistent and shares a price with the return, so the OLS slope is biased upward in finite samples. A positive slope is what you should expect even when the true slope is zero.
This project rebuilds the data from CRSP, replicates the paperβs Table 1, Table 2, and Figure 1 within a documented tolerance, and extends every exhibit through 2024 β where the problem turns out to be worse than in the paperβs own sample: the gap between the naive and the finite-sample p-value has grown from roughly threefold to tenfold.
Read the project site β overview, walkthrough notebook, interactive playground, and full report.
Exhibits#
Exhibit |
Content |
Headline result |
|---|---|---|
Table 1 |
Finite-sample properties of the OLS slope, by subsample |
Our finite-sample p-value 0.177 vs the paperβs 0.17 |
Table 2 |
Bayesian posteriors under four prior/likelihood specifications |
All sixteen cells within ~0.03 of the paper |
Figure 1 |
Ξ² vs Ο across methods and subperiods |
Reproduces the paperβs Ο > 1 overshoot in 1977β96 |
Updates |
Tables 1 and 2 on samples through 2024 |
Naive p 0.015 vs honest p 0.149 in 1997β2024 |
Plus two educational products: a guided walkthrough notebook and an interactive browser playground where two sliders show the bias growing with persistence and shrinking with sample size.
Repository structure#
.
βββ dodo.py # PyDoit build file β runs everything
βββ chartbook.toml # site configuration
βββ requirements.txt
βββ .env.example # template for your .env (WRDS_USERNAME)
βββ src/
β βββ settings.py # configuration: paths, dates, credentials
β βββ pull_CRSP_index.py # WRDS pull: CRSP monthly market index
β βββ pull_fama_french.py # WRDS pull: risk-free rate
β βββ calc_predictor_data.py # DATA CLEANING ONLY -> tidy monthly panel
β βββ monte_carlo.py # simulation engine (Table 1 Part A)
β βββ stambaugh_bias.py # bias-corrected estimators (Figure 1)
β βββ bayesian.py # conjugate posteriors (Table 2 specs A, B)
β βββ mcmc.py # Metropolis-Hastings (Table 2 specs C, D)
β βββ create_table_01_partC.py
β βββ create_table_01.py # Table 1 assembly -> _output/*.tex
β βββ create_table_02.py # Table 2 assembly -> _output/*.tex
β βββ create_figure_01.py # Figure 1 -> _output/figure_01.png
β βββ 01_walkthrough.ipynb.py # guided tour notebook (jupytext source)
β βββ test_*.py # unit tests
βββ reports/report.tex # the write-up; inputs the generated exhibits
βββ docs_src/ # site sources (edit these)
βββ docs/ # BUILT site β generated, served by Pages
βββ _data/ # pulled and cleaned data (git-ignored)
βββ _output/ # generated tables, figures (git-ignored)
Data cleaning lives in its own file, separate from all analysis. Raw data never enters the repository.
Setup#
conda create -n stambaugh python=3.12 -y
conda activate stambaugh
pip install -r requirements.txt
cp .env.example .env # then set WRDS_USERNAME=your_login
The first WRDS connection prompts for your password and offers to create a
.pgpass file so later runs are non-interactive. .env is git-ignored and must
never be committed.
Running it#
doit
That pulls from WRDS, builds the tidy panel, regenerates every table and figure, compiles the report, executes the notebook, rebuilds the site, and runs the tests. PyDoit tracks dependencies, so re-running rebuilds only what changed.
Individual stages:
doit pull # WRDS pulls
doit clean_data # tidy panel
doit table_01 figure_01 # paper-sample exhibits
doit table_01_updated # extended-sample exhibits
doit notebook # execute the walkthrough
doit compile_latex_docs # report PDF
doit build_chartbook_site # the published site
doit run_pytest # test suite
Note that table_02 runs eight Metropolis-Hastings chains and takes several
minutes.
Testing#
pytest -q src/
The suite is split deliberately. Simulation-based tests verify the bias mechanism itself β that the bias is positive when innovations are negatively correlated, vanishes when they are not, shrinks like 1/T, and matches the Kendall/Stambaugh analytical formula β and run anywhere, including CI without credentials. Data-dependent tests check our estimates against the paperβs published values within stated tolerances, and skip with an explanatory message when the panel has not been built.
Data sources#
CRSP Monthly Stock Market Indexes (
crsp.msi, WRDS) β value-weighted returns with (vwretd) and without (vwretx) dividends. Their difference gives the dividend series, which is how the dividendβprice ratio is reconstructed without a separate dividend file.FamaβFrench monthly factors (
ff.factors_monthly, WRDS) β the one-month risk-free rate, for continuously compounded excess returns.
Stambaugh uses a NYSE-only value-weighted index; our WRDS instance provides no pre-built NYSE-only monthly index carrying both return columns, so we use the CRSP total-market value-weighted index and document the choice in the report.
A note on docs/#
The built site is committed so GitHub Pages can serve it directly without a
build step. Edit docs_src/, never docs/ β the latter is regenerated by
doit build_chartbook_site and hand edits are lost.
Team#
Ashish Maheshwari
Omar Anabtawi
FINM 32900, Full-Stack Quantitative Finance, Summer 2026.