Quantitative research · Event study

Dead Cat Detector

When extreme stock selloffs reverse—and when they keep falling

Using 14,678 extreme one-day declines among point-in-time S&P 500 constituents from 2007–2026, this study tests whether severe stock crashes systematically mean-revert, and whether observable event-time features distinguish rebounds from continued underperformance.

The result

There is no average dead‑cat bounce.

We estimate a mean 20-day cumulative abnormal return against SPY of −0.265% (95% bootstrap CI [−0.403%, −0.124%], n = 14,647). Mean CAR is negative at every horizon from 1 to 60 trading days.

−0.27%Mean CAR20 vs SPY
47.1%End above benchmark
38.9%Regain pre-crash close
8.4%Std. dev. of CAR20
Mean cumulative abnormal return against SPY from the event close through 60 trading days, with a 95% bootstrap confidence band. The path drifts below zero and remains there.
Mean cumulative abnormal return against SPY from the event close through 60 trading days, with a 95% bootstrap band from 2,000 whole-event resamples. The path drifts below zero and stays there.
HorizonMean CAR95% CI
1d−0.101%[−0.141%, −0.064%]
5d−0.268%[−0.342%, −0.192%]
10d−0.187%[−0.283%, −0.089%]
20d−0.265%[−0.403%, −0.124%]
60d−0.425%[−0.658%, −0.200%]

This is not a trading signal. No transaction costs, borrow costs or capacity constraints are modelled. A −0.27% mean inside a 8.4% standard deviation is a statistical displacement, not an edge.

The dataset

Daily split- and dividend-adjusted OHLCV, with S&P 500 membership reconstructed point in time — the current constituent list rolled backwards through the index change log — so a stock is eligible for event detection only on dates it actually belonged to the index.

14,678Crash events (z ≤ −3σ)
600Tickers
2.98MTicker-days
19.7Years (2007–2026)

A crash is a day on which a stock's return falls at least three standard deviations below its own trailing 60-day distribution, with mean and volatility estimated from observations strictly preceding the event. A 20-trading-day ticker-specific cooldown collapses each episode to its first day, so consecutive breaches are not counted as independent observations.

Robustness

The result does not depend on how a crash is defined. Across 576 persisted specifications — crash threshold × volatility window × cooldown × outcome horizon × high-VIX definition — mean CAR20 is negative in 48 of 48 threshold × window × cooldown cells, and the recovery rate is below 50% in 48 of 48.

48 / 48Specs with negative mean CAR
48 / 48Specs with recovery below 50%
0Specs positive & significant
Heatmap of mean 20-day cumulative abnormal return across crash thresholds, volatility windows, cooldowns and horizons. Every cell is negative.
Mean CAR20 across the specification grid. Every cell is negative, and severity makes the drift worse rather than better: −0.19% at −2.5σ deepening to −0.49% at −4.0σ. Built from the complete grid, not a favourable subset.

Dispersion is the real story

The average effect is modestly negative. The variation around it is enormous: the standard deviation of CAR20 is 8.4% against a −0.27% mean — roughly thirty times the central tendency. Almost every individual crash resolves dramatically better or worse than the average.

Distribution of 20-day cumulative abnormal returns across all crash events, a wide and near-symmetric distribution centred just below zero.
The distribution of 20-day outcomes: wide, near-symmetric, and centred a fraction of a percent below zero. The economically interesting problem is this heterogeneity, not a universal bounce.

Hypotheses

Five hypotheses were pre-registered before estimation. Four failed, and they are reported as written.

HypothesisResult
H1Extreme declines do not universally mean-revertSupported
H2Broad-market crashes recover differently from idiosyncraticNot supported (p = 0.43)
H3Abnormal volume contains informationSuggestive, insufficient (q = 0.14)
H4Pre-crash momentum mattersNot supported (p = 0.478)
H5Market volatility changes the relationshipNot supported (p = 0.57)
Mean cumulative abnormal return by crash type. Broad-market, sector and idiosyncratic confidence bands overlap one another throughout.
Crash type does not separate 20-day outcomes: broad-market −0.35%, sector −0.24%, idiosyncratic −0.23%. Three confidence bands sitting on top of one another.

Can we predict recovery?

Three model families were trained to classify whether an event would end above the benchmark, using chronological splits with a 20-trading-day embargo between blocks. Hyper-parameters were chosen on validation only; the test block was scored once.

ModelTest ROC-AUCBrierBrier skillAccuracy
Logistic regression0.5040.2496−0.0020.524
Random forest0.4990.2499−0.0030.516
LightGBM0.4760.2564−0.0300.493
Constant base rate0.5000.2490−0.0000.531

Out-of-sample discrimination is indistinguishable from chance, and all three models have negative Brier skill — each is fractionally worse than predicting the training base rate for every event. The constant base-rate model also has the highest accuracy while making the same prediction every time, which is precisely why accuracy is reported last.

Calibration curves for the three models, tracking the diagonal closely but spanning a narrow range of predicted probabilities.
Calibration is good (ECE 0.028–0.063). The models are well calibrated around the base rate and carry no discriminating information — honest about probabilities while knowing nothing about which event is which.

Why the interpretability plots do not rescue this

Logistic coefficients, permutation importance and SHAP each rank a different feature first. Disagreement among interpretation methods, combined with chance-level predictive performance, is consistent with unstable feature ranking over noise rather than with real signal the models failed to exploit. Attributions extracted from a model with 0.476 test AUC describe what an uninformative fit latched onto in-sample; they are not substantive economics.

Side-by-side comparison of logistic coefficients and permutation importance, which rank different features at the top.
Three interpretation methods, three different rankings. An interpretability plot inherits the credibility of the model beneath it.

Multiple testing

0 of 32 exploratory coefficients survive correction.

Pooling every exploratory coefficient from the extended OLS and logistic models into a single Benjamini–Hochberg family, none survives at q ≤ 0.05.

This is the study's most important statistical guard. Read without correction, the exploratory models offer several coefficients at nominal p < 0.05 — a tempting set of “findings”. But 32 tests against a near-null outcome are expected to produce roughly that many false positives by construction. Reporting them would have been an artefact of the number of tests performed rather than evidence about markets.

The same logic governs abnormal volume. Its nominal p of 0.028 becomes q = 0.14 after correction. The sign is directionally consistent in 97.9% of specifications, which makes it a genuine lead — but the honest description is suggestive but statistically insufficient, not that volume predicts recovery.

Data quality

Before any analysis, two screens removed 54 tickers, each rejection manually inspected. Both drop whole tickers rather than winsorising individual returns, because the defect is series identity: when a price history is structurally inconsistent or represents more than one security, winsorising extreme returns does not repair the problem — it launders it into plausible-looking numbers.

Before screening, mean CAR20 computed to +77.7% with a standard deviation near 58. That is not a research result — it is a data-quality failure found during audit. Six corrupt events on two tickers were overwhelming fourteen thousand real ones.

The NVDA false positive

An early version of the screen used an absolute price-level rule and incorrectly flagged legitimate NVDA history, whose split-adjusted median close over the window is $0.80. Absolute price thresholds are unsafe in adjusted data because splits, scale changes and genuine high-growth securities all produce legitimately tiny early prices. The corrected rules test internal time-series consistency instead — round-trip level discontinuities, extreme-move frequency and value diversity — none of which depend on where a price sits in absolute terms.

Limitations

  1. Survivorship bias is reduced, not solved. Point-in-time membership removes look-ahead in universe selection, but only 35.0% of historical-only members have usable price history (99.4% of current members do). If the missing names disproportionately include failed firms, the surviving sample is biased upward — against the negative finding reported here. That makes it unlikely this bias created the result, but its magnitude is not identified.
  2. No delisting returns. A stock removed mid-window stops contributing rather than realising a terminal value.
  3. Market-adjusted, not risk-adjusted. CAR subtracts SPY, not a factor model; a volatility factor could absorb the one surviving coefficient.
  4. Close-to-close only. A stock that fell 30% intraday and closed flat is not an event here.
  5. One market, one era. Large-cap US equity, 2007–2026. Conclusions hold within this design and do not extend automatically elsewhere.
  6. A near-null is not proof of no effect. Fundamentals, news text, options-implied data and order flow are unexamined.

Reproduce it

Every number on this page is read from persisted result files at build time, so the presentation cannot drift from the analysis. The page itself runs no computation.

git clone https://github.com/Gariyuuu/dead-cat-detector
cd dead-cat-detector
make setup      # uv venv (Python 3.12) + dependencies

make verify     # check all 42 documented claims against persisted results
make test       # 43 tests, no network required
make analysis   # rebuild every analysis stage from cached data
make all        # full pipeline including the data download

All randomness is seeded, and a config fingerprint is stamped into every persisted result. The test suite is concentrated on the leakage surface; its load-bearing case mutates every return from the event date onward and asserts the rolling statistics that define the crash are unchanged.

View the full study on GitHub →