What Fifty Years of Testing Found Wrong With CAPM
A theory is only as good as its ability to survive contact with data, and CAPM has been tested against more historical return data than perhaps any other idea in finance. The results forced a genuine reckoning, one that quietly rewired how professional investors think about what actually drives stock returns.
How a CAPM test actually works
Testing CAPM sounds simple in principle: rank a large sample of stocks by their estimated beta, group them into portfolios, and check whether the portfolios with higher average beta actually delivered higher average realized return over some long historical period, in the proportion the security market line predicts. In practice, the standard approach, developed by Eugene Fama and James MacBeth in 1973, runs in two stages. First, betas are estimated for individual stocks or portfolios using a rolling window of past returns regressed against a market index. Second, in each subsequent period, average stock returns are regressed cross-sectionally against those pre-estimated betas, and the resulting slope coefficient is compared against the market risk premium the theory predicts it should equal.
This two-stage design was itself an innovation, since it separates the estimation of risk (beta) from the test of whether that risk is priced, reducing a statistical problem called errors-in-variables bias that would otherwise contaminate a simpler one-stage test. It remains the backbone of empirical asset pricing methodology even today, applied constantly to test newer multifactor models, not just the original CAPM.
Researchers rarely test individual stocks directly, and for good statistical reason. A single company's monthly returns are dominated by firm-specific noise, an earnings surprise, a product launch, a lawsuit settlement, which swamps the systematic component the test is actually trying to isolate. The standard workaround groups stocks into portfolios, typically twenty five or so, sorted first by an initial beta estimate and periodically reformed, before running the cross-sectional test on the portfolios rather than on individual names. Grouping diversifies away much of the firm-specific noise while preserving the spread in average beta across groups, giving the test meaningfully more statistical power than running it stock by stock ever could.
The flat-line finding, worked through
Suppose CAPM theory, using a risk-free rate of 4% and a market risk premium of 6%, predicts that portfolios sorted into beta deciles should show average returns rising in a straight line: a beta-0.4 portfolio should earn 4% + 0.4 × 6% = 6.4%, and a beta-1.8 portfolio should earn 4% + 1.8 × 6% = 14.8%, a spread of 8.4 percentage points across the full range of beta. What large-sample empirical studies have repeatedly found instead is a much flatter realized relationship: the low-beta decile earning closer to 8%, well above its predicted 6.4%, and the high-beta decile earning closer to 12%, well below its predicted 14.8%, a realized spread of only about 4 points instead of 8.4.
This pattern, sometimes called the low-volatility anomaly or the betting-against-beta effect, is one of the most robust and widely replicated findings in empirical asset pricing, showing up across different countries, time periods, and asset classes, including bonds and currencies, not just equities. It directly contradicts the security market line's core prediction that expected return should rise proportionally, one-for-one, with beta.
Several explanations have been proposed for why this flattening occurs, none fully settled. One candidate is leverage constraints: investors who are barred from borrowing to buy more of a low-beta, low-return portfolio may instead reach for the same expected higher return by simply buying higher-beta stocks outright, bidding up their price and compressing their future expected return in the process, exactly the mechanism the leverage-constrained zero-beta model formalizes. A second candidate is a behavioral preference for lottery-like payoffs, evidence suggests many investors are drawn to volatile, high-beta stocks for the same reason people buy lottery tickets, a small chance of a very large gain, which can push these stocks' prices up and their subsequent realized returns down relative to what pure risk compensation would predict.
Roll's critique: can CAPM even be tested?
In 1977, economist Richard Roll raised a critique that still shapes how seriously academics take any single CAPM test. CAPM's equilibrium requires that the market portfolio used in the formula be mean-variance efficient, meaning no other combination of the same assets could deliver a higher expected return for the same risk. But the theoretical market portfolio must include every risky asset in existence: every stock and bond worldwide, real estate, private businesses, human capital, even collectibles. No researcher has ever observed that portfolio, let alone tested its efficiency directly. Every empirical test instead substitutes a proxy, typically a broad domestic stock index, and Roll showed mathematically that using a proxy that is not itself mean-variance efficient can produce a rejection of CAPM even if the true, unobservable model is perfectly correct, or can produce an apparent confirmation even if the true model is false.
Roll's critique does not prove CAPM is right; it proves that no test can prove it wrong with full certainty either, since every test is really a joint test of CAPM plus the specific market proxy chosen. This is a genuinely uncomfortable conclusion for a discipline that prizes falsifiable theory, and it partly explains why the profession moved toward comparing models by their relative explanatory power rather than searching for one final, definitive test of CAPM alone.
The anomalies that would not go away
Beyond the flat security market line, researchers through the 1980s and 1990s catalogued a growing list of patterns in average stock returns that beta could not explain. Smaller companies, measured by total market capitalization, earned higher average returns over long periods than their betas predicted, a pattern known as the size effect. Companies trading at low prices relative to their book value, so-called value stocks, similarly outearned what beta alone would justify, the value effect. Stocks that had recently performed well tended to keep performing well over subsequent months, the momentum effect, a pattern with essentially no clean explanation inside the original CAPM framework at all. Each of these findings survived out-of-sample testing across different time periods and different national markets, which is the standard academic finance uses to distinguish a genuine, structural pattern from a coincidental result mined from one particular dataset.
By the early 1990s, the accumulated weight of these anomalies had become difficult to dismiss as measurement error or a temporary fluke, setting the stage for a direct challenge to CAPM's single-factor structure itself, rather than merely a patch to one of its assumptions.
Skeptics of any single anomaly are right to raise a specific concern known as data snooping: with enough researchers testing enough candidate variables against enough historical stock return datasets, some fraction of "discoveries" will be statistically significant purely by chance, the equivalent of finding a coin that flipped heads ten times in a row somewhere among a million coins flipped. The size, value, and momentum effects have partially answered this concern by surviving out-of-sample tests, in markets and time periods the original researchers did not use to discover the pattern, which is a meaningfully higher bar than simply fitting one historical dataset well. Not every academically documented anomaly has cleared this bar, and part of the ongoing work in empirical asset pricing is separating the patterns that represent genuine, persistent risk compensation from those that were, in retrospect, closer to statistical accidents.
The shift toward multifactor thinking
The turning point came in 1992 and 1993, when Eugene Fama and Kenneth French published research showing that a three-factor model, adding a size factor and a value factor alongside the original market factor, explained a substantially larger share of the cross-section of average stock returns than beta alone. The intellectual significance went beyond the specific factors chosen: it validated a broader shift in how the field approached asset pricing, from searching for the single correct model to building and comparing increasingly rich multifactor descriptions of the data, with subsequent research adding profitability, investment, and momentum factors on top of the original three.
This did not amount to a wholesale rejection of CAPM's logic. The market factor remains, in nearly every multifactor model built since, the single largest and most consistently priced source of systematic risk; what changed was the recognition that it is necessary but not sufficient to explain the full pattern of average returns observed in real markets.
The debate over interpretation continues to this day, and it splits roughly along two lines. One camp reads the size, value, and momentum premiums as compensation for genuine, additional dimensions of systematic risk that a single market beta simply fails to capture, small and cheaply valued companies, on this view, are riskier in ways that matter to investors, more exposed to credit and liquidity shocks, more fragile during recessions, even if that extra risk does not show up cleanly in their historical correlation with a broad stock index. A second camp reads the same premiums as evidence of persistent behavioral mispricing, investors systematically underpricing unglamorous, cheap, out-of-favor companies and overpricing exciting, popular ones, a gap that arbitrage has not fully closed because exploiting it requires patience and a tolerance for looking wrong for uncomfortably long stretches. Both camps agree on the empirical pattern; they disagree, sometimes sharply, on whether it represents a fair price for risk or a correctable market inefficiency.
What this debate means for a real portfolio
For an individual investor, the practical upshot of fifty years of testing is not that beta is useless, but that it is incomplete. A portfolio built entirely around chasing high-beta stocks in the hope of proportionally higher return is working against one of the most replicated findings in the empirical literature, the flattening of the realized security market line. Conversely, an investor tilting deliberately toward smaller, cheaper-valued companies is leaning into patterns with genuine, if imperfect and not guaranteed, historical support, though these tilts also come with periods of significant underperformance lasting a decade or longer, a fact the underlying academic papers document as clearly as the long-run premium itself.
Actionable breakdown
- Do not chase beta as a return strategy on its own.
- High beta has not delivered proportionally higher realized returns historically.
- Take documented factor premiums seriously, but skeptically.
- Size and value tilts have decades of out-of-sample support.
- Expect long stretches, sometimes a decade, where a factor underperforms.
- Remember every test has a blind spot.
- No test can fully rule CAPM in or out, per Roll's critique.
- Favor diversified, low-cost exposure over single-factor bets.
- A broad index fund already captures the market factor efficiently.
Common pitfalls
A common error is treating the flattened security market line as proof that risk does not matter, when the correct reading is that beta alone is an incomplete measure of risk, not that risk is unpriced. A second pitfall is forgetting Roll's critique and citing a single empirical study as definitive proof CAPM is dead, when the same study's conclusions depend heavily on the specific market proxy and time period chosen. A third pitfall is assuming a documented factor premium, size or value, will persist unchanged forever simply because it appeared in decades of past data; several well-known premiums have narrowed or reversed for extended periods since being popularized and widely traded.
The bottom line
Fifty years of testing confirmed that market risk matters but is not the whole story, pushing academic finance from a single-factor model toward richer multifactor descriptions of what actually drives stock returns.
Related reading: factor investing, stock analysis, the core CAPM formula and mechanism, multifactor models, an overview, the Fama-French three-factor model.