EMPIRICAL EVIDENCE ON SECURITY RETURNS

What Decades of Testing the CAPM and APT Actually Found

A pricing model can look elegant on a whiteboard and still buckle when checked against decades of real returns. This article works through how those tests are actually run and what the honest, decades-long verdict on beta-based pricing has been, because it determines how much weight to put on any fund's "risk-adjusted" performance claim.

Advanced13 min readUpdated 2026

The testing problem

Both the Capital Asset Pricing Model and Arbitrage Pricing Theory make the same basic claim: expected return depends only on a security's sensitivity (beta) to one or more priced risk factors, not on anything else about the security. That claim sounds simple to check. It is not. The true market portfolio that CAPM theorizes about includes every risky asset in existence, human capital, real estate, private businesses, art, none of which trade on a public exchange with observable prices. Every real test substitutes a proxy, usually a broad public stock index, for that unobservable true market. This creates a structural problem: if a test rejects the model, you cannot tell whether the model itself is wrong or merely the proxy used to stand in for the market was a poor one. Multifactor extensions like APT sidestep some of this by not requiring a single "true" market portfolio, but they introduce a different problem: the factors themselves (size, value, momentum, and others) are typically identified by sifting through the same historical data being used to test them, raising the risk of finding patterns that fit the past by chance rather than by economic law.

There is a second, deeper layer to the testing problem, sometimes called the joint hypothesis issue. Any statistical test of an asset pricing model is simultaneously testing two things at once: whether the model's economic logic is correct, and whether the specific proxy chosen to represent its factors was the right one. If a test rejects a pricing relationship, that rejection could mean the theory itself is wrong, or it could mean the theory is right but the researcher picked a flawed stand-in for the true factor. There is no clean statistical way to separate these two possibilities using a single test, which is why empirical asset pricing has never produced a single, universally accepted verdict of "CAPM is true" or "CAPM is false." Instead, the field has accumulated decades of tests under different proxies, different time periods, and different portfolio groupings, and drawn conclusions from the overall weight and consistency of that evidence rather than from any one study.

Key idea Rejecting a beta-based model in a test does not prove beta is meaningless. It proves that beta, measured against a specific proxy, over a specific period, failed to fully explain returns in that sample. The distinction matters more than it sounds.

The two-stage test, worked through

The standard empirical test runs in two stages. First, a time-series regression estimates each portfolio's beta against the chosen factor, using historical returns. Second, a cross-sectional regression takes those estimated betas as the input variable and each portfolio's average return as the output variable, checking whether portfolios with higher betas actually earned higher average returns, and whether the resulting slope matches the theoretical risk premium.

Suppose five diversified portfolios were sorted into beta buckets, and their estimated betas and average monthly excess returns over a sample period were: 0.6 and 0.30%; 0.8 and 0.45%; 1.0 and 0.55%; 1.2 and 0.62%; and 1.5 and 0.70%. Running a simple regression of average return on beta across these five points: the mean beta is 1.02 and the mean excess return is 0.524%. Summing the cross products of each point's deviation from its mean gives 0.2116, and summing the squared beta deviations gives 0.488. The slope, which should equal the market risk premium if CAPM holds exactly, is 0.2116 ÷ 0.488 = 0.434% per month (about 5.2% annualized). The intercept, which theory says should be zero, works out to 0.524% minus (0.434% × 1.02) = 0.082% per month. That small but positive intercept is the textbook signature of the "flat security market line" finding: low-beta portfolios earn more than a pure beta story predicts, and the line connecting risk to return is shallower and higher-starting than theory says it should be.

The second worked example shows why a single cross-sectional regression, like the one above, is never treated as conclusive on its own. Real tests, known as Fama-MacBeth-style procedures, run this cross-sectional regression separately in every month of the sample, then average the resulting slope across all months, treating the standard deviation of those monthly slopes as the basis for a statistical test. Suppose the estimated market risk-premium slope came out to 0.6% in month one, negative 0.2% in month two, and 0.9% in month three. The average is (0.6 − 0.2 + 0.9) ÷ 3 = 0.43%. The sample standard deviation of those three estimates is about 0.57%, so the standard error of the average, dividing by the square root of 3, is roughly 0.33%. The resulting t-statistic is 0.43 ÷ 0.33 ≈ 1.32, well below the roughly 2.0 threshold conventionally used for statistical significance. With only three months of data, this result proves essentially nothing either way, which is exactly why real published tests use decades of monthly observations, hundreds of data points, before drawing conclusions.

It is worth walking through why averaging across many months, rather than trusting any single month's cross-sectional regression, is the right procedure at all. In any given month, the estimated slope reflects a mix of the true underlying risk premium plus that month's idiosyncratic noise, driven by whatever news, surprises, or shifts in sentiment happened to move markets over those particular weeks. A single month's slope, whether it comes out unusually high, unusually low, or even negative, tells you almost nothing on its own. Averaging across many independent months lets the noise partially cancel while the persistent underlying signal, if one exists, accumulates, which is exactly the logic behind using a t-statistic built from the mean and standard deviation of many monthly estimates rather than trusting any single month's regression output.

What decades of tests found

Extending this procedure across many decades of U.S. stock data produced a consistent pattern, echoed across multiple independent studies using different sample periods and portfolio groupings: the security market line implied by the data is flatter than the theoretical line, with a smaller-than-predicted slope and a larger-than-zero intercept, similar in shape to the small illustrative example above. High-beta stocks have, on average, underperformed what CAPM would predict, and low-beta stocks have outperformed it. This single pattern has spawned an entire line of research into low-volatility investing.

Multifactor tests fare better on pure statistical fit. Adding size and value factors to a plain market-beta regression typically raises the share of cross-sectional variation in average returns that the model explains, sometimes from roughly 70% under a single-factor model to 90% or higher once several factors are included. But higher statistical fit is not the same as economic proof: some of that improvement plausibly reflects factors that were identified partly because they fit the historical sample well, meaning their out-of-sample performance has in many cases been weaker and more erratic than their in-sample track record suggested. Momentum and value premia, for instance, have both experienced extended multi-year stretches of underperformance even while remaining statistically significant across the full long-run sample.

Researchers have also documented that the specific set of test portfolios used matters more than intuition suggests. Early tests often sorted stocks into portfolios by industry, and found relatively weak support for beta as a return predictor. Later tests, sorting instead by characteristics like size and book-to-market ratio, found much stronger patterns, partly because those characteristics happen to spread average returns out more widely across the resulting portfolios, giving any candidate factor model more variation to actually explain. This sensitivity to test design is itself an important, if uncomfortable, empirical finding: it means a model's apparent success or failure can depend meaningfully on how the researcher chose to slice the data, which is one more reason no single published test should be read as the final word.

Key idea A model that explains 90% of average returns in a backward-looking sample is not guaranteed to explain the next decade's returns nearly as well. Fit and forecast accuracy are related but distinct properties.

Using this in a real portfolio

For a working investor, the practical implication is not that beta is useless, it clearly still correlates with realized returns on average, but that treating any single-factor risk model as a precise pricing formula overstates its reliability. Portfolio managers who report a fund's "alpha" relative to CAPM are implicitly claiming the single market factor fully captures risk, a claim the data only partly support. A fund with a large positive CAPM alpha may simply be tilted toward value stocks, small stocks, or low-beta stocks, exposures that a multifactor model would price directly rather than crediting to manager skill. Before accepting a stated alpha at face value, it is worth asking which benchmark and how many factors were used to compute it.

This has a direct, practical corollary for anyone comparing fund performance reports across providers. Two funds with genuinely identical underlying holdings can be presented with very different headline alpha figures purely because one report benchmarks against a plain market index and the other benchmarks against a multifactor model, since the multifactor benchmark absorbs return that the single-factor benchmark would have credited to the manager. Neither number is dishonest, exactly, but neither is complete on its own; the useful habit is to ask which benchmark produced a stated alpha before treating it as evidence of manager skill, and to be especially cautious of marketing materials that quietly switch benchmarks between the return chart and the risk-adjusted performance table.

Actionable breakdown

  • Reading a factor-model test
    • Check whether the test is in-sample or out-of-sample.
    • Note how many factors are used before trusting alpha.
    • Ask what proxy stood in for the market portfolio.
  • Interpreting a flat security market line
    • Expect low-beta assets to modestly outperform CAPM.
    • Expect high-beta assets to modestly underperform CAPM.
    • Don't assume the CAPM slope is exactly the risk premium.
  • Evaluating a manager's claimed edge
    • Ask which multifactor benchmark produced the alpha figure.
    • Distrust alpha computed against a single-factor model only.
    • Require a track record spanning multiple market cycles.

Common pitfalls

Confusing a rejected model with a useless one: a flatter-than-predicted security market line still shows a positive relationship between risk and return; it just is not the exact line theory drew.

Overtrusting in-sample fit: a factor model built and tested on the same historical data will always look better than it performs going forward, since some of its apparent explanatory power comes from fitting noise specific to that sample.

Treating three or six months of data as evidence: as the Fama-MacBeth example shows, small samples produce wildly unstable slope estimates; meaningful statistical tests require years, not months.

Assuming a wider factor list is automatically better: each additional factor added to a model also adds estimation noise, and a model with too many factors relative to the sample size can fit historical data almost perfectly while explaining nothing reliable about the future.

The bottom line

Decades of empirical testing find that beta and other factor exposures predict average returns directionally but imperfectly, with a persistently flatter-than-theory relationship between risk and reward that every serious investor should build into their expectations, holding models as useful, working approximations rather than exact, immutable laws of markets.

The Capital Asset Pricing Model · Arbitrage Pricing Theory · A multifactor APT · The Fama-French three-factor model · Factor investing guide

All articles · The deep guides