MUTUAL FUNDS AND OTHER INVESTMENT COMPANIES

The Right Way to Compare Two Mutual Funds' Returns

Comparing mutual funds by their headline one-year return is one of the most common mistakes retail investors make, because a bare return number says nothing about the benchmark it should be measured against, the risk taken to earn it, or whether the period examined was representative. A properly informed comparison looks very different from a ranked list of trailing returns.

Intermediate13 min readUpdated 2026

The core principle and mechanism

A raw return number is meaningless in isolation because it answers only one question, "how much did this fund make," while leaving unanswered the two questions that actually determine whether that return reflects skill: how much risk was taken to earn it, and how did that return compare to what a passive, low-cost alternative in the same market would have delivered over the identical period. Judging a fund correctly requires holding both of those factors constant before drawing any conclusion.

The first correction is benchmarking: measuring a fund's return against an index representing the same market segment and style the fund actually invests in, not a generic broad-market average that may bear little resemblance to the fund's actual holdings. A small-cap value fund should be measured against a small-cap value index, not the S&P 500, because a fund can look impressive next to a mismatched benchmark while quietly trailing the benchmark it is actually competing against.

The second correction is risk adjustment. Two funds can post identical returns while taking very different amounts of risk to get there, and a fund that earned its return by taking on substantially more volatility has not necessarily demonstrated skill, only a willingness to accept more uncertainty. Beta, a measure of how much a fund's returns move relative to its benchmark, with a beta of 1.0 meaning it moves roughly in line with the market, is the standard starting point for adjusting a return for the risk taken to produce it.

Key idea A fund returning 11% sounds attractive on its own, but tells you nothing until you know what its benchmark returned, and what level of risk the fund carried, over the same exact period. Return without context is not evaluable.

The math: alpha and expected return

Example 1, computing alpha. Alpha measures a fund's return in excess of what its risk level alone would predict, using the formula alpha = actual return minus (risk-free rate plus beta times (market return minus risk-free rate)). Suppose the risk-free rate is 4%, the fund's benchmark market returned 13% over the period, and a fund with a beta of 1.0 actually returned 11%. Expected return given that risk level is 4% plus 1.0 times (13% minus 4%), which is 4% plus 9%, or 13%. The fund's actual return of 11% falls short of that 13% expectation by 2 percentage points, so the fund produced negative 2% alpha, meaning it underperformed what its own risk level alone would have predicted, despite an 11% absolute return that might look respectable in isolation.

Example 2, a higher-beta fund that still underperforms on a risk-adjusted basis. A different fund carries a beta of 1.3, meaning it tends to move 30% more than the market in either direction, and posted a 16% return over the same period, with the risk-free rate still at 4% and the market still at 13%. Expected return given that higher risk level is 4% plus 1.3 times (13% minus 4%), which is 4% plus 11.7%, or 15.7%. The fund's actual 16% return just barely clears that higher bar, producing alpha of only 0.3%, a nearly negligible risk-adjusted edge despite a headline return five percentage points above the first fund's. The higher-beta fund's larger absolute return mostly reflects the extra risk it carried, not meaningfully superior skill.

Example 3, the effect of the measurement window. A fund posted returns of 4%, 6%, 3%, 38%, and 5% across five consecutive years, driven by one exceptional year in the middle of that stretch. The simple five-year average is (4 plus 6 plus 3 plus 38 plus 5) divided by 5, which is 56 divided by 5, or 11.2% a year, a figure that looks strong and is technically accurate, but is dominated almost entirely by a single outlier year; excluding that one year, the fund averaged only 4.5% across the other four, a far more representative picture of what an investor entering after that exceptional year should actually expect going forward.

Example 4, the Sharpe ratio as a second risk-adjusted lens. Alpha adjusts for market risk specifically, but the Sharpe ratio, calculated as (fund return minus risk-free rate) divided by the fund's standard deviation, offers a complementary view based on the fund's total volatility rather than just its correlation with the market. Fund X returns 10% with a standard deviation of 12%, and the risk-free rate is 4%: its Sharpe ratio is (10% minus 4%) divided by 12%, or 6 divided by 12, which is 0.50. Fund Y returns a higher 13% but with a standard deviation of 22%: its Sharpe ratio is (13% minus 4%) divided by 22%, or 9 divided by 22, which is roughly 0.41. Despite Fund Y's higher absolute return, Fund X delivered more return per unit of total risk taken, a distinction a raw return comparison alone would completely miss.

What the evidence and market history show

Broad, long-running academic and industry studies tracking actively managed fund performance across decades have consistently found weak persistence in outperformance: funds that beat their benchmark in one multi-year period show only a modest and inconsistent tendency to keep beating it in the following period, and the funds that outperformed most dramatically in one window are, if anything, somewhat more likely than average to underperform in the next, a pattern consistent with mean reversion around a manager's true long-run skill level, or lack thereof, once an unusually strong stretch has passed.

A related and well-documented finding is survivorship bias in fund performance databases. Fund families routinely close or merge away their weakest-performing funds rather than let them continue reporting poor results indefinitely, which means the universe of funds still available to examine at any point in time is a skewed, better-than-average sample of everything that was ever actually launched. Studies that correct for this bias by including the full historical universe, closed funds included, consistently find average active fund performance looks meaningfully worse than a survivorship-biased snapshot alone would suggest.

A further pattern worth noting: funds ranked in the top decile of their peer category over a trailing period frequently rank in the top decile purely by virtue of a stylistic or sector tilt that happened to be in favor during that specific window, rather than by any repeatable process, and such funds have historically shown a tendency toward below-median performance once that particular tilt falls out of favor.

The dispersion of outcomes within a fund category also tends to be wider than most investors expect, and that dispersion itself is informative. In categories where security selection genuinely matters and markets are less thoroughly analyzed, such as certain small-cap or emerging-market segments, the spread between the best and worst performing active funds in a given year can be large, which is sometimes cited as evidence that skilled managers can add value there. But a wide dispersion cuts both ways: it also means the penalty for picking a below-average manager in that same category is correspondingly large, and identifying the above-average manager in advance remains just as difficult as in any other category, so a wider opportunity for outperformance is inseparable from a wider risk of underperformance.

Key idea Ranking funds against their peer category, rather than against a passive benchmark, can make an entire category of underperforming active funds look competitive with each other while all of them trail the low-cost index they are effectively competing against once fees are included.

How it applies in real portfolios

Applying this correctly means resisting the pull of any single-year return ranking, however prominently displayed on a fund screener or a brokerage's "top performers" list, and instead pulling up rolling multi-year returns against the fund's actual benchmark, not a generic market index and not simply its peer category average. Most reputable fund research platforms display a fund's beta and standard deviation alongside its return, which is enough raw material to sanity check whether an attractive return came with a commensurate increase in risk.

For a long-term retirement portfolio, the practical implication is to weight the decision toward funds and strategies with genuinely long track records, ideally spanning at least one full market cycle including a meaningful drawdown, and to treat any single standout year, positive or negative, as one data point among many rather than the headline the fund's own marketing will likely present it as. A fund's fee level, which is knowable in advance and does not require any statistical adjustment to interpret, remains a more reliable input to a selection decision than a trailing performance ranking that may reflect nothing more than a favorable few years.

High earners with taxable accounts alongside their retirement accounts have an additional reason to weight risk-adjusted, benchmark-relative evaluation heavily: a fund that appears to justify its higher fee on a pretax basis can lose that justification entirely once the after-tax comparison from turnover and distributions, covered separately, is layered on top of a marginal alpha edge that was already thin before taxes.

Actionable breakdown

  • Compare a fund's return only to its specific, style-matched benchmark.
  • Check beta before judging whether a return reflects skill or extra risk.
  • Calculate alpha, not just raw return, when comparing similar funds.
  • Examine rolling multi-year returns, not one selected window.
  • Ask whether one exceptional year is driving an otherwise average average.
  • Remember peer-category rankings can hide underperformance versus an index.
  • Weight fee level heavily; it needs no statistical adjustment to interpret.

Common pitfalls

The most common pitfall is judging a fund by its most recent one-year return, a period far too short to distinguish genuine skill from ordinary noise, and one that fund marketing materials are specifically designed to highlight when it happens to be favorable.

A second pitfall is comparing a fund only to its peer category rather than to a true passive benchmark, allowing a fund to look competitive by outranking similarly underperforming peers while still trailing the index it is effectively being paid to beat.

A third pitfall is ignoring survivorship bias entirely, drawing conclusions about "how active management performs" from only the funds still in existence today, a sample that has already had its worst performers quietly removed.

A fourth pitfall is relying on a single risk-adjusted metric in isolation. Alpha, Sharpe ratio, and simple relative return each capture a slightly different slice of performance, and a fund that looks strong on one measure can look mediocre on another, so cross-checking more than one metric before drawing a firm conclusion guards against being misled by whichever number happens to flatter a given fund most.

A fifth pitfall is failing to account for the specific market environment a fund's track record was built in. A fund concentrated in a style or sector that happened to be favored throughout its entire reported history has an untested record in the environments least favorable to that style, which only becomes visible once conditions actually shift.

Common mistake Sorting a fund screener by one-year or three-year trailing return and picking the top result is a near-guaranteed way to select for a fund that benefited from a recent, possibly non-repeatable tailwind rather than for durable skill.

The bottom line

A fund's performance only means something in the context of its specific benchmark, its risk level, and a time horizon long enough to smooth out noise, so never judge a fund by a single year's absolute return.

All articles · The deep guides · Stock analysis · Mutual Funds · Information on Mutual Funds · Alpha · Active management