Sharpe, Treynor, and Jensen: Judging a Manager Correctly
Two managers with the same return can look identical on a performance chart and yet deserve completely different verdicts once risk is accounted for. The three conventional risk-adjusted measures, built for three different situations, can rank the same pair of managers in opposite order, and picking the wrong one produces the wrong answer.
- The core mechanism: three measures for three situations
- The math: when Sharpe, Treynor, and Jensen disagree
- A second example: M-squared in plain percentage terms
- What the evidence shows about manager rankings
- Choosing the right measure for your own situation
- Actionable breakdown
- Common pitfalls
- The bottom line
The core mechanism: three measures for three situations
Comparing raw returns across managers is misleading because two managers can produce an identical return while taking very different amounts of risk to get there, and a manager who simply took on more risk is not exhibiting more skill. The conventional toolkit fixes this with three distinct risk-adjusted measures, each answering a different question about how that manager's portfolio fits into an investor's broader financial life.
The Sharpe ratio, (portfolio return - risk-free rate) / portfolio standard deviation, measures return earned per unit of total risk, and it is the correct tool when the portfolio being evaluated represents an investor's entire wealth, since total risk, including diversifiable risk, is exactly what matters when there is nothing else in the portfolio to diversify it away. The Treynor measure, (portfolio return - risk-free rate) / portfolio beta, measures return per unit of systematic risk alone, and is the correct tool when the portfolio being evaluated is just one holding inside an otherwise diversified collection of assets, since any unsystematic risk it carries will be diversified away by the rest of the holdings and should not count against it. Jensen's alpha, portfolio return - [risk-free rate + beta × (market return - risk-free rate)], measures the raw excess return earned above what the capital asset pricing model would predict given the portfolio's beta, expressed directly in percentage points rather than as a ratio.
It is worth being explicit about why these three measures can point in different directions even when applied to identical return data. Sharpe divides by total standard deviation, which mixes systematic risk and diversifiable, portfolio-specific risk into a single number. Treynor divides by beta alone, discarding diversifiable risk from the calculation entirely. Jensen's alpha does not divide by any risk measure at all; it simply asks how much return was earned beyond what beta alone would predict, which makes it useful for comparing managers on a common percentage-point scale but silent on how much risk, total or systematic, was taken to earn that excess return. None of the three measures is more "correct" than the others in the abstract; each is correct for a specific evaluation question.
The math: when Sharpe, Treynor, and Jensen disagree
Consider two managers, both delivering an average return of 11% over a year in which the risk-free rate was 3% and the broad market returned 9%, for a market risk premium of 6%. Manager C runs a broad, diversified portfolio with a standard deviation of 14% and a beta of 1.2. Manager D runs a more concentrated portfolio carrying more diversifiable, stock-specific risk, with a higher standard deviation of 20% but a lower beta of 0.9.
Sharpe ratios: Manager C is (11% - 3%) / 14% = 8/14 = 0.571; Manager D is (11% - 3%) / 20% = 8/20 = 0.400. On total risk, Manager C is clearly superior. Treynor measures: Manager C is (11% - 3%) / 1.2 = 8/1.2 = 6.67%; Manager D is (11% - 3%) / 0.9 = 8/0.9 = 8.89%. On systematic risk alone, Manager D wins decisively. Jensen's alpha: Manager C is 11% - [3% + 1.2 × 6%] = 11% - [3% + 7.2%] = 11% - 10.2% = 0.8%; Manager D is 11% - [3% + 0.9 × 6%] = 11% - [3% + 5.4%] = 11% - 8.4% = 2.6%. Manager D wins on Jensen's alpha too.
The verdict flips entirely depending on which measure is used. If either manager represents your entire portfolio, Manager C's lower total volatility for the same return makes it the better choice by Sharpe ratio. If either manager is instead a single satellite holding inside an already diversified portfolio, Manager D's higher return per unit of undiversifiable risk, and its larger alpha, make it the better addition, because the rest of your holdings will absorb and diversify away the extra stock-specific risk that Manager D's higher standard deviation reflects.
A second example: M-squared in plain percentage terms
The Sharpe ratio is useful for ranking but awkward to interpret directly, since it is expressed as a ratio rather than a percentage return. The M-squared measure converts a Sharpe ratio back into a directly comparable percentage return by asking: if this manager's portfolio were leveraged up or down to match the market's own volatility, what return would it have delivered? The formula is M2 = risk-free rate + Sharpe ratio × market standard deviation.
Using the market's standard deviation of 15%, the same 3% risk-free rate, and the Sharpe ratios calculated above: Manager C's M-squared is 3% + 0.571 × 15% = 3% + 8.57% = 11.57%. Manager D's M-squared is 3% + 0.400 × 15% = 3% + 6.00% = 9.00%. Compared directly against the market's actual 9% return, Manager C, once volatility-adjusted to match the market, would have outperformed the market by 2.57 percentage points, while Manager D, volatility-adjusted the same way, would have matched the market exactly, delivering zero excess return in these terms.
This confirms and sharpens the Sharpe ratio finding from the first example: expressed as an actual, intuitive percentage figure rather than an abstract ratio, Manager C's total-risk-adjusted edge over the market is concrete and sizable, while Manager D, despite its superior Treynor measure and Jensen's alpha, would not have beaten the market at all once leveraged to the market's own risk level.
What the evidence shows about manager rankings
Studies comparing manager rankings across different risk-adjusted measures have consistently found that rankings are sensitive to the choice of measure, the benchmark used, and the time period examined, and that a manager ranked in the top decile by one measure is not reliably ranked in the top decile by another, particularly for managers whose portfolios carry meaningfully different levels of diversification from one another. This is not a flaw in any single measure; it reflects the genuinely different risk concept each measure is built to isolate.
The broader body of research on performance persistence, how well a manager's past risk-adjusted performance predicts future performance, has generally found the relationship to be weak over most of the performance distribution, with somewhat more persistence evident among a small subset of the very best and very worst performers than among the broad middle of the distribution. This finding, explored in depth elsewhere in the context of mutual fund and analyst performance, is a strong argument for treating any single period's risk-adjusted measure, however carefully calculated, as one data point rather than a definitive verdict on manager skill.
A related body of evidence examines how sensitive these measures are to the specific benchmark chosen for beta and alpha calculations. A manager evaluated against a broad market index may show a very different alpha than the same manager evaluated against a style-matched benchmark, such as a small-company or value-oriented index that better reflects the actual universe of securities the manager selects from. This sensitivity is part of the motivation behind multifactor evaluation approaches that use more than one benchmark simultaneously, since a single-factor comparison can attribute what is really a style tilt, holding more small or more cheaply valued companies than the broad market, to manager skill when it is actually just a systematic exposure the manager could have been given credit or blame for from the start.
Choosing the right measure for your own situation
The practical decision most investors actually face is whether a manager or fund under consideration will be the entire portfolio or one piece of a larger, already diversified one. An investor allocating their full retirement savings to a single actively managed fund should weigh that fund's Sharpe ratio heavily, since every bit of its volatility, diversifiable or not, is volatility the investor will actually experience. A professional adding a satellite allocation, perhaps a sector fund or a specialist strategy, alongside a core portfolio of broad index funds should weigh Treynor and Jensen's alpha more heavily, since the core holdings will absorb much of the satellite's stock-specific risk, leaving only its systematic contribution to worry about.
This distinction matters concretely when evaluating a specialist manager, a hedge fund allocation, or an actively managed sector fund pitched as a complement to an existing broad portfolio. Asking "what is this fund's Sharpe ratio" is often the wrong question entirely in that context; the better question is "what is this fund's Treynor measure and Jensen's alpha relative to what my core holdings already provide," since that is the risk the fund is actually adding to a portfolio that is not starting from zero.
Actionable breakdown
- Match the measure to the manager's role in your portfolio.
- Use Sharpe ratio when a fund represents your entire portfolio.
- Use Treynor and Jensen's alpha for a single satellite holding.
- Convert ratios into intuitive percentage terms when possible.
- M-squared restates Sharpe ratio as a directly comparable return.
- Compare M-squared directly against the market's own return.
- Treat any single period's measure as one data point.
- Manager rankings shift meaningfully across measures and periods.
- Performance persistence is weak outside the extreme tails.
- Ask the right question before allocating to a specialist fund.
- Ask what risk the fund adds to your existing core holdings.
- Do not judge a satellite holding on total-risk measures alone.
Common pitfalls
The most common pitfall is applying the Sharpe ratio to a fund that will only ever be one piece of a diversified portfolio, penalizing it for diversifiable risk that the rest of the portfolio will absorb and remove at no additional cost.
A second pitfall is comparing a fund's alpha to a poorly matched benchmark, which can make a genuinely mediocre manager look skillful, or a genuinely skillful manager look mediocre, simply because the benchmark's own risk characteristics do not resemble the fund being evaluated.
A third pitfall is treating short track records as statistically reliable. Sharpe ratios, Treynor measures, and alpha estimates calculated from a small number of return observations carry wide statistical uncertainty, and a handful of unusually good or unusually bad months can shift the calculated figure substantially without reflecting any real change in underlying skill.
A fourth pitfall is ignoring how a manager's beta itself was estimated. A beta calculated over a short or unusually calm historical window can understate a portfolio's true sensitivity to a subsequent market downturn, which quietly distorts both the Treynor measure and Jensen's alpha, since both depend directly on the accuracy of that beta estimate rather than on the manager's actual, forward-looking risk exposure.
The bottom line
Sharpe, Treynor, and Jensen's alpha can rank the same two managers in opposite order, so the question to answer first is not which manager performed better, but which measure actually fits the role that manager plays in your portfolio.
Related reading: understanding portfolio risk, stock analysis fundamentals, the capital asset pricing model, performance measurement for hedge funds, style analysis.