Sharpe ratio: what it is, why we chase it, and why one good year proves nothing
The desk's scoreboard and goal (return over cash per unit of bounce): how Bellwether measures it with an honest interval, what a Sharpe of 1.0, 1.5 or 2.0 actually feels like in a year, and why it is never one of the signals.
- Who it is for
- Everyone on the book
- Reading time
- 8 min read
- Updated
The desk's mentor put it in one line: "you want a 2 Sharpe ratio." Bellwether treats that number as the scoreboard the whole desk is graded on and the objective the Quant desk optimises, and deliberately not as one of the trading . It has no weight in the and no rank among the . This article says what the number is, where you see it, how honestly it is measured, and what chasing 2.0 does and does not mean.
What it is
The Sharpe ratio is the return you earned above cash, divided by how much the returns bounced around on the way. A book that made 18% in a year while a Treasury bill paid 3%, with returns that swung about 12% a year, has a Sharpe of (18 − 3) ÷ 12 = 1.25. A manager who made 15% at 8% volatility scores (15 − 3) ÷ 8 = 1.5: less return, but much more return per unit of risk, and risk is the thing you can scale.
Rules of thumb: below 0.5 is indistinguishable from noise over any horizon you will live through; around 1 sustained for years is genuinely good; 2 sustained for years is rare and is the desk's goal; a backtest that prints 3 is almost always flattering itself.
Two cousins sit beside it on the scoreboard. Sortino counts only the down-swings as risk, so a book that only ever jumps upward is not penalised for it. Treynor divides the return over cash by the book's beta to SPY instead of by its total bounce; it asks whether the return came from anything other than simply holding the market.
Where you see it
- Performance › Backtest evidence and the Track record open with the risk-adjusted scoreboard: this book against SPY and a 60/40 mix over exactly the same funded days (Sharpe with its 95% interval, Sortino, max , Treynor and beta, the share of up months) plus two "what would it take" lines and the goal.
- Lab › Sharpe explorer models what a Sharpe of 1.0, 1.5 and 2.0 (and any target you type) would actually feel like over a year at your 's volatility, in percent and in dollars for your book.
- Strategy › Profile prints, under each posture preset, what the desk's measured Sharpe implies at that preset's volatility target: an estimate, labelled as one.
- Every backtest in Lab › Backtests and every algorithm card (for example Jim's) prints its Sharpe next to a 95% interval and a deflated Sharpe.
- Every Quant desk suggestion in the Inbox states its objective in one line: the expected change in out-of-sample Sharpe with an interval, the deflated Sharpe of the proposed mix, and the evidence run behind it.
How Bellwether measures it honestly
Monthly returns by default. The scoreboard compounds the book's daily returns into calendar months and annualises with √12. Measuring over longer intervals makes volatility look smaller and the Sharpe look bigger; a monthly basis with the count of months printed beside it cannot be gamed that way. The daily basis is one click away and says so.
Over cash, and it says which cash. The line under the table names the risk-free rate every Sharpe on the page was computed against: the 3-month US Treasury bill from FRED (series DTB3) when the nightly Quant run could fetch it, otherwise a configured constant; always labelled, never silent. Backtests and algorithm cards keep the convention they were built with (cash rate 0, so strictly an information ratio against cash) and say that too.
With an interval, always. A Sharpe measured on a handful of months is mostly luck, so every value carries a 95% interval from Lo's standard error, √((1 + SR²/2) / n), with n shrunk when consecutive months are correlated (momentum and mean-reversion books have streaky months that would otherwise inflate the figure; Lo 2002). Under three months of data the number is shown with an insufficient sample label and nothing on the page treats it as a judgement.
"What would it take." Under the table, two sentences say how many years of returns like these would be needed before the measured Sharpe sits clearly above zero, and above 1 (the minimum track record length of Bailey and López de Prado). A true Sharpe of 1 needs roughly three years of months to tell from luck; a Sharpe of 0.5 needs more than a decade. That is why one good year proves nothing.
Deflated for the number of tries. When a Lab sweep tries sixty configurations, the best of them would show a healthy Sharpe even if none had any skill. The deflated Sharpe is the probability the winner's true Sharpe beats what the luckiest skill-less try would show; below about 0.95 the winner is not distinguishable from selection luck (Bailey and López de Prado 2014). Sweeps carry their real trial count; a single card counts itself plus its sensitivity variants.
What 1.0, 1.5 and 2.0 feel like
The Sharpe explorer runs twenty thousand seeded one-year paths (the same inputs always give the same table) and reports, for each Sharpe: the expected return over cash, how often a year still loses money, the one-in-twenty bad year, the typical and the bad drawdown, and the share of losing months; then the same things in dollars for your book size. At 10% volatility a true Sharpe of 1.0 means about +10% a year over cash, a losing year roughly one year in six, and a typical worst dip near 8%; a true Sharpe of 2.0 at 15% volatility means about +35% a year, a losing year about one in thirty-five, and a typical worst dip near 9%. Even at Sharpe 2, more than a quarter of all months lose money. It is a model with normal, independent days (real returns have fatter tails and longer streaks, so real drawdowns run deeper) and the page says so.
Beside the table, what your signals imply estimates the Sharpe the desk's own signals could support if traded cleanly: how well the blend score has ranked the next session's returns (the information coefficient) times the square root of how many independent calls it makes in a year (breadth): Grinold's fundamental law. It is an optimistic ceiling from a short history, labelled an estimate, there to compare the goal against.
Why it is the objective, not a signal
An algorithm's weight is the only input to the blend score, a is the only thing that can block a name, and is cosmetic; nothing else (the strategy model, ADR-0011). Sharpe sits above all of that: it is how the result is judged. The Quant desk's weight and enablement suggestions are chosen to improve out-of-sample Sharpe (the walk-forward recommender picks the target mix by the Sharpe of years it had not seen) inside the code bounds on how far a weight may move per run and how much may change in total, and never touching a posture limit. Each suggestion's card prints the expected change in Sharpe with a paired-bootstrap interval (the same resampled days for before and after, so shared market noise cancels), the deflated Sharpe of the proposed mix, and the run that produced the evidence. When the interval straddles zero the card says the change is a nudge, not a proven improvement. Nothing applies itself outside the Quant desk's bounds, and every applied change is revertible.
The goal line
Each book has a target Sharpe (2.0 by default), edited in one place, the goal line of the scoreboard on Backtest evidence (0.5 to 3.0, audited). It is a yardstick only: it moves no limit, no weight and no trade. The sentence beside it says how far the measured value is from the goal and, when the goal sits inside the measured interval, that the data cannot yet tell them apart. Until the book has three months of history the posture outlook uses the goal as a stand-in and labels it "a goal, not a measurement".
Pitfalls the page guards against
- Longer intervals flatter. Quarterly or annual returns hide the bounce; the scoreboard defaults to monthly and prints the basis.
- Cherry-picked windows. The window is a named range (since inception, year to date, three months), the same for the book and both benchmarks, printed above the table.
- Streaks inflate it. Positively autocorrelated returns make √12 (or √252) overstate the annual figure; the interval uses the autocorrelation-adjusted sample size and the Lab prints the adjusted value.
- Fat tails. Sharpe treats a strategy that earns a little most months and occasionally loses a lot ("picking up nickels in front of a steamroller") too kindly. Read max drawdown and the bad-year column beside it, and Sortino when the two disagree.
- Many tries. The best of many backtests is lucky by construction; the deflated Sharpe and the trial count travel with every sweep.
- One good year. See "what would it take": at realistic Sharpe ratios, telling skill from luck takes years, not months. The 's limits, not the scoreboard, are what protect the book in the meantime.
Sources
Public references only: William F. Sharpe, "The Sharpe Ratio", Journal of Portfolio Management (1994); Investopedia, "Sharpe Ratio: Definition, Formula, and Examples" (the operator's reference, including the 1.25 and 1.5 examples above and the pitfalls list); Andrew W. Lo, "The Statistics of Sharpe Ratios", Financial Analysts Journal 58(4), 2002 (standard errors, serial correlation, time aggregation); David H. Bailey and Marcos López de Prado, "The Sharpe Ratio Efficient Frontier", Journal of Risk 15(2), 2012 (minimum track record length) and "The Deflated Sharpe Ratio", Journal of Portfolio Management 40(5), 2014; Richard C. Grinold, "The Fundamental Law of Active Management", Journal of Portfolio Management 15(3), 1989; Frank A. Sortino and Lee N. Price, "Performance Measurement in a Downside Risk Framework", Journal of Investing (1994); Jack L. Treynor, "How to Rate Management of Investment Funds", Harvard Business Review (1965). The risk-free series is the Federal Reserve Bank of St. Louis FRED "3-Month Treasury Bill Secondary Market Rate" (DTB3).