What you will learn
- Use return and risk-adjusted metrics to evaluate a strategy
- Compute the profit factor
- Understand why win rate alone deceives
- Know that metrics on an overfit backtest are fiction
Once a strategy has been backtested or traded, its performance must be evaluated, and doing so rigorously requires the right metrics. This lesson brings the risk-adjusted measures of Unit 6 into the quantitative-trading context and adds several metrics specific to evaluating strategies. The recurring theme is that no single number tells the whole story, and that metrics computed on overfit backtests are fiction, however good they look.
The Sharpe ratio explained by a quant trader (Wall Street Quants)
A quant's take on the Sharpe ratio for evaluating strategies. Focus on return per unit of risk.
Return and risk-adjusted metrics
The most basic measures are returns: the total return over a period and the annualized return, which expresses performance on a yearly basis for comparison. But raw return is meaningless without accounting for the risk taken to achieve it, the central lesson of Unit 6. This is why risk-adjusted metrics are essential, chief among them the Sharpe ratio, which measures excess return per unit of volatility and remains the standard yardstick for comparing strategies. The Sortino ratio, which penalizes only downside volatility, and the Calmar ratio, which measures return relative to maximum drawdown, offer complementary risk-adjusted views, each defining risk differently as Unit 6 explained.
- gross profit = total of all winning trades
- gross loss = total of all losing trades (as a positive number)
Key terms
- Sharpe ratio
- Excess return per unit of volatility, the standard risk-adjusted yardstick.
- Maximum drawdown
- The largest peak-to-trough decline, the worst loss endured.
- Profit factor
- Gross profit divided by gross loss, above 1 means the wins outweigh the losses.
- Win rate
- The percentage of trades that are profitable. Misleading on its own.
Profit factor and win rate together
A strategy's winning trades total 18,000 dollars and its losing trades total 12,000 dollars, from a mix of many trades. What is the profit factor?
- Apply the formula. Gross profit over gross loss is 18,000 / 12,000.
- Compute. That is 1.5.
Why it matters: A profit factor above 1 means the strategy made more than it lost. But this says nothing about win rate: it could be many small wins and few big losses, or the reverse. Metrics must be read together.
Compute a profit factor
A strategy's gross profit is 30,000 dollars and its gross loss is 20,000 dollars. What is its profit factor?
Drawdown metrics
Drawdown metrics, from Unit 6, capture the downside pain a strategy inflicts. The maximum drawdown, the largest peak-to-trough decline, measures the worst loss an investor would have had to endure, while the drawdown duration measures how long the strategy stayed underwater before recovering. These metrics matter a lot because, as Unit 6 stressed, large drawdowns test discipline and can force abandonment at the worst moment, and the asymmetry of recovery means deep drawdowns hurt more than their size suggests. A strategy with an attractive return but a brutal maximum drawdown may be untradeable in practice, because no one could stick with it through the pain.
Trade-level metrics and their subtleties
- The win rate, the percentage of trades that are profitable, is intuitive but can be deeply misleading on its own, since a high win rate can mask rare catastrophic losses.
- The profit factor, gross profits divided by gross losses, summarizes how much is won relative to how much is lost.
- The average win compared to the average loss reveals the payoff structure, showing whether the strategy makes large gains and small losses or the reverse.
- The number of trades matters for statistical significance, since, by the law of large numbers from Unit 5, a strategy needs enough trades for its metrics to be meaningful rather than a product of a lucky small sample.
Win rate versus payoff
A crucial subtlety is the relationship between win rate and payoff structure, which reveals why a single metric deceives. A strategy can have a low win rate yet be highly profitable if its occasional wins are large and its frequent losses are small, the characteristic profile of trend-following and momentum strategies that lose often in small amounts but win big on the rare strong trend. Conversely, a strategy can have a high win rate yet be ruinous if its frequent small wins are punctuated by rare enormous losses, the dangerous profile of naked option selling from Unit 7, picking up pennies in front of a steamroller. The win rate alone tells you nothing about which of these a strategy is, only by examining the win rate together with the average win and average loss does the true character emerge.
A high win rate can hide a catastrophe, and a low win rate can hide a fortune. No single metric reveals a strategy's true character.
Evaluating performance with the Sharpe ratio in Python (compsci)
Computes risk-adjusted performance in code. Reinforces reading metrics as a panel, not one number.
The deceptive win rate
Strategy A wins 90 percent of its trades, Strategy B wins only 40 percent. Which is more profitable?
Win rate alone cannot determine profitability. Strategy A's 90 percent wins could be tiny while its rare 10 percent losses are enormous, the picking-up-pennies profile that can be ruinous. Strategy B's 40 percent wins could be large while its frequent losses are small, the profitable trend-following profile. Only by reading win rate together with the average win and average loss does the true character emerge.The honest use of metrics
The disciplined evaluation of a strategy rests on several principles that echo the skepticism of this entire unit. No single metric tells the whole story, so a strategy must be judged on a panel of measures spanning return, risk-adjustment, drawdown, and trade structure, never on one cherry-picked number. Risk-adjusted and drawdown metrics generally matter more than raw return, because they reveal the risk behind the performance. And most importantly, metrics computed on an overfit backtest are fiction: a spectacular Sharpe ratio or win rate from a strategy that has been overfit to historical data tells you nothing about future performance, because the underlying results are an artifact of fitting noise. The Sharpe ratio of a backtest is not the Sharpe ratio you will achieve live. Performance metrics are valuable tools for evaluation, but only when applied to honestly derived results and read as a complete picture, with the same vigilance against self-deception that governs every other part of quantitative trading.
A panel, not a number
In your own words, explain why a strategy must be judged on a panel of metrics, and why an impressive Sharpe ratio from a backtest can be meaningless.
Write an answer before comparing it with the model response.
Model answer
No single number captures a strategy's true character, because each metric reveals only one facet. Raw return ignores the risk taken, the Sharpe ratio adjusts for total volatility but not for skew or drawdown, win rate ignores the size of wins versus losses, and the number of trades affects whether any of these are even statistically meaningful. A high win rate can hide rare catastrophic losses and a low win rate can hide a fortune, so I have to read win rate alongside average win, average loss, drawdown, and a risk-adjusted ratio to see what is really going on. On top of that, any metric is only as honest as the results it is computed from: if the backtest was overfit to historical noise, its Sharpe ratio and win rate are artifacts of fitting that noise and tell me nothing about the future. The backtest Sharpe is not the live Sharpe. So I evaluate on a panel of measures, weight risk-adjusted and drawdown metrics over raw return, and only trust metrics that come from honestly derived, out-of-sample results.