Performance Metrics

Lesson 16 of 20, about 16 minutes

What you will learn

  • Use return and risk-adjusted metrics to evaluate a strategy
  • Compute the profit factor
  • Understand why win rate alone deceives
  • Know that metrics on an overfit backtest are fiction

Once a strategy has been backtested or traded, its performance must be evaluated, and doing so rigorously requires the right metrics. This lesson brings the risk-adjusted measures of Unit 6 into the quantitative-trading context and adds several metrics specific to evaluating strategies. The recurring theme is that no single number tells the whole story, and that metrics computed on overfit backtests are fiction, however good they look.

The Sharpe ratio explained by a quant trader (Wall Street Quants)

A quant's take on the Sharpe ratio for evaluating strategies. Focus on return per unit of risk.

Return and risk-adjusted metrics

The most basic measures are returns: the total return over a period and the annualized return, which expresses performance on a yearly basis for comparison. But raw return is meaningless without accounting for the risk taken to achieve it, the central lesson of Unit 6. This is why risk-adjusted metrics are essential, chief among them the Sharpe ratio, which measures excess return per unit of volatility and remains the standard yardstick for comparing strategies. The Sortino ratio, which penalizes only downside volatility, and the Calmar ratio, which measures return relative to maximum drawdown, offer complementary risk-adjusted views, each defining risk differently as Unit 6 explained.

Formula
Profit factor = gross profit / gross loss
  • gross profit = total of all winning trades
  • gross loss = total of all losing trades (as a positive number)

Key terms

Sharpe ratio
Excess return per unit of volatility, the standard risk-adjusted yardstick.
Maximum drawdown
The largest peak-to-trough decline, the worst loss endured.
Profit factor
Gross profit divided by gross loss, above 1 means the wins outweigh the losses.
Win rate
The percentage of trades that are profitable. Misleading on its own.
Worked example

Profit factor and win rate together

A strategy's winning trades total 18,000 dollars and its losing trades total 12,000 dollars, from a mix of many trades. What is the profit factor?

  1. Apply the formula. Gross profit over gross loss is 18,000 / 12,000.
  2. Compute. That is 1.5.
Result: The profit factor is 1.5.

Why it matters: A profit factor above 1 means the strategy made more than it lost. But this says nothing about win rate: it could be many small wins and few big losses, or the reverse. Metrics must be read together.

Calculation

Compute a profit factor

A strategy's gross profit is 30,000 dollars and its gross loss is 20,000 dollars. What is its profit factor?

Need a hint?

Profit factor = gross profit / gross loss.

Drawdown metrics

Drawdown metrics, from Unit 6, capture the downside pain a strategy inflicts. The maximum drawdown, the largest peak-to-trough decline, measures the worst loss an investor would have had to endure, while the drawdown duration measures how long the strategy stayed underwater before recovering. These metrics matter a lot because, as Unit 6 stressed, large drawdowns test discipline and can force abandonment at the worst moment, and the asymmetry of recovery means deep drawdowns hurt more than their size suggests. A strategy with an attractive return but a brutal maximum drawdown may be untradeable in practice, because no one could stick with it through the pain.

Trade-level metrics and their subtleties

  • The win rate, the percentage of trades that are profitable, is intuitive but can be deeply misleading on its own, since a high win rate can mask rare catastrophic losses.
  • The profit factor, gross profits divided by gross losses, summarizes how much is won relative to how much is lost.
  • The average win compared to the average loss reveals the payoff structure, showing whether the strategy makes large gains and small losses or the reverse.
  • The number of trades matters for statistical significance, since, by the law of large numbers from Unit 5, a strategy needs enough trades for its metrics to be meaningful rather than a product of a lucky small sample.

Win rate versus payoff

A crucial subtlety is the relationship between win rate and payoff structure, which reveals why a single metric deceives. A strategy can have a low win rate yet be highly profitable if its occasional wins are large and its frequent losses are small, the characteristic profile of trend-following and momentum strategies that lose often in small amounts but win big on the rare strong trend. Conversely, a strategy can have a high win rate yet be ruinous if its frequent small wins are punctuated by rare enormous losses, the dangerous profile of naked option selling from Unit 7, picking up pennies in front of a steamroller. The win rate alone tells you nothing about which of these a strategy is, only by examining the win rate together with the average win and average loss does the true character emerge.

A high win rate can hide a catastrophe, and a low win rate can hide a fortune. No single metric reveals a strategy's true character.

Evaluating performance with the Sharpe ratio in Python (compsci)

Computes risk-adjusted performance in code. Reinforces reading metrics as a panel, not one number.

Decision scenario

The deceptive win rate

Strategy A wins 90 percent of its trades, Strategy B wins only 40 percent. Which is more profitable?

The honest use of metrics

The disciplined evaluation of a strategy rests on several principles that echo the skepticism of this entire unit. No single metric tells the whole story, so a strategy must be judged on a panel of measures spanning return, risk-adjustment, drawdown, and trade structure, never on one cherry-picked number. Risk-adjusted and drawdown metrics generally matter more than raw return, because they reveal the risk behind the performance. And most importantly, metrics computed on an overfit backtest are fiction: a spectacular Sharpe ratio or win rate from a strategy that has been overfit to historical data tells you nothing about future performance, because the underlying results are an artifact of fitting noise. The Sharpe ratio of a backtest is not the Sharpe ratio you will achieve live. Performance metrics are valuable tools for evaluation, but only when applied to honestly derived results and read as a complete picture, with the same vigilance against self-deception that governs every other part of quantitative trading.

Reflection

A panel, not a number

In your own words, explain why a strategy must be judged on a panel of metrics, and why an impressive Sharpe ratio from a backtest can be meaningless.

Write an answer before comparing it with the model response.

Quiz

This lesson ends with a 5-question quiz. Create a free account or sign in to take it, save your progress and earn points.