Skip to content
Lab & evidence

Backtests versus the track record

What the Lab's four tabs do, how queued backtests and sweeps carry their own honesty numbers, and why the live record on Performance is the number that counts.

Who it is for
Everyone on the book
Reading time
4 min read
Updated

Lab is a safe place to ask "what if". It places no trades, changes no setting and accepts no configuration: the only things you can carry out of it are a saved scenario and a draft you choose to open on Strategy › Guidance. The line under the title shows your book and the two limits that it, so every result is read against the same rules that run live.

The four tabs

TabThe question it answers
Stress testIf the market fell tomorrow, what would my book lose, and would the system halt?
BacktestsHow would this mix of have done since 2006, honestly measured?
ReplayRe-run a past stretch of history against today's rules; where would it have stopped me?
ProjectResample the next months many times: odds of a loss, of a halt, and the AI cost
Lab › Stress test: move the market by a chosen amount and read what the book would lose, position by position.
Lab › Stress test: move the market by a chosen amount and read what the book would lose, position by position.

Stress test applies a what-if move to your current positions: the whole market, one name, a sector, or wider swings. It shows today after the move, the room left before a halt, how far under the book's best-ever value it would sit, and which positions would be through their stop or their . Results are instant.

Replay walks today's book day by day through a real episode with the same halt rules. It shows what holding through would have done, what selling at the halt would have done, the difference, and the first day the system would have stopped trading.

Project resamples real daily moves into hundreds or thousands of paths over 3 to 24 months, with halts applied on every run and AI cost taken off. A compare switch runs two rule sets side by side.

Backtests: queued, persisted, honest

The Backtests tab has a New run builder and a table of past runs. Three kinds:

Lab › Backtests: queued and finished runs with their headline numbers; the builder and the comparison of two runs open from here.
Lab › Backtests: queued and finished runs with their headline numbers; the builder and the comparison of two runs open from here.
  • Backtest: one engine run of the you choose (any registered algorithm with a weight, the universe, the window, costs, sizing, rebalance and the limit profile every simulated order must pass).
  • Sweep: vary one or two parameters over a grid and run one backtest per cell. The parent reports the grid, the best cell, and the honesty numbers for having tried N things.
  • Sim: replay the PM (shown as on ) day by day through the production . Scripted is free; the Claude-driven variant needs the model key on the worker and a confirmed spend.

Runs are queued for a worker; leave and come back. Every result carries an Honesty panel, the numbers that argue against it: Deflated Sharpe (the probability the true Sharpe beats what the luckiest of N skill-less configurations would show; below about 0.95 the result is not distinguishable from selection luck), a hold-out split of the first 70% against the last 30% of sessions, a one-year bootstrap of outcomes against SPY, and for sweeps the probability of overfitting and a walk-forward curve where each year's cell was chosen on the prior years only. Pick two runs to compare them side by side.

Nothing you find here changes the blend. To act on it, go to Algorithms and move the weights yourself.

The track record

Performance is the other side of the ledger: what the book actually did in paper, after AI cost if you flip the switch. It chains daily returns (deposits stripped out) into an equity line against the same money in SPY and a 60/40 mix, draws the underwater chart of drawdowns, grids returns by month, and keeps the Trade journal of every round trip with win rate, expectancy per trade and profit factor. The equity line starts on the first funded day; the journal, win rate and breakdowns start with the first position bought and fully sold. There is nothing to flatter before then.

Performance › Track record: the live, forward-only record of what the desk proposed and what happened.
Performance › Track record: the live, forward-only record of what the desk proposed and what happened.

Research adds a second live measure: every brief's checkable calls are scored from daily closes at the horizon, with no model grading, and Research shows hits over scored calls per brief (no combined hit rate across briefs yet).

A backtest tells you what a rule would have done on the winners of the last 20 years. The track record tells you what this system did, with this data, this month. When the two disagree, believe the second. See what counts as proof.