Backtests versus the track record
What the Lab's four tabs do, how queued backtests and sweeps carry their own honesty numbers, and why the live record on Performance is the number that counts.
- Who it is for
- Everyone on the book
- Reading time
- 4 min read
- Updated
Lab is a safe place to ask "what if". It places no trades, changes no setting and accepts no configuration: the only things you can carry out of it are a saved scenario and a draft you choose to open on Strategy › Guidance. The line under the title shows your book and the two limits that it, so every result is read against the same rules that run live.
The four tabs
| Tab | The question it answers |
|---|---|
| Stress test | If the market fell tomorrow, what would my book lose, and would the system halt? |
| Backtests | How would this mix of have done since 2006, honestly measured? |
| Replay | Re-run a past stretch of history against today's rules; where would it have stopped me? |
| Project | Resample the next months many times: odds of a loss, of a halt, and the AI cost |

Stress test applies a what-if move to your current positions: the whole market, one name, a sector, or wider swings. It shows today after the move, the room left before a halt, how far under the book's best-ever value it would sit, and which positions would be through their stop or their . Results are instant.
Replay walks today's book day by day through a real episode with the same halt rules. It shows what holding through would have done, what selling at the halt would have done, the difference, and the first day the system would have stopped trading.
Project resamples real daily moves into hundreds or thousands of paths over 3 to 24 months, with halts applied on every run and AI cost taken off. A compare switch runs two rule sets side by side.
Backtests: queued, persisted, honest
The Backtests tab has a New run builder and a table of past runs. Three kinds:

- Backtest: one engine run of the you choose (any registered algorithm with a weight, the universe, the window, costs, sizing, rebalance and the limit profile every simulated order must pass).
- Sweep: vary one or two parameters over a grid and run one backtest per cell. The parent reports the grid, the best cell, and the honesty numbers for having tried N things.
- Sim: replay the PM (shown as on ) day by day through the production . Scripted is free; the Claude-driven variant needs the model key on the worker and a confirmed spend.
Runs are queued for a worker; leave and come back. Every result carries an Honesty panel, the numbers that argue against it: Deflated Sharpe (the probability the true Sharpe beats what the luckiest of N skill-less configurations would show; below about 0.95 the result is not distinguishable from selection luck), a hold-out split of the first 70% against the last 30% of sessions, a one-year bootstrap of outcomes against SPY, and for sweeps the probability of overfitting and a walk-forward curve where each year's cell was chosen on the prior years only. Pick two runs to compare them side by side.
Nothing you find here changes the blend. To act on it, go to Algorithms and move the weights yourself.
The track record
Performance is the other side of the ledger: what the book actually did in paper, after AI cost if you flip the switch. It chains daily returns (deposits stripped out) into an equity line against the same money in SPY and a 60/40 mix, draws the underwater chart of drawdowns, grids returns by month, and keeps the Trade journal of every round trip with win rate, expectancy per trade and profit factor. The equity line starts on the first funded day; the journal, win rate and breakdowns start with the first position bought and fully sold. There is nothing to flatter before then.

Research adds a second live measure: every brief's checkable calls are scored from daily closes at the horizon, with no model grading, and Research shows hits over scored calls per brief (no combined hit rate across briefs yet).
A backtest tells you what a rule would have done on the winners of the last 20 years. The track record tells you what this system did, with this data, this month. When the two disagree, believe the second. See what counts as proof.