Evidence before money
Two honest halves (a deterministic backtest report and the forward-only live track record), plus what each does and does not prove, and how to export, print or email them.
- Who it is for
- Everyone on the book
- Reading time
- 5 min read
- Updated
Before anyone funds a book, Bellwether shows two kinds of evidence and never blends them: a backtest of the rule-based stack that code alone computes, and the live record of what the whole system (the PM agent included) has actually done in this book. The PM agent is not backtested: it is a language model that has read every one of those years, so an in-window "backtest" of it could not tell skill from memory. Its only honest record is forward.

Two halves, kept apart
- Backtest evidence (Performance › Backtest evidence) replays every rule-based and your over the last five years with the same that runs live: whole shares, the per-order cap, the daily-loss and halts, orders per day, costs on every fill. It is reproducible to the cent from the seed and the dataset fingerprint printed on the page.
- Live track record (Performance › Track record) is what the system has done in this book since inception, day by day, by proposal. Nothing is replayed; nothing is chosen with hindsight.
Reading the backtest page
- Full window: the compound growth rate, volatility, Sharpe ratio and drawdown, next to SPY and a 60/40 mix over the same days with the same money.
- Sharpe 5–95%: the range the Sharpe plausibly lands in if the same years were shuffled and replayed thousands of times. An interval that reaches below zero cannot rule out "no edge at all".
- Walk-forward, out of sample: for each algorithm, parameters picked on the past only, then scored on the following year; the stitched line is the honest one. The blend has no such line: nothing about it was chosen by the machine, so its year-by-year panel is the same configuration restarted each year and is labelled "not out of sample". The blend's weights, and the defaults every grid is centred on, were chosen by people who had seen these years; no statistic on the page can undo that.
- Deflated Sharpe / P(overfit): how much of the best cell's Sharpe is selection luck, counting only the grid cells this report tried. Below about 0.95 / above 0.5 respectively, treat the cell as noise. A value marked † (or "P(Sharpe > 0)") had one counted trial: nothing was deflated.
- Costs ×3: the same run with spread and impact tripled. If the edge vanishes there, it lived inside the cost assumption.
- Lookahead probe: every price after a cut date is replaced by noise and each strategy is rerun; a decision before the cut that changes means the strategy read the future. The page prints the result per algorithm, and runs a strategy that cheats on purpose through the same probe to show it gets caught.
- Do the knobs matter? Each console setting (blend weights, algorithm modes, , individual caps, sizing/breadth/cadence) moved on its own from your baseline. The bar is the Sharpe difference with its interval; within noise means five years cannot tell better from worse, dead here means the setting changed nothing at all in the rule-based stack (it may still act on the PM agent live; the row says what), not modelled means a backtest cannot see it by construction.
- Same rules, different balance: the limits are dollars, not percent: a $150 daily-loss fires on ordinary red days for a $100k book, and a $500 per-order cap cannot buy one share of a $700 stock in a $10k book. The table shows exactly how often each rule bit.
- What this proves / does not prove: read it. The universe is today's mega-caps (survivorship-biased upward), fills are at the open plus a stated spread, dividends are frictionless, and no operator, or is modelled.

Does the PM add anything?
On the track-record page, PM vs the blend alone puts three lines on the same money and days: what the rule-based blend would have done on its own under your weights and caps (simulated on the committed bars), what the book actually did, and the book with operator-initiated trades taken out (the PM alone). Underneath: PM minus blend in basis points, per session and in total, with an interval. Read the verdict literally: "not yet" and "not significant" mean exactly that; only "has added … interval excludes zero" is a claim, and it needs at least twenty sessions.
Sharing it
An owner or operator can share a book read-only from the account menu (People & access → Add someone → Read only) or from the Share read-only button on the track-record page. A viewer sees both evidence pages and every number in the console, and can act on none of it.
Reports, exports and email
Every evidence page has PDF, CSV, Print and Email this…. The PDF is rendered by the server (the same file an email attaches); CSVs carry the raw rows (proposals with outcomes, round trips, daily equity, backtest metrics and curves). The monthly statement is one calendar month: start and end equity, flows, what trading made after model spend, positions at the last close, every round trip.
Email this… sends one message per recipient with the PDF attached; people who can already open the book also get a read-only link, everyone else gets the PDF only. An email never widens who can open the console. Sending asks for your , is written to the audit trail, and is limited to 20 recipient-sends per book per hour. Scheduled reports sends the weekly letter on Fridays after the close and the monthly statement on the 1st to up to ten addresses; adding an address asks for your code, switching a report off does not. Viewers and read-only roles can download but never see recipient addresses.
Lab and simulation results are not evidence
Anything the Lab or the simulation harness produces is a behaviour test, not evidence of returns: it checks that the machinery does what it should on the days it is given. Returns are claimed only on the two pages above.