Skip to content
Lab & evidence

Reading the Evidence page

One page that separates what is proven about the desk today (the machinery) from what only time can prove (the live record), with the sample and date behind every figure and a plain list of what none of it proves.

Who it is for
Everyone on the book
Reading time
8 min read
Updated

Evidence (reached from Performance's lead line, "Evidence: proven vs pending") is the page to open when someone asks "does it work?". It answers in three parts and never mixes them: what is proven about the desk today, what only time can prove, and what none of it proves. Every figure on the page carries the amount of history behind it and the date it was measured. A figure that does not have enough history yet says how far along it is instead of printing a number, and a figure that was computed but is not fair to publish says withheld and why.

Who can open it
Everyone signed in to the book, including people you shared it with as read only
What it reads
This book's own daily results, its rules twin if one exists, and the platform's committed test records
What it changes
Nothing. It is a read-only page
The one sentence
The box under the title: the date of the newest machinery record (each card carries its own) and how old the live record is

What is proven today

Four cards, each a fact about the machinery as of the date printed on the card, each re-computable from the platform's own committed records (we can re-run any of them in front of you). A card whose record is not loaded says so by name and states nothing.

  • Every idea tested stays on the record. Every the platform ever backtested is on a register, the dead ones included, measured on the stocks as they were listed at the time and after trading costs. The card says how many were measured, how many were retired, how many survive, and how many configurations were ever tried, because every extra try makes a lucky result more likely and the register corrects for it: the best edge anyone measured is printed with its deflated Sharpe, the figure that discounts all those tries (0.95 or more would make luck an unlikely explanation). The latest research round's sentence is built from the round's own counts (a base signal plus the additions judged against an equal-weight basket of the same names), so the day a signal survives the card says that instead.
  • Every role is examined. Each desk role (the PM, the , the , the Trader, , and the whole desk together) runs a competence suite: does it read what it must before it acts, use its tools correctly, handle a failed call or a stale price, stay grounded in the numbers it was given. Scripted rows measure the harness with a stand-in; the live row measures the real model on the same recorded day. When the live suite's own gate reads fail, the card says fail and names the check.
  • Readiness is a list of checks. The readiness ledger defines "safe on paper", "production quality" and "real money" as executable checks. The card shows how many pass today, how many are known gaps, how many are checked and failing today ("not yet"), and how many have no proof written yet, by area, and its sentence names each count in words. "No proof yet" always not proven. Operators can open the full ledger from the card.
  • The live mandate run. The latest live run puts the model through every on one recorded market day and counts sessions that stayed inside the posture's limits and any breach; the card's title is built from those counts ("stayed inside its mandate in 40 of 40 sessions", or "breached its mandate …" when it did). Under it: how a change to a role's behaviour ships. It is trialled live, several runs with the change against the same number without it, and only a proven change ships; the inconclusive and dead trials are counted too.

None of these four is a claim about money. They say the instrument is built correctly and reports honestly.

What only time proves

Four figures, each in its own card with a state pill; the fourth (up and down capture) prints once SPY has had at least ten up days and ten down days in the record:

  • Return vs SPY at the book's volatility. The book's return since its first funded day next to SPY's return over the same days, with SPY scaled so its day-to-day swings match the book's, and the difference between the two with its 95% interval (whole weeks of daily differences resampled, the same method the weekly Desk minus Rules figure uses). A direction is read only from that interval and only after 63 trading days: the "Reads:" sentence says "not yet distinguishable from SPY at the same risk" while the interval includes zero, "ahead" or "behind" only when the whole interval sits on one side, and "too short to judge" before 63 days whatever the tiles show. The note under it says what SPY itself returned, the scaling used and how many weeks the interval resampled.
  • Worst fall from a high vs SPY. The deepest peak-to-trough fall of the book against SPY's over the same days, and the gap between them. Two single paths carry no interval, so this card never draws a conclusion from the gap at any age.
  • Desk minus Rules, per week. If a rules twin runs beside this book (a rules-only paper copy trading the same signals mechanically, with no AI PM), the weekly difference between the two is the one honest measure of what the AI's judgement adds. It prints as a percent a week with its 95% interval and one of three readings: not yet distinguishable from zero, the desk is ahead, the rules are ahead. Whether the AI keeps the Portfolio Manager job is decided at 13 counted weeks; the card says so in those words.
  • Up and down capture vs SPY: on the days SPY rose, how much of its gain the book kept on average; on the days it fell, how much of its loss the book took. Two averages with no interval, so no verdict; 100% on both would mean the book simply moved with the market. (Computed on daily returns; the monthly convention needs a year of months.)

The states

StateWhat the card showsWhy
collecting"Collecting: 18 of 30 trading days" with a progress line, no numberBelow the floor a comparison would be noise. SPY comparisons print after 30 trading days; Desk minus Rules after 8 counted weeks
shown, not judgedThe numbers and the interval, with an amber pill, and "Reads: too short to judge"Between 30 and 63 trading days (about three months) the figure is real but too young to read anything into, so no direction word prints
measured, earlyThe numbers, neutral pill, "under twelve months: early"A year's record is 252 trading days; the win condition is stated over a rolling year
measuredThe numbersEnough history to quote
withheld"Withheld: week 2026-W40 is not counted: the rules twin ran no session that week" (or "had no signal table to act on that week (a platform fault on the twin, not a result)", or "recorded no decision on signals that qualified") with "2 of 8 weeks counted; this week withheld"A week the rules twin could not or did not trade would compare the desk with cash, not with the rules, so it is computed, shown as withheld, and never enters the average. Below the 8-week floor the card reads withheld and keeps its progress; at or above the floor only the week is withheld: the figure stays measured on its counted weeks and the pill says "measured; this week withheld". The SPY comparisons are withheld when SPY closing prices are not loaded for enough of the book's days (the card says how many, and states the book's own return as a fact, not a comparison), and all three are withheld when the book's value has not moved at all in the window (all cash or never traded): cash is not downside protection and 0% of SPY's loss is not a result
not set up"No rules twin runs beside this book yet"The comparison does not exist for this book; it starts the week a twin is created

Every card ends with its sample line ("71 trading days, Sep 15 to Dec 23, 2026" or "9 weeks counted") and "as of" the last day in the sample.

What this does not prove

The last section is four plain sentences that carry no number and no register fact, so they can never contradict the cards above: a backtest is not a live result; passing evals prove correct behaviour, not profitable judgement; the scoreboard is the forward record, reads collecting or early until it has its sample, and reads a direction only from an interval; paper fills are simulated and live-broker slippage is measured separately before real money.

NoteNothing on the page is a forecast or a target. If a sentence on it ever reads like a promise, that is a defect; tell us through the feedback button.

Where the numbers come from

  • The live figures are computed from this book's own end-of-day results and SPY's daily closes over the same days, the same definitions Performance uses. The matched-volatility line uses one volatility estimate for the whole platform (an exponentially weighted one over at least 20 sessions).
  • Desk minus Rules is the latest weekly attribution report written into this book; the Live track record and Backtest evidence pages are the two older halves this page sits on top of.
  • The machinery cards summarise committed records (the register, the eval suites, the live mandate run and its trials) as of the date printed on each card, and the readiness counts come from the latest readiness run loaded for this deployment (Admin › Readiness for operators). The register summary is regenerated in the same command that regenerates the register, and the platform's own change gate refuses a summary that is older than its sources.