How the desk learns
The Learning tab: what each role is graded on and why most tiles say "not enough data yet" at first, repeated mistakes, the lessons timeline, and hiding a lesson from the desk.
- Who it is for
- Owners & operators
- Reading time
- 6 min read
- Updated
The desk is meant to get better at its own jobs from its own history, a little every day, without anyone editing prompts by hand. Your desk › Learning is where you watch that happen. Everything on it is counted by code after the close, never by the model grading itself, and none of it changes trading on its own.
- Owns
- Nothing an agent reads: the grades, repeated-mistake counts and the timeline are for you
- Reads
- Sessions, , fills, daily prices, the 's findings, the automatic session review, the lessons the agents wrote
- Runs
- After the close, in the same slot as the : settle outcomes, grade each role, refresh , queue memory clean-ups
- Change it at
- One lever: hide a lesson from the desk on the Lessons timeline (lesson visibility· Your desk › Learning › Lessons timeline › Hide from the desk)
The loop, station by station
The strip at the top reads left to right:
- Record. Every session is written down with the exact instructions, memory, model and rubric it ran under, so a grade can say which version it describes.
- Settle. After the close, code works out what happened to each decision (proposals, including the ones the rejected or that expired; position ; the 's checkable calls) from daily prices and the book's own fills. A decision stays "open" until its horizon has passed, then it is settled once and never re-scored.
- Grade. Each role gets a small set of grades with honest sample sizes (below).
- Tidy memory. A nightly pass looks for duplicate, stale or contradictory lessons and queues up to three clean-ups for the PM to answer in its own end-of-day session. Code never rewrites a lesson; the agent does, in its own words.
- Draft fixes. Repeated mistakes become suggested instruction changes you approve with your code. This station arrives in a later release and reads "later" until then.
A station is red when it was expected after the last close and left nothing; the day report says so too (Review).
Why most tiles say "not enough data yet"
A grade over seven decisions is noise dressed up as a number. Every tile is held to a floor (about 20 sessions for rates, 30 settled decisions for confidence accuracy, 20 flagged proposals for the Critic), and below it the tile prints exactly how many it has and how many it needs, for example (needs 30 decisions, has 7). No value, no colour, no arrow. At a normal pace the Portfolio Manager's row fills first, within a few weeks; the Researcher, Critic and rows fold to one line until something is measured.
Above the floor a tile shows the value, a range (a 95% interval that accounts for decisions made on the same day moving together), and what it was counted over. Tiles marked context (ideas that worked, brief calls that came true) describe results, not skill, and never trigger anything.
Under each role's tiles, its one headline measure is compared this month against last month. When an exact statistical test says the two months differ by more than chance, the line reads . That is deliberately neutral: it may be noise or a real shift, and the right response is to read the sessions, not to change a setting.
What each role is graded on
- Portfolio Manager: confidence accuracy (how close its stated confidence was to what happened, beside the score of just guessing the base rate), whether it honoured its own stops, whether buy ideas named a stop, slip-ups per session from the automatic review, sessions with a slip-up (the headline), repeated known mistakes, cost per decision, sessions that acted.
- Researcher: briefs with an uncited claim (the headline), sources per brief, slip-ups per brief, and brief calls that came true (context).
- Critic: objections that proved right: how much more often a trade went bad when the Critic objected than when it did not (the headline), time to answer a finding, and how often the PM overruled it and the trade then went bad.
- Consult: questions left unanswered (the headline) and suggestions you accepted.
The agents never see any of this. A model that can read its own scoreboard learns to please the scoreboard; the PM's carries at most one line about the Critic's track record, and only once there are twenty flagged proposals behind it.
Repeated mistakes
When the automatic session review or the Critic catches the same slip-up more than once (a proposal without data references, a moved goalpost, a brief claim without a source), it is counted here per role, with the instruction it points at and a 30-day sparkline. A row says elevated only when the last two weeks are higher than that desk's own longer-run rate by an exact test; otherwise it is a plain count ("seen 3 times in 14 days · usually about 1"). These counts come from code, never from the lessons, so retiring a lesson cannot make a mistake look fixed.
The lessons timeline, and your one lever
Lessons are notes an agent writes to itself after the close (memory). The timeline shows every change with the before and after: learned, noticed again, pinned, merged (two lessons folded into one, in the agent's own restated words), reworded, kept after review, retired.
You do not write, edit, merge or restore lessons: they are the agent's own experience, and an operator-authored "lesson" would be a in disguise (write guidance instead). Your one lever is : the lesson stops appearing in the state card, the row is untouched, and the agent sees hidden by operator with your reason in its next end-of-day session, where it must either retire the lesson or say why it should stay; it may not quietly write the same rule again. Hiding needs a reason and no code, because it can only reduce a lesson's influence. If the agent later folds a hidden lesson into another one, your hide follows the merged lesson until you lift it. Show to the desk again lifts it. Both are in the audit log.
What each role runs under
The versions band shows, per role, the version of its context it currently runs under and why the last one opened: instruction modules changed, an override was enabled or rolled back, the Critic rubric changed, the model policy changed. It links to Portfolio Manager › Configure, the one place those are edited. Memory changing every evening does not open a new version; that is the point of memory.
Queued for the Portfolio Manager tonight
The last block lists what the nightly tidy-up has queued: two lessons that say the same thing (with the longer wording on file as a starting point), a lesson that quotes a limit the book no longer has, one nobody has seen in six weeks, one you hid. The Portfolio Manager answers each in its end-of-day reflection (fold two into one in its own words, reword, retire, or keep with a reason), and the answer shows up on the lessons timeline the next morning. A merge or rewording that adds anything the original lessons did not say is refused and the item stays queued; so is any "lesson" that claims an edge ("wins 70% of the time") or talks about a score. Each item expires after two weeks if unanswered, and an unanswered item a slip-up in the session review. These are the agent's to answer, so there are no buttons here and they never appear in the Inbox.