Skip to content

Model matrix, safer deploy plumbing, and honest signal evidence

Model choice is now backed by measurement.

What changed for people using Bellwether

  • Model choice is now backed by measurement. Each desk role has a recorded verdict from live evaluations: for Consult you can pick among the supported models; for the roles that run your book unattended the most capable model stays fixed until the numbers say a cheaper one is just as reliable. The evidence behind each verdict is kept and re-run as models change.
  • The signal-blend evidence page is stricter and clearer. The comparison between today's blend and the candidate blend now uses the same yardstick for both, requires a minimum history before it can recommend anything, checks hit rate bucket by bucket, and says plainly when a universe is a stand-in or not point-in-time. Today's verdict is unchanged: the candidate blend stays in the background and does not drive trades.
  • Behind the scenes, the checks that run after every update got stricter (a missing credential now fails loudly instead of passing quietly), which is groundwork for hands-off releases.

Release 2f3e3af