Claude gets a new version. Should Brainback upgrade?
The answer is almost never immediately.
The reason: our prompts and guards are tuned to specific model behavior. A new model handles our system prompts differently, sometimes better, sometimes worse. Before we roll out, we run our full evaluation suite against the new model — a set of about 200 tutor turns, 50 hint requests, 50 explain-backs, and 50 known-answer leak tests — and only upgrade if the numbers hold across every category.
For the recent Sonnet upgrade, tutor quality went up, hint quality was flat, and leak rate went down. Ship. For the last Haiku upgrade, classifier accuracy dropped by 3 points. Hold and rewrite the prompt before rolling out.
This is boring work. It's also the work that keeps the product's stance intact.
Filed by Engineering on Oct 19, 2026. Under Engineering.