Being direct about what we actually ship, because the internet is full of products that imply one model and serve another: a short casual turn on LADLE goes to Claude Haiku 4.5. A longer prompt, a code block, a thread past its fifth message, an analysis request, or anything after you pick Deep goes to Claude Sonnet. The exact rule is on /models. This page is about the obvious next question — if Haiku is cheap and good, why not use it for all of it?
What we actually do today
Short casual turnHaiku 4.5Under 800 chars, first five messages, no code
Prompt over 800 charactersSonnet
Sixth message onward in a threadSonnet
Contains a code block or reads as code/analysisSonnet
Any thread that already used SonnetSonnetNever drops back mid-conversation
You pick DEEP in the composerSonnetAlways available, on every plan
The quality gap is real, and it shows up unevenly
Haiku is roughly a fifth of Sonnet's token cost and noticeably faster. It is also less capable on multi-step reasoning, long-context recall, nuanced writing register, and complex code. The thing is, that gap barely shows on "what's a good substitute for buttermilk" and shows immediately on "read this contract and tell me what's unusual." So the routing rule is an attempt to spend the expensive model where the gap actually matters — which is also why the rule leans heavily on length, code signals, and thread depth rather than trying to guess difficulty from the topic.
Where the rule gets it wrong
It will get it wrong sometimes, and we'd rather say so than pretend the heuristic is clever. A genuinely hard question asked in a short sentence — "is this proof valid?" — looks casual to the router and goes to Haiku. That's the known weak spot. The mitigation is the Deep toggle, which is one tap in the composer and applies to the turn you're about to send. If you're doing serious work, set Deep and stop thinking about it.
EVERY RESPONSE REPORTS THE MODEL THAT PRODUCED IT
What Haiku is genuinely good at
- Short conversational turns where the answer is short — which is a large share of real chat usage.
- High-throughput API workloads: classification, extraction, transformation.
- Real-time interactive agents where latency is the primary constraint.
- First-pass filtering before routing to a bigger model.
Why not Haiku for everything, then
Because the turns people subscribe for are the substantive ones. Quality degradation is the single biggest driver of AI-subscription churn, and routing a 2,000-word document analysis to the small model to save a cent is how you lose the subscriber who was going to fund meals for a year. Running Haiku across the board might let us raise the meal-fund earmark from $8 to $12 per subscription — but only if subscribers stay, and they wouldn't.
Why not Sonnet for everything
That's the goal, and we'd rather state it as a goal than imply we're already there. It's an economics problem, not a technical one: at $20/mo with $8 committed to the meal fund and inference paid out of the remainder, Sonnet on every turn doesn't currently clear. If Anthropic's prices fall or our subscriber volume grows enough to absorb it, routing to Haiku is the first thing we retire — and it'll be announced on /changelog, not quietly flipped.
GOAL · NOT A SHIPPED FEATURE
If you specifically want Haiku on everything
Set FAST in the composer and it sticks for the turns you send. You'll get faster responses and the same meal-fund earmark — the $8 doesn't move with inference cost, so a cheaper turn doesn't fund more meals, it just widens the margin that keeps the $8 safe.