Why Haiku isn't the default for everything.
APRIL 2026 · MODELHaiku is cheaper and faster. Sonnet is better at the hard turns. LADLE routes between them per turn, publishes the rule, and lets you override it — because the alternative is implying one model and serving another.
Anthropic ships two tiers of Claude that LADLE uses: Sonnet (larger, slower, more expensive) and Haiku (smaller, faster, cheaper). Plenty of products in this category pick one, advertise it, and route to the cheap one when nobody's looking. We'd rather publish the rule.
Here it is. A turn goes to **Sonnet** if any of these is true:
- the prompt is longer than 800 characters - it's the sixth message or later in the thread - it contains a fenced code block - it reads like a code or analysis request (write, fix, refactor, debug, analyze, explain this error) - any earlier turn in that chat already used Sonnet — once a conversation goes heavy it stays heavy - you picked **Deep** in the composer
Everything else goes to **Haiku**. In practice that means short casual turns at the start of a thread.
## Why not Haiku for everything
By raw economics, routing every reply to Haiku would cut inference cost substantially. We don't, because the turns people subscribe for are the substantive ones.
Haiku is genuinely excellent at what it's built for: short, well-scoped tasks with clear instructions. Draft a subject line. Answer this factual question from the context I gave you. On tasks like that it's fast enough to feel instant and reliable enough to feel like the right tool.
Haiku is not as good at the tasks that make an assistant feel like an assistant. Long-form drafting where you want the model to read tone and mirror it. Multi-step reasoning where the second step depends on getting the first step right. Code review where the interesting problem is what the code *should* do. Rewrite-in-my-voice work that needs a lot of context held at once.
When you type "help me think through this," you expect the model that thinks better. Routing that to the small model to save a cent is how you lose the subscriber who was going to fund meals for a year.
## Why not Sonnet for everything
That's the goal. Stating it as a goal rather than implying we're already there: at $20/month with $8 committed to the meal fund and inference paid out of what's left, Sonnet on every turn doesn't currently clear. If Anthropic's prices fall or our subscriber volume grows enough to absorb it, routing to Haiku is the first thing we retire — and it'll be announced on /changelog, not quietly flipped.
## Where the rule gets it wrong
It's a heuristic, not a difficulty classifier, and it has a known weak spot: a genuinely hard question asked in a short sentence looks casual to the router. "Is this proof valid?" is 22 characters and goes to Haiku.
The mitigation is the Deep toggle — one tap in the composer, applies to the turn you're about to send, available on every plan. If you're doing serious work, set Deep and stop thinking about it. And every response reports which model produced it, so you're never guessing.
## The other places Haiku shows up
**Fast mode.** The composer toggle that forces Haiku. Use it when you want speed over depth.
**Graceful degradation.** When a subscriber approaches their monthly compute budget, the router falls back to Haiku rather than blocking sends outright. Deep model paused; Fast still yours. That's the honest version of graceful degradation — you keep getting help, just from the smaller model.
**Chat titling.** Thirty tokens out, deterministic pattern. Sonnet would be waste.
We haven't published a figure for what share of turns lands on each model, because LADLE hasn't launched and we don't have one. When there's a real number it'll go on /open with the rest of the metrics, and it'll be measured rather than estimated.
If Haiku closes the quality gap on assistant tasks — each release has been notably better — we'll move the whole rule toward it and say so. Until then: two models, one published rule, and a toggle that overrides it.