LADLE meters usage by what your turns cost us in inference, measured in dollars over your billing cycle — not by a token count over a rolling day. Each plan carries a compute allowance derived from its price, and there is a separate short-window cap that stops a single afternoon on the Deep model from draining the month. Approaching the allowance pauses the Deep model; exhausting it pauses sends until your cycle resets. The specifics are below, including the part where you do get blocked.
Per-plan compute allowance (per billing cycle)
Your allowance is the part of your subscription left after the meal-fund earmark, payment processing, and an operations reserve. We publish the derivation because a cap you cannot check is just a number. Base is generous for most workloads; Max tiers exist for heavy days.
What happens at the cap
Three steps, and we are not going to soften the third one: at 100% you are blocked from sending until your cycle resets. We do not queue the request, we do not slow it down as a hint, and nothing is silently held and replayed later. The reason to be blunt is that a product which promises you will never be blocked and then blocks you has told you something false at the worst possible moment.
What counts toward the cap
Everything the model processes counts toward the dollar allowance, priced at the real per-token rate for whichever model handled the turn — so a Haiku turn costs roughly a fifth of the same turn on Sonnet. That includes re-processing of persistent Project files on every message. Cached input is billed at a tenth of the normal rate, which is why long threads cost less per turn than you would expect.
How to see current consumption
Settings → Usage shows your spend against this cycle's allowance and against the 5-hour Deep window. Usage is recorded just after each response finishes streaming, so a turn you just sent may take a few seconds to appear.
How to use less
The single biggest lever: attach files to Projects rather than chats. A project file attached to a 20-message chat costs ~1× the file's tokens; attached to 20 separate chats it costs ~20×. Second lever: interrupt bad responses early (Esc during streaming) — this stops the model mid-generation and reduces output tokens. Third: use ⌘R to regenerate rather than paraphrase your prompt, so the input tokens don't grow.
Which plan fits your usage
Base fits most individual professional workloads. Max 5x is appropriate if you regularly attach large PDFs to multiple long-running chats, or if you keep Deep switched on all day. Max 20x is for people running the model constantly. Try Base first and upgrade only if you actually hit the wall more than once or twice a month — and if you hit it in week one, tell us, because our allowances are set from modelling rather than from real usage data and we would rather hear it than have you cancel.
Docs: File types and limits · Projects
Task guides in /help: What happens when you hit the cap. · Base vs Max 5x vs Max 20x.