This is the fine-print version of /docs/usage-limits. Same allowances, more precision about how they are measured. If the two pages ever disagree, this one is wrong and we want to know.
Per-plan compute allowance
There is no message-count cap. The meter is the dollar cost of inference over your billing cycle, and the allowance is what is left of your subscription after the meal-fund earmark, payment processing, and an operations reserve.
How the cap is measured
- The cycle allowance runs with your billing period, not a calendar month and not a rolling day.
- Every turn is priced at the real per-token rate for the model that handled it, so a Fast (Haiku) turn costs roughly a fifth of the same turn on Deep (Sonnet).
- Cached input is billed at a tenth of the normal input rate, which is why a long thread costs less per turn than you would guess.
- The Deep-model short window is a separate rolling 5-hour sum, and only Sonnet turns count toward it.
- We deliberately do not publish a message-count equivalent, because it would vary by a factor of ten depending on attachments and thread length. Settings → Usage shows the real figure.
What counts toward the cap
- Your prompt text — input tokens.
- Attached file contents — input tokens (this is the big one; a 100K-token attached file eats 100K of your daily budget every time you send a message with it attached).
- System instructions from custom instructions or project instructions — input tokens.
- The model's response — output tokens.
- Chat history sent back to the model with each new message — input tokens (grows over long chats).
What doesn't count
- Failed messages that error out server-side (we don't charge for our own failures).
- Messages you sent but the model didn't complete (network dropped, etc.).
- Reading past chats or scrolling — no model calls involved.
- File uploads themselves (upload doesn't call the model; only sending a message that includes the file does).
What happens when you hit the cap
Three stages. At 80% a quiet chip appears near the composer and nothing else changes. At 90% the Deep model pauses and Fast keeps working — “Deep model paused. Fast is still yours.” At 100% new sends stop until your cycle resets, and the message names the reset date. You can still open past chats, read them, search them, and export everything. There is no overage billing; the caps are hard and we do not queue or replay blocked messages.
If you consistently hit caps
- First: check whether large attached files are the driver — Projects with heavy attached files eat budget fast.
- Second: consider upgrading to the next tier.
- Third: set Fast as your default for routine turns and save Deep for the work that needs it — that alone stretches the allowance a long way.
- Fourth: if Max 5x isn't enough for one person, that's genuinely unusual — email us with what you're doing and we'll work out whether Max 20x fits or whether the workflow needs restructuring. Our allowances are modelled rather than observed, so this feedback is genuinely useful to us.
Fair-use enforcement
Caps are self-enforcing (you hit them, you stop). We do not throttle preemptively or slow down based on load. Automation that generates messages programmatically (bots, scripts) against the web UI violates the AUP (see /legal/aup) even if under the cap — the caps assume a human at the keyboard.