I'm getting rate limited.
UPDATED 2026-08-05Two different rate limits can hit you.
**1. LADLE's per-cycle compute allowance.** Your plan includes a dollar allowance of inference per billing cycle (about $9.12 on Base — the derivation is at /docs/usage-limits). It steps down in three stages:
- **80%** — a quiet chip near the composer. Nothing else changes. - **90%** — the Deep model pauses. You see "Deep model paused. Fast is still yours." and Fast (Haiku) keeps answering normally. - **100%** — sends are blocked until your cycle resets, and the message tells you the reset date. Reading, search, your library and exports all keep working; it is only new messages that stop.
Being straight with you: at 100% you are blocked. Nothing is queued in the background and replayed later, and we do not slow responses down as a hint. If you need to keep going before the reset, upgrading takes effect immediately.
**1b. The short-window cap.** Separately there is a rolling 5-hour cap on Deep (Sonnet) turns only — roughly $0.57 on Base. It stops one intense afternoon from eating the whole month. If you hit it you will see "You have spent a lot in the last few hours," and it clears as the window rolls forward. Fast turns never count toward it.
**1c. The flood guard.** More than ten messages in a minute gets a brief pause with a countdown. That one is about abuse rather than your allowance, and it clears in seconds.
**2. Anthropic's model-level 429s.** Occasionally Anthropic returns a "High demand right now" error during a peak-load spike across the entire platform. This is unrelated to your account's cap — it affects everyone. LADLE shows the copy "High demand right now — try again in a moment." Wait 15-30 seconds and retry; these usually clear fast.
**How to tell which one you're hitting:** - LADLE's cap: banner reads "RATE-LIMITED · RESUMES IN N MIN" with a specific countdown. - Anthropic's 429: the reply just says "High demand right now — try again in a moment" without a countdown.
**Reducing your cap usage:** - Attach files to Projects rather than individual chats. A project file is stored once and re-processed on every chat in the project; a chat-attached file is re-processed on every message in that chat. - Interrupt bad responses early with Esc — the tokens generated so far still count. - Use ⌘R to regenerate rather than paraphrasing your prompt — regeneration doesn't grow the input. - Turn off SEARCH for chats that don't need it. Search retrieval adds input tokens. - Turn off THINK for casual conversations. Thinking tokens are billed as output.
**Upgrading:** Base is generous for most professional workloads. If you regularly hit the cap several times a week, Max 5x is likely the right move. Max 20x is for extremely heavy use (several hours a day with large file attachments).