A Claude AI rate exceeded error means Claude has stopped accepting requests from you for a while, and the fix depends entirely on which limit tripped. In the chat app it is usually your plan's usage budget or a temporary capacity squeeze on Anthropic's side. On the API it is an HTTP 429 or 529.
Read the exact wording before you do anything. The wording tells you whether to wait a few minutes or change something on your side.
What the Claude AI rate exceeded error is telling you
"Rate exceeded" is the phrase people search for. Claude itself uses several different messages, and each one points to a different cause.
Anthropic's guide to Claude error messages lists them. Here is how they map to a response:
| What you see | What tripped | What to do |
|---|---|---|
| Approaching 5-hour limit | Your plan's usage budget, nearly spent | Finish the important task first |
| 5-hour limit reached, with a reset time | Your plan's usage budget, fully spent | Wait for the reset, or use usage credits |
| Due to unexpected capacity constraints, Claude is unable to respond | Demand across all users | Retry in a few minutes |
| Your message will exceed the length limit for this chat | The conversation's context window | Start a new chat or attach less |
HTTP 429 rate_limit_error | Your API organization's limits | Honour retry-after and back off |
HTTP 529 overloaded_error | API load across all users | Retry with exponential backoff |
The length limit sneaks onto that list because it feels the same from the keyboard. It is a different thing entirely. The context window, meaning the maximum amount of text Claude can hold in one conversation, has nothing to do with time, so waiting will not fix it. A fresh chat will.
The rest of this article works through the other rows, since those are the ones that keep people staring at a spinner at 11 p.m. and refreshing like it might help.
Claude AI usage limits and the five-hour window
Usage limits are a budget. Every plan gets a certain amount of Claude over a window of time, and when the budget runs out you wait until it refills.
According to Anthropic's help centre, one budget covers the chat app and Claude Code together. What drains it is conversation length and the model you choose, among others. On paid plans the window that bites is five hours, and the limit message tells you exactly when it resets.
The detail that surprises people is how conversations are billed against that budget. Each reply resends the entire conversation to the model. Turn forty of a long chat therefore costs roughly forty times what turn one did, even when your new message is a single line.
So the person who "barely used it today" often did one enormous thing. A single sprawling thread with three PDFs attached can outspend a dozen short, focused chats, which is deeply unfair and entirely consistent with the maths.
When you hit the limit, you can wait for the reset or keep working on usage credits. Credits bill at standard API rates, so set a monthly spending ceiling before you enable them. Our breakdown of what each Claude plan costs and how much usage it buys has the numbers for Pro and both Max tiers.
Capacity constraints: when the limit belongs to Anthropic
Sometimes the limit has nothing to do with you. When demand spikes, Claude declines some requests even from accounts with plenty of budget left.
The message reads "Due to unexpected capacity constraints, Claude is unable to respond to your message. Please try again soon." Anthropic describes this as managing high demand, separate from a technical outage. So the status page may show no incident while you are staring at the error.
The fix is waiting a few minutes and sending the message again. Upgrading your plan will not help here, because capacity constraints apply across the board. Buying Max to escape a demand spike is like buying a faster car to escape a traffic jam you are already sitting in.
If the problem lasts well beyond a few minutes, or you cannot log in at all, check Anthropic's status page at status.claude.com. A genuine incident shows up there with progress updates.
Rate limit errors on the Claude API: 429 and 529
On the API, limits become status codes. A 429 means your organization crossed one of its own limits. A 529 means the API is overloaded for everyone.
Anthropic's rate limits documentation measures limits per model in requests per minute and tokens per minute, with input and output tokens counted separately. The API uses a token bucket algorithm, meaning capacity refills continuously instead of resetting on the minute. A limit of 60 requests per minute can therefore be enforced as one per second, so a short burst trips a 429 even when your minute total looks fine.
A 429 carries a retry-after header with the number of seconds to wait. Retrying earlier fails. There is one important exception: a 429 with no retry-after header means your organization has reached its tier's monthly spend cap, and every retry will fail until access resumes on the first of the next month or you move up a tier.
Anthropic's API errors reference also notes that a sharp jump in traffic can trigger 429s from acceleration limits, so ramp new workloads up gradually. The official SDKs already retry twice with exponential backoff. When you want control over that, a minimal version in Python looks like this:
import random, time
import anthropic
client = anthropic.Anthropic(max_retries=0)
def ask(messages, attempts=5):
for attempt in range(attempts):
try:
return client.messages.create(
model="claude-sonnet-5", max_tokens=1024, messages=messages)
except anthropic.RateLimitError as e:
wait = e.response.headers.get("retry-after")
if wait is None: # spend cap reached: retrying cannot help
raise
time.sleep(float(wait) + random.uniform(0, 1))
except anthropic.APIStatusError as e:
if e.status_code < 500:
raise
time.sleep(min(2 ** attempt, 30) + random.uniform(0, 1))
raise RuntimeError("Claude API still unavailable after retries")
The random jitter matters more than it looks. Without it, every client that failed at the same moment retries at the same moment, and together they recreate the spike that caused the error.
Prompt caching is the cheapest way to raise your effective ceiling. For most current models, tokens read from the cache do not count toward the input-tokens-per-minute limit, so a large shared system prompt stops eating your rate budget on every call.
Why Claude Code and AI agents hit limits first
Agents make many requests where a person makes one. That is why Claude Code users meet usage limits long before chat users do.
A single instruction to an agent becomes a loop of reading files and calling tools, repeated until the task is done. Every step carries the growing context forward. An afternoon of Claude Code can spend the same five-hour budget that a week of casual chat would, because the budget is shared between the two.
On the API the same loop shows up as bursts. Ten parallel sub-agents firing at once is exactly the short spike a token bucket punishes. Cap how many run at once, and give every loop a stopping condition. An agent with no stopping condition will retry a failing tool until a rate limit stops it for you, which is a stopping condition of sorts, though an expensive one.
If you are building these loops yourself, Agentic AI Engineering covers reasoning loops and tool design at the level where request volume actually gets decided.
How to stop seeing the error
Most rate exceeded errors come down to habits more than plan size. A short checklist covers the common cases:
- Start a new conversation when a thread gets long, and paste in a short summary of what matters.
- Lower the effort level and turn off tools you are not using, such as web search or Research, for routine questions.
- Check Settings, then Usage, before you start a long session so the reset time does not ambush you mid-task.
- Treat a capacity message as a five-minute coffee break and nothing more.
- On the API, honour
retry-afterwith jitter added. - Alert on any 429 that arrives with no
retry-after, since that one means the spend cap. - Cache stable prompt content and ramp new traffic gradually.
Upgrade only when the five-hour reset keeps interrupting paid work after you have fixed the habits above. If you are still getting to grips with how Claude's models and Claude Code fit together, Claude AI for Beginners walks through them from the ground up.