Error 529: overloaded_error — Overloaded (Anthropic)
Anthropic's servers are temporarily saturated — nothing in your code is wrong. Retry with exponential backoff and jitter, and let the SDK's built-in retries do their job.
Updated
The error
anthropic.InternalServerError: Error code: 529 - {'type': 'error', 'error': {'type': 'overloaded_error', 'message': 'Overloaded'}}What it means
Status 529 is Anthropic's "servers are busy" signal. Demand briefly exceeded capacity, and your request was shed to protect the service. It is transient, it is not caused by your code, your account or your prompt, and the same request will succeed when pressure drops — often seconds later.
The one thing 529 asks of you is to retry politely. Hammering a saturated service extends the saturation.
Why it happens
Traffic to LLM APIs is spiky — big product launches, viral moments, regional peaks. During spikes, some requests get 529 instead of queueing forever. Long or complex requests are neither more nor less likely to be shed; it is a load signal, not a judgement of the request.
How to fix it
1. Rely on the SDK's automatic retries first. The official client already retries 529 (and similar transient statuses) with backoff. You can raise the attempt count:
import anthropic
client = anthropic.Anthropic(max_retries=5) # default is 2For many scripts this is the entire fix.
2. For batch jobs, add your own backoff with jitter around the call.
import random, time
import anthropic
client = anthropic.Anthropic()
def call_with_backoff(**kwargs):
delay = 2.0
for attempt in range(6):
try:
return client.messages.create(**kwargs)
except (anthropic.InternalServerError, anthropic.APIStatusError) as e:
status = getattr(e, "status_code", None)
if status not in (500, 502, 503, 529) or attempt == 5:
raise
time.sleep(delay + random.uniform(0, delay)) # jitter
delay *= 2Jitter — the random extra sleep — prevents every client from retrying in the same instant, which would re-create the spike.
3. Reduce concurrency during incidents. If a worker pool is sending 50 parallel requests, halve it when 529s start and restore gradually. Failing fast on the whole batch wastes the requests that would have succeeded.
4. For non-urgent volume, move work off-peak or into batch. Overnight windows see fewer spikes, and batch-style processing tolerates deferred execution by design.
5. Check the status page during long incidents. A sustained wall of 529s across many minutes is a platform event — status.anthropic.com will say so, and waiting beats engineering around it.
How to prevent it
Build retry-with-backoff into every production LLM call path from day one — 529, 500 and network blips all reuse it. Cap total retry time and surface a clean failure to users when the cap hits. Distinguish 529 from 429 in logs: one is their capacity, the other is your quota, and the dashboards you check differ.
Related errors
- RateLimitError 429 — your account's limits, not their servers
- APIConnectionError — when the request never arrives at all
- Your credit balance is too low (Anthropic)