Error database

Error 529: overloaded_error — Overloaded (Anthropic)

Anthropic's servers are temporarily saturated — nothing in your code is wrong. Retry with exponential backoff and jitter, and let the SDK's built-in retries do their job.

The message you saw
Error 529: overloaded_error — Overloaded (Anthropic)

By Updated

The error

Output
anthropic.InternalServerError: Error code: 529 - {'type': 'error', 'error': {'type': 'overloaded_error', 'message': 'Overloaded'}}

What it means

Status 529 is Anthropic's "servers are busy" signal. Demand briefly exceeded capacity, and your request was shed to protect the service. It is transient, it is not caused by your code, your account or your prompt, and the same request will succeed when pressure drops — often seconds later.

The one thing 529 asks of you is to retry politely. Hammering a saturated service extends the saturation.

Why it happens

Traffic to LLM APIs is spiky — big product launches, viral moments, regional peaks. During spikes, some requests get 529 instead of queueing forever. Long or complex requests are neither more nor less likely to be shed; it is a load signal, not a judgement of the request.

How to fix it

1. Rely on the SDK's automatic retries first. The official client already retries 529 (and similar transient statuses) with backoff. You can raise the attempt count:

python
import anthropic

client = anthropic.Anthropic(max_retries=5)   # default is 2

For many scripts this is the entire fix.

2. For batch jobs, add your own backoff with jitter around the call.

python
import random, time
import anthropic

client = anthropic.Anthropic()

def call_with_backoff(**kwargs):
    delay = 2.0
    for attempt in range(6):
        try:
            return client.messages.create(**kwargs)
        except (anthropic.InternalServerError, anthropic.APIStatusError) as e:
            status = getattr(e, "status_code", None)
            if status not in (500, 502, 503, 529) or attempt == 5:
                raise
            time.sleep(delay + random.uniform(0, delay))   # jitter
            delay *= 2

Jitter — the random extra sleep — prevents every client from retrying in the same instant, which would re-create the spike.

3. Reduce concurrency during incidents. If a worker pool is sending 50 parallel requests, halve it when 529s start and restore gradually. Failing fast on the whole batch wastes the requests that would have succeeded.

4. For non-urgent volume, move work off-peak or into batch. Overnight windows see fewer spikes, and batch-style processing tolerates deferred execution by design.

5. Check the status page during long incidents. A sustained wall of 529s across many minutes is a platform event — status.anthropic.com will say so, and waiting beats engineering around it.

How to prevent it

Build retry-with-backoff into every production LLM call path from day one — 529, 500 and network blips all reuse it. Cap total retry time and surface a clean failure to users when the cap hits. Distinguish 529 from 429 in logs: one is their capacity, the other is your quota, and the dashboards you check differ.