Error database

RateLimitError 429 from an LLM API

You sent requests or tokens faster than your account is allowed. Check whether the code says rate_limit_exceeded or insufficient_quota — retrying helps one of them and never helps the other.

The message you saw
RateLimitError 429 from an LLM API

By Updated

The error

Output
openai.RateLimitError: Error code: 429 - {'error': {'message': 'Rate limit reached for gpt-4o-mini in organization org-abc123 on tokens per min (TPM): Limit 200000, Used 199342, Requested 1200. Please try again in 162ms.', 'type': 'tokens', 'param': None, 'code': 'rate_limit_exceeded'}}
Output
anthropic.RateLimitError: Error code: 429 - {'type': 'error', 'error': {'type': 'rate_limit_error', 'message': 'Number of request tokens has exceeded your per-minute rate limit.'}}

A different problem that arrives with the same 429 status code:

Output
openai.RateLimitError: Error code: 429 - {'error': {'message': 'You exceeded your current quota, please check your plan and billing details.', 'type': 'insufficient_quota', 'param': None, 'code': 'insufficient_quota'}}

What it means

The provider refused the request because your account went over an allowance. Which allowance, and whether waiting will help, depends on the code field — and this is the first thing to read.

  • rate_limit_exceeded — you went too fast. The allowance refills continuously. Waiting works, and often the wait is measured in milliseconds.
  • insufficient_quota — you have no credit or have hit a spending cap. The allowance does not refill on its own. Retrying will never succeed, and a retry loop here burns your logs and your patience for nothing.

Providers reuse status 429 for both, which is why an error that looks like a traffic problem is sometimes a billing problem.

Why it happens

Limits are enforced on two axes at once, and you can hit either.

Requests per minute (RPM) counts calls. Twenty parallel workers each making three calls a second will cross a 60 RPM limit instantly, even though each individual call is tiny.

Tokens per minute (TPM) counts the text. Every request consumes your prompt tokens plus the output room you reserved with max_tokens. A RAG application sending 15,000 tokens of retrieved context per question uses TPM thirteen times faster than a chat app sending 1,100, at exactly the same request rate.

Limits are usually per organisation and per model, and new accounts start on a low tier that rises with usage history and spend. That is why the same code that ran fine in testing fails when you switch to a colleague's key, or when a batch job that used to process 50 rows starts processing 5,000.

Two aggravating factors are worth naming. Retries without backoff make things worse — every retry counts against the same limit that is already full. And burst patterns matter more than averages: 60 requests spread across a minute passes a 60 RPM limit, while the same 60 fired at once may not.

How to fix it

1. Read the code field before doing anything else.

python
from openai import OpenAI, RateLimitError

client = OpenAI()
try:
    resp = client.chat.completions.create(
        model="gpt-4o-mini",
        messages=[{"role": "user", "content": "hello"}],
    )
except RateLimitError as exc:
    print(exc.code)          # 'rate_limit_exceeded' or 'insufficient_quota'
    print(exc.message)

If it says insufficient_quota, stop here. Add credit or raise your spending limit in the provider's billing dashboard. No code change fixes this, and free trial credits do expire.

2. Retry with exponential backoff and jitter. For genuine rate limits, this alone resolves most production cases. Both major SDKs already retry a couple of times automatically, and you can raise that:

python
client = OpenAI(max_retries=5)          # the SDK backs off between attempts for you

When you need your own loop — around a whole pipeline step, say — the shape that works is:

python
import random
import time
from openai import OpenAI, RateLimitError

client = OpenAI(max_retries=0)          # take over retrying ourselves

def ask(messages, attempts=6):
    for attempt in range(attempts):
        try:
            return client.chat.completions.create(model="gpt-4o-mini", messages=messages)
        except RateLimitError as exc:
            if getattr(exc, "code", "") == "insufficient_quota":
                raise                                    # retrying cannot help
            if attempt == attempts - 1:
                raise
            delay = min(60, 2 ** attempt) + random.uniform(0, 1)   # jitter avoids lockstep
            print(f"rate limited, waiting {delay:.1f}s")
            time.sleep(delay)

The jitter matters. Without it, twenty parallel workers that all failed at the same moment retry at the same moment and fail together again.

3. Honour retry-after when the provider sends it. A stated wait is better than a guessed one.

python
resp = exc.response                       # httpx response on the SDK exception
wait = float(resp.headers.get("retry-after", 0)) if resp is not None else 0

4. Watch the headroom instead of waiting for failures. Every response carries the current state of your limits:

python
resp = client.chat.completions.with_raw_response.create(
    model="gpt-4o-mini",
    messages=[{"role": "user", "content": "hello"}],
)
h = resp.headers
print(h.get("x-ratelimit-remaining-requests"), h.get("x-ratelimit-remaining-tokens"))
print(h.get("x-ratelimit-reset-tokens"))

Slowing down when the remaining count gets low is far better than sprinting into a wall.

5. Cap your concurrency. Unbounded parallelism is the most common cause in batch jobs.

python
import asyncio
from openai import AsyncOpenAI

client = AsyncOpenAI(max_retries=5)
gate = asyncio.Semaphore(5)                # at most 5 in flight

async def ask(prompt: str):
    async with gate:
        resp = await client.chat.completions.create(
            model="gpt-4o-mini",
            messages=[{"role": "user", "content": prompt}],
        )
        return resp.choices[0].message.content

async def main(prompts):
    return await asyncio.gather(*(ask(p) for p in prompts))

Five concurrent requests that all succeed finish a job faster than fifty that spend their time being rejected and retried.

6. Use fewer tokens if TPM is the limit you are hitting. The message says which axis failed. For TPM, trim retrieved context, drop old conversation turns, and set max_tokens to something realistic. The reserved output space counts toward the budget. Asking for 4,000 tokens when you expect 200 wastes a fifth of a 20,000-token allowance on every call.

7. Spread the load. Route simple work to a smaller model with a separate limit. For anything not needed within seconds, a provider's batch endpoint has far higher limits and costs less. Caching identical requests removes them entirely — the cheapest call is the one you do not make.

How to prevent it

Put every API call behind one function that handles retries, backoff and quota detection, and let the whole codebase use it. Rate limiting handled in fifteen places is rate limiting handled in none.

Know your actual limits rather than guessing: they are listed per model in the provider's dashboard, and they change as your account tier changes. Size your concurrency from the RPM number and your batch size from the TPM number.

Log two things — how many requests were retried, and the remaining-token header. A steady rise in retries is the early warning that your traffic has grown into the limit, and it arrives days before the outage.

For anything user-facing, add your own queue in front of the provider. A short wait with a spinner is a better experience than an error page, and it keeps a burst of users from turning into a burst of 429s.