HTTP 429 from the API: read rate-limit headers and back off correctly
For developers debugging 429 responses from an API under real traffic. This runbook shows how to identify which limit you hit, read the response headers correctly, verify whether retries are making it worse, and apply concrete fixes in clients, workers, and gateways.
TL;DR — A 429 means your client is sending requests faster than the API currently allows, and the fastest path to resolution is to inspect the response headers on the 429 itself:
Retry-After,X-RateLimit-*, and any request/correlation ID. The most common fix is to stop immediate retries, add bounded exponential backoff with jitter, and reduce concurrency so you stay under the effective per-key or per-IP limit. Reading time: ~6 min
The scenario
You push a Tuesday afternoon deploy that parallelizes a few API calls to cut page latency. Five minutes later, your workers start logging 429 Too Many Requests, the queue depth climbs, and your "quick retry" code turns one bad minute into a self-inflicted storm. Product asks why some users see stale data while others work fine, and your dashboards show request volume only slightly higher than normal. You need to know whether you hit a documented rate limit, a burst/concurrency limit, or a proxy in front of the API is collapsing many clients behind one IP.
Symptoms
- Client logs show HTTP 429 responses, often verbatim like:
HTTP/1.1 429 Too Many Requests
- Application errors such as:
fetch failed: 429 Too Many Requests
AxiosError: Request failed with status code 429
requests.exceptions.HTTPError: 429 Client Error: Too Many Requests
- Response headers on the failing request include one or more of:
Retry-After: 17
X-RateLimit-Limit: 100
X-RateLimit-Remaining: 0
X-RateLimit-Reset: 1733241120
RateLimit-Limit: 100
RateLimit-Remaining: 0
RateLimit-Reset: 17
- Repeated retries amplify the problem; logs show the same request ID or endpoint failing in bursts:
POST /v1/jobs 429 req_id=8f2b... attempt=1
POST /v1/jobs 429 req_id=91ac... attempt=2
POST /v1/jobs 429 req_id=cc10... attempt=3
- End users see partial failures, stale data, or delayed jobs rather than a total outage.
- If requests traverse a shared egress proxy/NAT, multiple services start failing at once even though each service's own QPS looks low.
Likely causes
| Cause | How common | Quick check |
|---|---|---|
Client retries ignore Retry-After or retry immediately | Very common | `grep -RniE 'retry |
| Too much parallelism/burst traffic from a deploy, worker pool, or cron fan-out | Very common | kubectl top pods -A |
| You are hitting a different limit than expected: per-IP, per-token, per-endpoint, or concurrency limit | Common | curl -sS -D - -o /dev/null https://api.example.com/v1/resource |
| Shared NAT/egress IP causes unrelated clients to share one rate bucket | Common | curl -s https://ifconfig.me |
| Caching is missing, so identical reads hammer the API | Common | `grep -RniE 'Cache-Control |
| A gateway/CDN/WAF in front of the API is generating the 429, not the origin API | Less common | `curl -sS -D - -o /dev/null https://api.example.com/v1/resource |
Step-by-step diagnosis
- Inspect one real 429 response, headers included.
curl -sS -D - -o /tmp/body.out https://api.example.com/v1/resource
If this is your problem, the headers will show 429 Too Many Requests plus Retry-After and/or X-RateLimit-* or RateLimit-*. If you only see a bare 429 with gateway headers like server: cloudflare, via:, or x-cache, jump to Fixes → Gateway/CDN/WAF is generating the 429.
- Distinguish window-based limits from burst/concurrency limits.
curl -sS -D - -o /dev/null https://api.example.com/v1/resource | egrep -i '^(HTTP/|Retry-After:|X-RateLimit-|RateLimit-)'
If Remaining: 0 and Reset is a future time or seconds-until-reset, you hit a normal rate window; jump to Fixes → Hitting a documented per-key/per-IP/per-endpoint limit. If Remaining is not present but 429s correlate with spikes in in-flight requests, jump to Fixes → Too much parallelism or burst traffic.
- Check whether your client is making the storm worse with retries.
grep -RniE 'retry|backoff|Retry-After|429' .
If you find immediate retries, fixed sleep(1) loops, or libraries configured with retries > 0 but no 429-specific handling, jump to Fixes → Retries ignore Retry-After or retry immediately.
- Confirm whether many services share one egress IP.
curl -s https://ifconfig.me
Run that from at least two affected workloads. If they return the same public IP and failures started across unrelated services, jump to Fixes → Shared NAT/egress IP is collapsing clients into one bucket.
- Check whether requests are unnecessarily uncached.
grep -RniE 'ETag|If-None-Match|Cache-Control|max-age|stale-while-revalidate' .
If the client repeatedly fetches identical resources without conditional requests or local caching, jump to Fixes → Missing caching or request coalescing.
- Verify whether the 429 comes from a gateway/CDN/WAF instead of the API origin.
curl -sS -D - -o /dev/null https://api.example.com/v1/resource | sed -n '1,20p'
Output like this points to an intermediary:
HTTP/2 429
server: cloudflare
cf-ray: 8f1d...
content-type: text/html
Or:
HTTP/1.1 429 Too Many Requests
Server: nginx
Via: 1.1 varnish
X-Cache: MISS
If headers/body look like HTML or branded error pages instead of your API's normal JSON error shape, jump to Fixes → Gateway/CDN/WAF is generating the 429.
Fixes
Retries ignore Retry-After or retry immediately
Use bounded exponential backoff with jitter, and honor Retry-After when present.
Node.js example:
async function sleep(ms) { return new Promise(r => setTimeout(r, ms)); }
function parseRetryAfter(h) {
if (!h) return null;
const n = Number(h);
if (Number.isFinite(n)) return n * 1000;
const t = Date.parse(h);
return Number.isFinite(t) ? Math.max(0, t - Date.now()) : null;
}
async function requestWith429Backoff(doRequest, maxAttempts = 5) {
for (let attempt = 1; attempt <= maxAttempts; attempt++) {
const res = await doRequest();
if (res.status !== 429) return res;
const ra = parseRetryAfter(res.headers.get('retry-after'));
const backoff = Math.min(30000, 500 * 2 ** (attempt - 1));
const jitter = Math.floor(Math.random() * 250);
await sleep((ra ?? backoff) + jitter);
}
throw new Error('429 persisted after retries');
}
Python requests example:
import random, time, requests
from email.utils import parsedate_to_datetime
from datetime import datetime, timezone
def retry_after_ms(v):
if not v:
return None
try:
return int(v) * 1000
except ValueError:
dt = parsedate_to_datetime(v)
return max(0, int((dt - datetime.now(timezone.utc)).total_seconds() * 1000))
for attempt in range(5):
r = requests.get("https://api.example.com/v1/resource", timeout=10)
if r.status_code != 429:
break
ra = retry_after_ms(r.headers.get("Retry-After"))
backoff = min(30000, 500 * (2 ** attempt))
jitter = random.randint(0, 250)
time.sleep(((ra or backoff) + jitter) / 1000)
Verify it worked:
grep -E '429|Retry-After' app.log | tail -20
You should see fewer repeated attempts per request and a drop in 429 volume.
Too much parallelism or burst traffic
Cap concurrency at the client or worker level. If you fan out work, queue it instead of firing all requests at once.
Shell example with xargs reducing parallelism from 50 to 5:
cat ids.txt | xargs -n1 -P5 -I{} curl -fsS https://api.example.com/v1/items/{}
Node.js with a simple concurrency limiter:
import pLimit from 'p-limit';
const limit = pLimit(5);
await Promise.all(ids.map(id => limit(() => fetch(`https://api.example.com/v1/items/${id}`))));
Kubernetes worker deployment example:
kubectl scale deployment sync-worker --replicas=3 -n prod
Verify it worked:
watch -n 5 'grep -c " 429 " /var/log/app.log; kubectl get hpa,pods -n prod'
429s should fall as in-flight requests stabilize.
Hitting a documented per-key, per-IP, per-endpoint, or concurrency limit
Read the headers exactly; do not assume one global bucket. Common patterns:
X-RateLimit-Limit/Remaining/Reset: usually fixed-window or sliding-window metadata.RateLimit-Limit/Remaining/Reset: newer standard-style fields;Resetmay be seconds until reset, not a Unix timestamp.Retry-After: wait this many seconds, or until the given HTTP date.
Quick parser for a live response:
curl -sS -D - -o /dev/null https://api.example.com/v1/resource | awk 'BEGIN{IGNORECASE=1}/^HTTP\/|^Retry-After:|^X-RateLimit-|^RateLimit-/{print}'
If one endpoint is expensive, spread calls over time or batch where the API supports it. If the limit is per token, use separate credentials for distinct workloads only if your API terms allow it; do not rotate tokens to evade limits.
Verify it worked:
curl -sS -D - -o /dev/null https://api.example.com/v1/resource | egrep -i 'RateLimit|X-RateLimit|Retry-After'
Remaining should stay above zero during normal traffic.
Shared NAT/egress IP is collapsing clients into one bucket
If multiple services leave through one public IP, a per-IP limiter can throttle all of them together. Split egress paths or route the noisy workload through a dedicated NAT/egress IP.
In Kubernetes, first identify the noisy namespace/workload and isolate it behind separate node pools or egress where your platform supports that. At minimum, reduce its concurrency immediately:
kubectl top pods -A --sort-by=cpu
kubectl scale deployment noisy-worker --replicas=1 -n prod
Then confirm public IPs from each workload:
kubectl exec -n prod deploy/api -- curl -s https://ifconfig.me
kubectl exec -n prod deploy/noisy-worker -- curl -s https://ifconfig.me
Verify it worked:
for p in api noisy-worker; do kubectl exec -n prod deploy/$p -- curl -s https://ifconfig.me; echo; done
Different egress IPs, or reduced traffic from the noisy workload, should stop cross-service 429s.
Missing caching or request coalescing
For repeated GETs, send conditional requests and cache successful responses locally or at your edge.
Client example using ETag:
etag=$(curl -fsSI https://api.example.com/v1/catalog | awk -F': ' 'BEGIN{IGNORECASE=1}/^ETag:/{print $2}' | tr -d '\r')
curl -sS -D - -H "If-None-Match: $etag" -o /dev/null https://api.example.com/v1/catalog
Expected healthy shape:
HTTP/1.1 304 Not Modified
ETag: "abc123"
Nginx reverse-proxy microcache for safe GETs:
proxy_cache_path /var/cache/nginx levels=1:2 keys_zone=apicache:10m max_size=1g inactive=60m use_temp_path=off;
server {
location /v1/catalog {
proxy_cache apicache;
proxy_cache_valid 200 30s;
proxy_cache_methods GET HEAD;
proxy_pass https://api.example.com;
add_header X-Cache-Status $upstream_cache_status;
}
}
Verify it worked:
curl -sSI https://your-proxy.example.com/v1/catalog | egrep 'HTTP/|X-Cache-Status|ETag'
You want 304 on conditional requests or X-Cache-Status: HIT on repeated reads.
Gateway/CDN/WAF is generating the 429
If the 429 is not from the API origin, fix the intermediary rule or bypass it for trusted API traffic.
First compare direct-to-origin versus public endpoint if you have an origin hostname or internal service DNS:
curl -sS -D - -o /dev/null https://origin-api.internal/v1/resource | sed -n '1,15p'
curl -sS -D - -o /dev/null https://api.example.com/v1/resource | sed -n '1,15p'
If origin returns 200/401/403 but public returns 429 with gateway headers, adjust rate-limit/WAF rules in your provider dashboard for the API path. In your provider's dashboard, look for rate limiting or bot/WAF rules on the API hostname/path (for example, DNS/CDN provider dashboard → Security/WAF or Rate limiting rules). Prefer path-specific rules like /v1/* over hostname-wide rules that also catch browser traffic.
Verify it worked:
curl -sS -D - -o /dev/null https://api.example.com/v1/resource | sed -n '1,20p'
The response should now have your API's normal headers/body shape rather than the gateway's error page.
Prevention
- Add a 429-specific metric and alert on both rate and retry amplification.
# Prometheus-style recording rule example
- record: app:http_429_ratio_5m
expr: sum(rate(http_client_requests_total{status="429"}[5m])) / sum(rate(http_client_requests_total[5m]))
- Log request IDs and rate-limit headers on every non-2xx response.
{"level":"warn","status":429,"path":"/v1/resource","retry_after":"17","ratelimit_remaining":"0","request_id":"8f2b..."}
- Pin client retry behavior in code and test it in CI with a mock 429.
curl -sS -D - -o /dev/null http://localhost:8080/test-429
# Assert your client sleeps >= Retry-After and caps attempts <= 5
- Put a concurrency ceiling in config, not just code defaults.
API_MAX_CONCURRENCY=5
API_MAX_RETRIES=5
API_BACKOFF_BASE_MS=500
API_BACKOFF_MAX_MS=30000
- Cache safe reads and coalesce duplicate in-flight requests.
const inflight = new Map();
function oncePerKey(key, fn) {
if (inflight.has(key)) return inflight.get(key);
const p = fn().finally(() => inflight.delete(key));
inflight.set(key, p);
return p;
}
- Track egress IPs for workloads that call third-party APIs, especially after network changes.
kubectl get pods -A -o name | xargs -I{} sh -c 'echo {}; kubectl exec {} -- curl -s https://ifconfig.me; echo'
That gives you a fast way to spot when unrelated services are sharing one public IP and one rate bucket.
This article was written by an AI system and published pending human review. Verify anything you intend to act on.
Have a project in mind?
Get an instant AI price estimate for it, or talk directly to our team.
One email a month on what we learn building with AI