Diagnose model timeouts inside your service and fix them fast
For developers debugging AI-backed APIs that hang, return 504s, or die with client timeout errors. This runbook shows what model timeouts look like from inside your service, how to tell upstream latency from your own timeout budget, and the exact config and code changes to fix the common causes.
TL;DR — When a model call times out, your service usually sees one of three shapes: your HTTP client raises a read/deadline error, your reverse proxy returns 504/499, or your worker gets killed before the model responds. The most common fix is to align timeout budgets across your app, proxy, and platform, then lower model latency with smaller outputs/streaming so the upstream call finishes inside that budget. Reading time: ~6 min
The scenario
You ship a Tuesday afternoon change that adds a longer prompt and structured JSON output to an endpoint that calls a model provider. Five minutes later, your API error rate jumps: clients report intermittent 504s, your app logs show request IDs with 30-60 second gaps, and some requests never log a response body at all. The provider status page is green, CPU on your service is normal, and the only obvious clue is that short prompts still work while larger requests fail. You need to answer one question fast: is the model slow, or is your own stack timing out first?
Symptoms
- Client-facing status codes:
504 Gateway Timeout502 Bad Gatewayafter a long wait499 Client Closed Requestin nginx when the caller gives up first408 Request Timeoutif your app/framework emits it explicitly
- App log errors from common HTTP clients:
requests.exceptions.ReadTimeout: HTTPSConnectionPool(host='api.example-llm.com', port=443): Read timed out. (read timeout=30)httpx.ReadTimeout: The read operation timed outaiohttp.client_exceptions.ServerTimeoutError: Timeout on reading data from socketcontext deadline exceedednet/http: request canceled (Client.Timeout exceeded while awaiting headers) - Reverse proxy / ingress logs:
upstream timed out (110: Connection timed out) while reading response header from upstream10.0.4.12 - - [01/Oct/2026:14:22:11 +0000] "POST /v1/chat HTTP/1.1" 504 167 "-" "my-client/1.2.3" rt=60.001 uct=0.002 urt=60.000 - Platform/runtime signals:
- Request duration pinned at a hard limit like
30.0s,60.0s, or100.0s - Worker/container restarts if the process is killed by a platform request timeout or liveness failure
- Request duration pinned at a hard limit like
- What users see:
- Spinner until failure, then generic "Something went wrong"
- Partial streamed output stops mid-sentence if the upstream or proxy closes the connection
Likely causes
| Cause | How common | Quick check |
|---|---|---|
| App HTTP client timeout is shorter than model latency | Very common | `grep -RniE 'timeout= |
| Reverse proxy / ingress timeout is shorter than app timeout | Very common | `nginx -T 2>/dev/null |
| Platform/request deadline kills long requests | Common | kubectl describe ingress <name> -n <ns> |
| Prompt/output size increased latency beyond your budget | Common | jq '.messages, .max_tokens // .max_output_tokens // empty' failing-request.json |
| No streaming; you wait for full completion before first byte | Common | `grep -RniE 'stream["'"' ]*:[ ]*true |
| Retry policy amplifies latency and causes timeout cascades | Less common | `grep -RniE 'retry |
Step-by-step diagnosis
-
Check whether failures snap to a hard duration.
grep -E ' 504 |ReadTimeout|deadline exceeded|timed out' app.log | tail -n 50If many failures cluster at exactly
30s,60s, or another round number, this is almost always a configured timeout, not random network loss. Jump to the fix that matches the layer: app client, proxy, or platform. -
Reproduce one failing request with timing breakdown.
curl -sS -o /tmp/resp.json -D /tmp/headers.txt -w 'code=%{http_code} connect=%{time_connect} ttfb=%{time_starttransfer} total=%{time_total}\n' -X POST http://localhost:8080/v1/chat -H 'content-type: application/json' --data @failing-request.json cat /tmp/headers.txtIf
time_starttransferis near your timeout andcode=504, your service/proxy waited and then gave up. If your app returns a JSON error with a client-library timeout message, jump to### App HTTP client timeout is shorter than model latency. -
Inspect app timeout settings in code and env.
grep -RniE 'timeout=|Client.Timeout|context.WithTimeout|READ_TIMEOUT|REQUEST_TIMEOUT|HTTPX_TIMEOUT' .This is your problem if you find values like
30,60, or framework defaults lower than real model latency. Typical examples:client = httpx.Client(timeout=30.0)client := &http.Client{Timeout: 30 * time.Second}Jump to
### App HTTP client timeout is shorter than model latency. -
Inspect reverse proxy / ingress timeouts.
nginx -T 2>/dev/null | grep -E 'proxy_read_timeout|proxy_send_timeout|send_timeout'Typical bad output:
proxy_read_timeout 60s; proxy_send_timeout 60s; send_timeout 60s;If your app timeout is higher than this, nginx is cutting the request first. Jump to
### Reverse proxy / ingress timeout is shorter than app timeout. -
Check your platform or load balancer request deadline. In Kubernetes, inspect ingress/controller annotations and service docs for your environment:
kubectl get ingress -A -o yaml | grep -nE 'timeout|proxy-read-timeout|proxy-send-timeout'If your platform has a hard request cap lower than model latency, this is your problem. Jump to
### Platform/request deadline kills long requests. -
Compare failing and successful payload sizes.
jq '{input_chars: (.messages | tostring | length), max_tokens: (.max_tokens // .max_output_tokens // 0), stream: .stream}' failing-request.jsonIf failures correlate with larger prompts, higher
max_tokens, or strict JSON/schema output, latency increased beyond your budget. Jump to### Prompt/output size increased latency beyond your budget. -
Check whether you stream the response.
grep -RniE 'stream["'"' ]*:[ ]*true|text/event-stream|EventSource|ReadableStream' .If you do not stream and clients/proxies wait for the full body, long generations are more likely to hit idle/read timeouts. Jump to
### No streaming; you wait for full completion before first byte. -
Check retry behavior around model calls.
grep -RniE 'retry|backoff|max_retries' .If a single user request can trigger multiple long upstream attempts, your p95/p99 will explode and saturate workers. Jump to
### Retry policy amplifies latency and causes timeout cascades.
Fixes
App HTTP client timeout is shorter than model latency
Set separate connect/read/write timeouts instead of one low global timeout.
Python httpx:
import httpx
timeout = httpx.Timeout(connect=5.0, read=120.0, write=30.0, pool=5.0)
client = httpx.Client(timeout=timeout)
Python requests:
resp = requests.post(url, json=payload, timeout=(5, 120))
Go:
client := &http.Client{
Timeout: 0, // do not use a single total timeout here
}
ctx, cancel := context.WithTimeout(r.Context(), 120*time.Second)
defer cancel()
req = req.WithContext(ctx)
resp, err := client.Do(req)
Trade-off: raising read timeout blindly can tie up workers. If concurrency is high, pair this with streaming or async job handling.
Verify it worked:
curl -sS -o /dev/null -w 'code=%{http_code} ttfb=%{time_starttransfer} total=%{time_total}\n' -X POST http://localhost:8080/v1/chat -H 'content-type: application/json' --data @failing-request.json
Reverse proxy / ingress timeout is shorter than app timeout
For nginx, raise the upstream read timeout above your app/model budget.
location /v1/chat {
proxy_pass http://app_upstream;
proxy_connect_timeout 5s;
proxy_send_timeout 30s;
proxy_read_timeout 180s;
send_timeout 180s;
proxy_buffering off;
}
Reload safely:
nginx -t && sudo systemctl reload nginx
If you use Kubernetes ingress-nginx, set annotations on the Ingress:
metadata:
annotations:
nginx.ingress.kubernetes.io/proxy-read-timeout: "180"
nginx.ingress.kubernetes.io/proxy-send-timeout: "180"
Trade-off: longer proxy timeouts increase connection occupancy. If you expect long generations, prefer streaming so the connection is active and users see progress.
Verify it worked:
curl -I -sS https://your-service.example.com/v1/chat
Then send a known slow request and confirm you no longer get a 504 at the old cutoff.
Platform/request deadline kills long requests
If your platform enforces a hard request cap, move long model work off the synchronous request path.
Pattern:
- API accepts request and returns
202 Acceptedwith a job ID. - Worker performs model call.
- Client polls
/jobs/:idor consumes a websocket/SSE stream.
Minimal response shape:
{"job_id":"9f6c7c2e","status":"queued"}
If your platform supports configurable request timeouts, raise them in your provider dashboard or deployment manifest. In Kubernetes, also check probes so long requests do not trip liveness.
⚠️ Changing platform timeouts can increase user-visible waits and tie up instances. If your current hard cap is below real model latency, queueing is safer than just increasing the cap.
Verify it worked:
curl -sS -X POST http://localhost:8080/v1/chat -H 'content-type: application/json' --data @failing-request.json | jq
You should get 202 plus a job ID instead of a timeout.
Prompt/output size increased latency beyond your budget
Reduce generation cost before increasing timeouts.
Actions:
- Lower output cap:
{"max_tokens": 400} - Trim conversation history before sending upstream.
- Remove unnecessary schema constraints or huge tool definitions.
- Split one large generation into smaller calls.
Quick payload diff:
jq '{chars: (.messages | tostring | length), max_tokens: (.max_tokens // .max_output_tokens // 0)}' failing-request.json passing-request.json
Trade-off: lower token caps can truncate useful answers. If correctness matters, combine a lower cap with continuation logic or summarization of prior turns.
Verify it worked:
for f in passing-request.json failing-request.json; do echo "$f"; curl -sS -o /dev/null -w 'code=%{http_code} total=%{time_total}\n' -X POST http://localhost:8080/v1/chat -H 'content-type: application/json' --data @$f; done
No streaming; you wait for full completion before first byte
Enable streaming from the model provider and pass it through to the client.
Server response headers:
proxy_buffering off;
add_header X-Accel-Buffering no;
HTTP response:
Content-Type: text/event-stream
Cache-Control: no-cache
Connection: keep-alive
Node/Express shape:
res.setHeader("Content-Type", "text/event-stream");
res.setHeader("Cache-Control", "no-cache");
res.setHeader("Connection", "keep-alive");
res.flushHeaders();
Trade-off: streaming improves time-to-first-byte and avoids some idle timeouts, but complicates retries and response validation because the body is no longer atomic.
Verify it worked:
curl -N -H 'Accept: text/event-stream' -X POST http://localhost:8080/v1/chat -H 'content-type: application/json' --data @failing-request.json
You should see chunks arrive before the full completion finishes.
Retry policy amplifies latency and causes timeout cascades
Cap retries and do not retry long non-idempotent generations automatically.
Example policy:
{"max_retries":1,"backoff_ms":250,"retry_on":[429,500,502,503]}
Do not retry on local read timeout unless you know the upstream operation is safe to duplicate or you supply an idempotency key.
Trade-off: fewer retries reduce tail latency and worker pileups, but may lower success rate during brief provider blips. For interactive endpoints, latency usually matters more than squeezing out one extra retry.
Verify it worked:
grep -E 'retry|attempt=' app.log | tail -n 20
You should no longer see multiple long attempts for one request ID.
Prevention
-
Add per-hop latency logging with one request ID.
{"request_id":"abc123","model_ms":42871,"app_ms":42920,"status":200,"upstream_status":200}Alert when
model_msapproaches 80% of your smallest timeout budget. -
Pin timeout budgets in config, not scattered literals.
MODEL_CONNECT_TIMEOUT_MS=5000 MODEL_READ_TIMEOUT_MS=120000 PROXY_READ_TIMEOUT_S=180 REQUEST_DEADLINE_MS=125000In CI, fail if app deadline exceeds proxy/platform deadline.
-
Add a synthetic slow-request test in CI/staging.
curl -sS -o /dev/null -w 'code=%{http_code} total=%{time_total}\n' -X POST "$STAGING_URL/v1/chat" -H 'content-type: application/json' --data @tests/fixtures/slow-model-request.jsonFail the pipeline if it returns
504,502, or exceeds your SLO. -
Track prompt size and output cap as first-class metrics.
model_request_input_chars model_request_max_tokens model_request_stream_enabledCorrelate these with timeout rate; this catches regressions after prompt/template changes.
-
For long-running endpoints, default to streaming or async jobs.
{"mode":"stream"}Use synchronous full-body responses only for requests that reliably complete well inside your smallest timeout.
-
Put an explicit timeout matrix in the repo.
app client read timeout: 120s app request deadline: 125s nginx proxy_read_timeout: 180s platform hard cap: 300sKeep it in
docs/timeouts.mdand review it whenever prompt size, model, or response format changes.
This article was written by an AI system and published pending human review. Verify anything you intend to act on.
Have a project in mind?
Get an instant AI price estimate for it, or talk directly to our team.
One email a month on what we learn building with AI