Designing AI Features That Degrade Gracefully Instead of Hard Failing
For developers shipping AI-backed product features, this guide shows how to keep the user journey working when the model, vector store, rate limit, or upstream provider is slow or down. You’ll get a concrete architecture, one end-to-end example, and deployable patterns for timeouts, fallbacks, circuit breakers, and observability.
TL;DR — AI features fail in more ways than normal code: model timeouts, quota exhaustion, malformed output, retrieval misses, and provider outages. The most useful default is to treat AI as an optional dependency behind a strict deadline: cap latency, validate output, and fall back to a deterministic path or partial UX instead of letting the whole request fail. Reading time: ~7 min
What it is and where it sits
"Degrades instead of failing" means the product still completes the user’s primary task when the AI subsystem cannot. The feature may become less smart, less personalized, or less automated, but it does not turn a normal request into a 500, spinner, or broken page.
In practice, this sits at the boundary between your application and one or more probabilistic dependencies:
- LLM API
- embedding model
- vector database / search index
- moderation / classification model
- OCR / speech-to-text / translation service
- post-processing step that parses model output into structured data
The key architectural point: the AI path should usually be a sidecar path to the main business flow, not the only path.
Typical request flow:
Browser / Mobile
|
v
API gateway / app server
|
+--> deterministic business logic -------------------+
| |
+--> AI orchestrator --> retrieval --> model --> validator
| | |
+-------------+---------+
|
fallback selector
|
v
response assembler / UI state
What it replaces: usually hand-written heuristics, search, templates, ranking rules, or human-only workflows.
What talks to it: the app server, background jobs, queue consumers, or edge workers.
Where it lives: ideally behind an internal interface like Summarizer, Classifier, or AnswerService, not scattered across controllers. If every route directly calls an LLM SDK, you have no place to enforce deadlines, retries, schema validation, or fallback policy.
A useful mental model is "progressive enhancement for backend intelligence." The core action must survive without the AI result.
How it actually works
Use one concrete pattern: deadline + validation + fallback ladder.
Example feature: support ticket composer. User types a message; your app suggests a reply draft for the agent. If AI is unavailable, the agent should still see the ticket, canned responses, and relevant KB links.
End-to-end example
- Agent opens
/tickets/4821. - App fetches ticket data from Postgres and renders the page immediately.
- Frontend calls
GET /api/tickets/4821/reply-suggestion. - Backend starts a 1200 ms deadline for the entire AI path.
- Backend retrieves top 5 KB documents from Postgres full-text search or a vector index.
- Backend calls the model with the ticket text and retrieved docs.
- Backend validates the response against a schema:
tone,summary,draft_reply. - If any step misses the deadline, returns invalid JSON, or gets a 429/5xx, backend falls back.
- Fallback returns:
- top KB links
- 3 canned response templates based on deterministic rules
- status field
degraded: true
- Frontend renders either the AI draft or the fallback panel. No spinner forever, no blank module, no page-level failure.
Concrete failure modes and what happens:
- Retrieval index unavailable: skip retrieval, try model with only ticket text if budget remains.
- Model returns malformed JSON: reject it; do not try to "kind of parse" unless you can prove safety.
- Provider returns 429: open circuit for 30-60 seconds and stop sending more traffic there.
- Total deadline exceeded: return fallback immediately.
A realistic HTTP shape from an upstream model gateway under rate limit might look like:
$ curl -i https://ai-gateway.internal/v1/chat -H 'Content-Type: application/json' -d @req.json
HTTP/1.1 429 Too Many Requests
content-type: application/json
retry-after: 15
x-request-id: 6f2a4d1b7f
{"error":{"type":"rate_limit","message":"quota exceeded"}}
Your app should translate that into an internal typed error like ErrRateLimited, not leak provider-specific text through the stack.
Likewise, a timeout should be explicit in logs and metrics. Good log shape:
{"level":"warn","route":"/api/tickets/:id/reply-suggestion","ticket_id":4821,"ai_stage":"model","deadline_ms":1200,"elapsed_ms":1207,"fallback":"canned_templates","error":"context deadline exceeded","request_id":"9d8c2e7a"}
The mechanism is not just retries. Retries often make AI incidents worse by multiplying load against a struggling provider. Prefer this order:
- very short timeout per stage
- one total request deadline
- schema validation
- circuit breaker on repeated upstream failures
- fallback ladder
- background retry only if the result is non-interactive and can arrive later
When to use it (and when not to)
Use degradation when the AI output improves the experience but is not the only way to complete the user’s job.
| Scenario | Recommendation |
|---|---|
| Suggested reply, summary, tags, ranking hint, translation preview | Yes. Treat AI as optional enhancement with strict deadline and fallback UI. |
| Fraud decision, medical advice, legal filing, payment authorization | Usually no. Fail closed or require human review; do not silently degrade to a weaker decision path. |
| Search relevance boost on top of keyword search | Yes. Keep deterministic search as baseline; AI reranks if available. |
| OCR for document ingestion where no text means no workflow | Partial. Accept upload, queue OCR async, show "processing" state; don’t fake completion. |
| Internal admin tool used by a few operators | Maybe. Simpler fallback may be enough; full circuit-breaker machinery may be overkill. |
| Feature with p95 latency budget under 300 ms | Probably not with remote LLM calls in-band. Use local heuristics or async generation. |
You probably don’t need this if:
- the feature is already asynchronous and users naturally tolerate delay
- the AI result is non-critical and can simply be omitted with a small UI note
- traffic is low enough that manual recovery is acceptable
- the deterministic fallback is so poor that it creates user harm or false confidence
Trade-offs
Every benefit costs something.
- Higher availability for the user → costs more code paths. You now own primary path, AI path, and fallback path.
- Lower incident blast radius → costs observability work. You need metrics by stage: retrieval, model, validation, fallback rate, circuit state.
- Predictable latency → costs possibly worse AI quality. Tight deadlines mean you will drop slow-but-good responses.
- Provider outage resilience → costs duplicate implementations. A real fallback often means maintaining templates, rules, or search indexes.
- Safer structured outputs → costs rejected responses. Strict schema validation will discard some otherwise usable model output.
- Less vendor lock-in → costs abstraction discipline. If you normalize request/response contracts, you may avoid provider-specific features.
- Controlled spend → costs product complexity. Caching, token budgets, and selective invocation require policy decisions, not just SDK calls.
Two edge cases experienced teams ask about:
Silent degradation can hide incidents
If the fallback rate jumps from 2% to 40% and the UI still works, your team may miss a real outage. Alert on fallback rate and circuit-open duration, not just 5xx rate.
Fallback quality can become stale
Canned templates and deterministic rules rot. Put them under the same product review as the AI path, or users will learn that "fallback mode" means "bad mode."
In practice
Example 1: Node.js/TypeScript endpoint with deadline, validation, and fallback
import express from "express";
import { z } from "zod";
const app = express();
const ReplySchema = z.object({
tone: z.enum(["empathetic", "neutral", "direct"]),
summary: z.string().min(1).max(500),
draft_reply: z.string().min(1).max(4000)
});
async function fetchKbLinks(ticketId: string) {
return [{ title: "Reset password", url: "/kb/reset-password" }];
}
async function cannedTemplates(ticketText: string) {
if (/refund/i.test(ticketText)) {
return ["I understand your concern about the refund. Let me check the order details."];
}
return ["Thanks for contacting support. I’m reviewing the details now."];
}
async function callModel(prompt: unknown, signal: AbortSignal) {
const res = await fetch("http://ai-gateway.internal/v1/chat", {
method: "POST",
headers: { "content-type": "application/json" },
body: JSON.stringify(prompt),
signal
});
if (res.status === 429) throw new Error("rate_limited");
if (!res.ok) throw new Error(`upstream_${res.status}`);
return res.json();
}
app.get("/api/tickets/:id/reply-suggestion", async (req, res) => {
const deadlineMs = 1200;
const ac = new AbortController();
const timer = setTimeout(() => ac.abort(), deadlineMs);
const ticketText = String(req.query.ticket_text || "");
try {
const kb = await fetchKbLinks(req.params.id);
const raw = await callModel({ ticketText, kb }, ac.signal);
const parsed = ReplySchema.parse(raw);
res.json({ degraded: false, source: "ai", ...parsed, kb_links: kb });
} catch (err) {
const kb = await fetchKbLinks(req.params.id);
const templates = await cannedTemplates(ticketText);
res.status(200).json({
degraded: true,
source: "fallback",
reason: String(err),
kb_links: kb,
templates
});
} finally {
clearTimeout(timer);
}
});
app.listen(3000);
What it does: enforces a hard 1200 ms budget, validates model output, and returns a typed fallback payload instead of a 500. Gotcha: AbortController only works if every downstream call honors cancellation; if your DB client or SDK ignores it, your server can still burn resources after the response is sent.
Example 2: nginx upstream timeouts to protect the app from hanging AI calls
upstream ai_gateway {
server 10.0.12.15:8080 max_fails=3 fail_timeout=30s;
keepalive 32;
}
server {
listen 443 ssl http2;
server_name app.internal;
location /internal/ai/ {
proxy_pass http://ai_gateway/;
proxy_http_version 1.1;
proxy_set_header Connection "";
proxy_set_header X-Request-Id $request_id;
proxy_connect_timeout 200ms;
proxy_send_timeout 2s;
proxy_read_timeout 2s;
send_timeout 2s;
proxy_next_upstream error timeout http_502 http_503 http_504;
proxy_next_upstream_tries 1;
}
}
What it does: puts hard network bounds around your internal AI gateway so requests fail fast enough for the application to fall back. Gotcha: proxy_next_upstream can accidentally retry non-idempotent requests; keep tries low and know whether your AI endpoint charges or mutates state per call.
Example 3: Prometheus alert on fallback-rate spike
groups:
- name: ai-degradation
rules:
- alert: AIFallbackRateHigh
expr: sum(rate(ai_requests_total{result="fallback"}[5m])) / sum(rate(ai_requests_total[5m])) > 0.15
for: 10m
labels:
severity: warning
annotations:
summary: "AI fallback rate above 15%"
description: "Users are being served degraded responses for more than 15% of AI requests over 10 minutes."
What it does: catches the case where the app still returns 200s but users are mostly seeing fallback behavior. Gotcha: segment by route and tenant if one noisy feature or customer can dominate the metric.
Further reading
- Martin Fowler — Circuit Breaker
- Google SRE Book — Handling Overload
- MDN HTTP docs — 429 Too Many Requests
- OpenTelemetry Specification — Semantic Conventions
- RFC 9110 — HTTP Semantics
This article was written by an AI system and published pending human review. Verify anything you intend to act on.
Have a project in mind?
Get an instant AI price estimate for it, or talk directly to our team.
One email a month on what we learn building with AI