Webhook delivered twice: diagnose retries and build idempotent handlers
For developers debugging duplicate webhook processing in production. This runbook shows how to tell apart provider retries, redirect/rewrite mistakes, multiple subscriptions, and race conditions, then harden your handler with idempotency and fast-ack patterns.
TL;DR — If a webhook "fired twice", the most common reality is that the sender retried because your endpoint did not return a fast 2xx, or your app processed the same event twice because the handler is not idempotent. First confirm whether the duplicate requests share the same event ID; then fix the transport issue or dedupe on that ID with a unique constraint and return 2xx before slow work. Reading time: ~6 min
The scenario
You merge a routine deploy after lunch, and thirty minutes later support pings you: customers are getting two receipts, two Slack messages, or two CRM updates for a single action. Your app logs show the webhook route hit twice within seconds, but the upstream provider dashboard insists it sent one event successfully. You grep the logs and find one request took 12 seconds because it waited on a downstream API, while another completed in 40 ms. Now you need to answer two questions fast: did the sender retry, or did your system process the same event twice?
Symptoms
- Two side effects for one business action: duplicate emails, duplicate rows, double shipment creation, repeated message posts.
- Access logs show the same webhook path twice, often within seconds:
203.0.113.10 - - [01/Oct/2026:14:03:11 +0000] "POST /webhooks/provider HTTP/1.1" 500 72 "-" "Webhook-Client/1.0"
203.0.113.10 - - [01/Oct/2026:14:03:16 +0000] "POST /webhooks/provider HTTP/1.1" 200 2 "-" "Webhook-Client/1.0"
- Or both requests are
200, but your app created the side effect twice because two workers raced on the same event. - Provider delivery log shows wording like
timeout,connection reset,non-2xx response,retrying, or multiple delivery attempts for the same event ID. - App logs contain duplicate event identifiers, signatures, or request IDs:
received webhook event_id=evt_9f2c3 type=invoice.paid
received webhook event_id=evt_9f2c3 type=invoice.paid
- Redirects on the webhook endpoint:
HTTP/1.1 301 Moved Permanently
Location: https://api.example.com/webhooks/provider/
- Reverse proxy/app server timeout messages:
upstream timed out (110: Connection timed out) while reading response header from upstream
- Database errors after adding dedupe correctly:
ERROR: duplicate key value violates unique constraint "webhook_events_provider_event_id_key"
That last one is usually good news: it means the duplicate was blocked.
Likely causes
| Cause | How common | Quick check |
|---|---|---|
| Sender retried because your endpoint did not return a fast 2xx | Very common | `grep -E 'POST /webhooks/provider |
| Handler is not idempotent; same event processed twice by app/workers | Very common | psql "$DATABASE_URL" -c "\d webhook_events" |
Webhook endpoint redirects (301/302/307/308) or rewrites path/scheme | Common | curl -i -X POST https://api.example.com/webhooks/provider -d '{}' |
| Multiple webhook subscriptions/endpoints point to the same handler | Common | In your provider's dashboard, open Webhooks/Developers and count active endpoints for the same event type |
| Queue/consumer redelivery after worker crash or ack-after-work pattern | Less common | `grep -E 'requeued |
| Upstream/load balancer timeout shorter than handler runtime | Less common | curl -s -o /dev/null -w 'time_total=%{time_total} code=%{http_code}\n' https://api.example.com/webhooks/provider |
Step-by-step diagnosis
- Check whether the duplicates share the same provider event ID.
grep -R "event_id=" /var/log/app.log | tail -n 100
If you see the same event_id twice, this is a retry or your app reprocessed the same event. Jump to Fixes → Sender retried because your endpoint did not return a fast 2xx or Fixes → Handler is not idempotent. If the IDs differ, jump to step 4.
- Check the HTTP status and latency for the webhook route.
grep 'POST /webhooks/provider' /var/log/nginx/access.log | tail -n 20
This is your problem if you see >=300, 499, 500, 502, 503, 504, or request times near your proxy timeout. Example shape:
203.0.113.10 - - [01/Oct/2026:14:03:11 +0000] "POST /webhooks/provider HTTP/1.1" 200 2 "-" "Webhook-Client/1.0" rt=11.982 ua="Webhook-Client/1.0"
A 200 after ~12s can still trigger retries if the sender timed out earlier. Jump to Fixes → Sender retried because your endpoint did not return a fast 2xx.
- Test for redirects on the exact webhook URL.
curl -i -X POST https://api.example.com/webhooks/provider -H 'Content-Type: application/json' -d '{}'
This is your problem if the response starts with 301, 302, 307, or 308, especially with a Location: header changing host, scheme, or trailing slash.
HTTP/1.1 308 Permanent Redirect
Location: https://api.example.com/webhooks/provider/
Jump to Fixes → Webhook endpoint redirects or rewrites path/scheme.
-
Check for multiple active subscriptions for the same event type. Open your provider's dashboard and inspect the webhook endpoint list. If you find two endpoints hitting the same URL, or one old endpoint plus one new endpoint after a migration, this is your problem. Jump to Fixes → Multiple webhook subscriptions/endpoints point to the same handler.
-
Verify whether your app has durable dedupe storage.
psql "$DATABASE_URL" -c "\d+ webhook_events"
This is your problem if the table does not exist, or there is no unique index on provider + event ID. Jump to Fixes → Handler is not idempotent; same event processed twice by app/workers.
- Check queue/worker redelivery behavior.
grep -E 'requeued|redeliver|nack|visibility timeout|consumer cancelled' /var/log/app.log | tail -n 50
If you see the same event/job requeued after a crash or timeout, jump to Fixes → Queue/consumer redelivery after worker crash or ack-after-work pattern.
- Compare handler runtime to proxy/LB timeouts.
curl -s -o /dev/null -w 'connect=%{time_connect} starttransfer=%{time_starttransfer} total=%{time_total} code=%{http_code}\n' https://api.example.com/healthz
Then inspect your proxy config for proxy_read_timeout, fastcgi_read_timeout, ingress timeout, or LB idle timeout. If the handler routinely runs close to or beyond those limits, jump to Fixes → Upstream/load balancer timeout shorter than handler runtime.
Fixes
Sender retried because your endpoint did not return a fast 2xx
Return 2xx immediately after signature verification and durable enqueue/dedupe write. Move slow work out of the request path.
Postgres dedupe table:
CREATE TABLE IF NOT EXISTS webhook_events (
provider text NOT NULL,
event_id text NOT NULL,
received_at timestamptz NOT NULL DEFAULT now(),
payload jsonb NOT NULL,
PRIMARY KEY (provider, event_id)
);
Node/Express pattern:
app.post('/webhooks/provider', express.raw({ type: '*/*' }), async (req, res) => {
const sig = req.get('X-Signature');
verifySignatureOrThrow(req.body, sig);
const event = JSON.parse(req.body.toString('utf8'));
const inserted = await db.query(
`INSERT INTO webhook_events(provider, event_id, payload)
VALUES ($1, $2, $3)
ON CONFLICT DO NOTHING`,
['provider', event.id, event]
);
res.status(200).send('ok');
if (inserted.rowCount === 0) return; // duplicate delivery
await queue.publish('webhook-jobs', { provider: 'provider', eventId: event.id });
});
If enqueue must happen before 200, write to the database first, then let a background poller publish jobs from the table.
Verify it worked:
grep -R "event_id=evt_" /var/log/app.log | tail -n 20
You should see duplicate deliveries logged once as duplicate ignored and no duplicate side effects.
Handler is not idempotent; same event processed twice by app/workers
Add a unique constraint keyed by the provider event ID, and make downstream writes idempotent too.
For an existing table:
ALTER TABLE webhook_events ADD COLUMN IF NOT EXISTS provider text;
ALTER TABLE webhook_events ADD COLUMN IF NOT EXISTS event_id text;
CREATE UNIQUE INDEX CONCURRENTLY IF NOT EXISTS webhook_events_provider_event_id_key
ON webhook_events(provider, event_id);
For business rows, use upserts instead of blind inserts:
INSERT INTO invoices(provider_event_id, customer_id, amount_cents)
VALUES ($1, $2, $3)
ON CONFLICT (provider_event_id) DO NOTHING;
If two workers can race, claim work atomically:
UPDATE webhook_events
SET processing_started_at = now()
WHERE provider = $1 AND event_id = $2 AND processing_started_at IS NULL
RETURNING *;
Only the worker that gets a row proceeds.
Verify it worked:
psql "$DATABASE_URL" -c "INSERT INTO webhook_events(provider,event_id,payload) VALUES ('provider','evt_test','{}') ON CONFLICT DO NOTHING RETURNING event_id;"
The second run should return zero rows.
Webhook endpoint redirects or rewrites path/scheme
Use the exact final URL in the provider config and stop redirecting webhook requests.
nginx example: serve the webhook path directly and avoid trailing-slash rewrites.
server {
listen 443 ssl http2;
server_name api.example.com;
location = /webhooks/provider {
proxy_pass http://app_upstream;
proxy_set_header Host $host;
proxy_set_header X-Forwarded-Proto $scheme;
}
}
If you have a global HTTP→HTTPS redirect, keep it, but register the HTTPS URL in the provider. Do not rely on the sender following redirects; some do not, some retry, some convert method semantics badly.
Verify it worked:
curl -i -X POST https://api.example.com/webhooks/provider -d '{}'
The first line should be HTTP/1.1 200 or 204, not 3xx.
Multiple webhook subscriptions/endpoints point to the same handler
Delete or disable duplicate endpoints in the provider dashboard, or split them by event type and destination.
If you cannot change the provider immediately, log and reject unexpected sources by secret or signature key. Keep one secret per endpoint so you can identify which subscription sent the request.
Example env layout:
WEBHOOK_SECRET_PRIMARY=...
WEBHOOK_SECRET_OLD=...
Then log which secret matched and remove the old endpoint after traffic drains.
Verify it worked: In the provider dashboard, trigger one test event and confirm exactly one delivery attempt reaches your access log.
Queue/consumer redelivery after worker crash or ack-after-work pattern
Acknowledge the queue message only after the side effect is made idempotent, not merely after starting work. If your queue supports visibility timeout, set it longer than worst-case processing time or split the job.
Pseudo-flow:
webhook request -> insert dedupe row -> 200 OK -> enqueue job
worker -> claim event atomically -> perform idempotent side effect -> ack queue message
If the worker crashes after the side effect but before ack, the redelivery should hit the dedupe guard and no-op.
Verify it worked:
grep -E 'duplicate ignored|already processed' /var/log/app.log | tail -n 20
A forced worker restart should not create duplicate side effects.
Upstream/load balancer timeout shorter than handler runtime
Raise the relevant timeout only if you cannot move work out of band immediately. The better fix is still fast-ack plus background processing.
nginx example:
location = /webhooks/provider {
proxy_connect_timeout 5s;
proxy_send_timeout 30s;
proxy_read_timeout 30s;
proxy_pass http://app_upstream;
}
If you run behind a cloud load balancer or ingress, align its idle/request timeout with the proxy and app server. Mismatched layers cause confusing retries.
⚠️ Raising timeouts can increase connection occupancy and amplify overload during incidents. Do this only after confirming the handler cannot be shortened today.
Verify it worked:
grep 'POST /webhooks/provider' /var/log/nginx/access.log | tail -n 20
Request times should stay comfortably below the shortest timeout in the path.
Prevention
- Add a database-backed dedupe key in CI migrations, not as an ad hoc hotfix:
CREATE TABLE IF NOT EXISTS webhook_events (
provider text NOT NULL,
event_id text NOT NULL,
payload jsonb NOT NULL,
received_at timestamptz NOT NULL DEFAULT now(),
processed_at timestamptz,
PRIMARY KEY (provider, event_id)
);
- Add a metric for duplicate deliveries and non-2xx webhook responses. Example log-based counters to alert on:
webhook_requests_total{route="/webhooks/provider",status="5xx"}
webhook_duplicate_events_total{provider="provider"}
webhook_handler_duration_seconds
Alert if 5xx > 0 for 5 minutes or p95 duration exceeds the provider retry threshold.
- Add an integration test that posts the exact same signed payload twice and asserts one side effect:
./scripts/send-webhook.sh fixtures/invoice-paid.json
./scripts/send-webhook.sh fixtures/invoice-paid.json
psql "$DATABASE_URL" -c "SELECT count(*) FROM invoices WHERE provider_event_id='evt_fixture_123';"
Expected count: 1.
- Pin the exact webhook URL and reject redirects in smoke tests:
test "$(curl -s -o /dev/null -w '%{http_code}' -X POST https://api.example.com/webhooks/provider -d '{}')" != "301"
- Log provider event ID, signature verification result, request latency, and dedupe decision in one structured line:
{"route":"/webhooks/provider","event_id":"evt_9f2c3","verified":true,"dedupe":"inserted","latency_ms":42,"status":200}
Without that, you are guessing during the incident.
- Keep webhook handlers boring: verify signature, persist, ack, enqueue. Any direct call to SMTP, third-party APIs, PDF generation, or long DB transactions in the request path should fail code review unless there is a written reason.
This article was written by an AI system and published pending human review. Verify anything you intend to act on.
Have a project in mind?
Get an instant AI price estimate for it, or talk directly to our team.
One email a month on what we learn building with AI