Set up queryable structured logs for incident response fast
For developers who need logs they can filter under pressure, not grep through line noise. This shows how to emit JSON logs from an app, ship them with a standard collector, and query them locally or in your log backend with stable field names during an incident.
TL;DR — If your logs are plain text, incidents turn into regex archaeology. Emit one JSON object per line with fixed keys like
timestamp,level,service,trace_id,request_id, andmessage, then ship stdout/stderr with a collector that preserves those fields; the most common fix is to stop embedding JSON inside strings and log actual JSON lines. Reading time: ~5 min
Goal
When you finish, your service writes one valid JSON log event per line, your collector forwards those fields without flattening them into an opaque message, and you can run a fielded query during an incident such as "all level=error events for service=api with trace_id=... in the last 15 minutes" and get usable results immediately.
Prerequisites
- Shell access to the host or container running your app
- Permission to change app logging config and restart the service
- Docker Engine if you want the local demo path — check with
docker --version jq1.6+ for validation — check withjq --versioncurl7.70+ for test traffic — check withcurl --version- A place to send logs: either a local OpenSearch/Elasticsearch-compatible endpoint, Loki, or your existing log backend
- Your service name, environment name, and one request path you can hit safely in production or staging
- If using systemd/journald:
journalctlaccess
Steps
Step 1: Emit one JSON object per line from the app
Use a stable schema. Do not log free-form prefixes like INFO [api] before the JSON.
For a Node.js service using pino:
npm install pino
import pino from 'pino';
export const logger = pino({
level: process.env.LOG_LEVEL || 'info',
base: {
service: process.env.SERVICE_NAME || 'api',
environment: process.env.NODE_ENV || 'production'
},
timestamp: pino.stdTimeFunctions.isoTime,
formatters: {
level: (label) => ({ level: label })
}
});
logger.info({ request_id: 'req-123', trace_id: '4bf92f3577b34da6a3ce929d0e0e4736' }, 'user login started');
logger.error({ err: { type: 'Error', message: 'db timeout' }, request_id: 'req-123' }, 'login failed');
Example output shape:
{"level":"info","timestamp":"2026-10-01T12:00:00.123Z","service":"api","environment":"production","request_id":"req-123","trace_id":"4bf92f3577b34da6a3ce929d0e0e4736","message":"user login started"}
{"level":"error","timestamp":"2026-10-01T12:00:01.456Z","service":"api","environment":"production","err":{"type":"Error","message":"db timeout"},"request_id":"req-123","message":"login failed"}
What you should see when this succeeds: each log line is valid JSON by itself, with no text before or after the object.
Step 2: Validate the log format before shipping it
If you already have logs on stdout, validate the last 20 lines.
For Docker:
docker logs --tail 20 your-container-name 2>&1 | jq -c . >/dev/null && echo OK
For systemd:
journalctl -u your-service -n 20 -o cat | jq -c . >/dev/null && echo OK
If a line is not valid JSON, jq exits non-zero and prints an error like:
parse error: Invalid numeric literal at line 1, column 5
What you should see when this succeeds: OK and no parse errors.
Step 3: Add request and trace identifiers at ingress
If your app sits behind nginx, pass through a request ID and any incoming trace header.
proxy_set_header X-Request-Id $request_id;
proxy_set_header Traceparent $http_traceparent;
log_format json_combined escape=json '{"timestamp":"$time_iso8601","remote_addr":"$remote_addr","request_id":"$request_id","method":"$request_method","uri":"$request_uri","status":$status,"request_time":$request_time,"upstream_response_time":"$upstream_response_time"}';
access_log /var/log/nginx/access.json json_combined;
Reload nginx:
nginx -t && sudo systemctl reload nginx
What you should see when this succeeds: nginx: configuration file ... test is successful and new access log lines are JSON with a request_id field.
Step 4: Ship logs with a collector that preserves JSON fields
If you want a generic local path, run Fluent Bit and forward JSON logs to stdout first so you can verify parsing before sending to your backend.
Create fluent-bit.conf:
[SERVICE]
Flush 1
Daemon Off
Log_Level info
[INPUT]
Name tail
Path /var/log/app/*.log
Parser json
Tag app.*
Read_from_Head true
[PARSER]
Name json
Format json
Time_Key timestamp
Time_Format %Y-%m-%dT%H:%M:%S.%L%z
Time_Keep On
[OUTPUT]
Name stdout
Match app.*
Format json_lines
Run it:
docker run --rm -v $(pwd)/fluent-bit.conf:/fluent-bit/etc/fluent-bit.conf -v /var/log/app:/var/log/app cr.fluentbit.io/fluent/fluent-bit:3.1 -c /fluent-bit/etc/fluent-bit.conf
You should see output shaped like:
{"date":1769323200.123,"service":"api","environment":"production","level":"error","request_id":"req-123","message":"login failed"}
What you should see when this succeeds: collector output is still structured JSON, not a single log field containing escaped JSON text.
Step 5: Point the collector at your log backend
Use the backend you already have. The key requirement is that fields arrive as fields.
For OpenSearch/Elasticsearch-compatible HTTP output in Fluent Bit, replace the output section with:
[OUTPUT]
Name es
Match app.*
Host logs.example.internal
Port 9200
Index app-logs
Logstash_Format On
Suppress_Type_Name On
Generate_ID On
For Loki output in Fluent Bit:
[OUTPUT]
Name loki
Match app.*
Host logs.example.internal
Port 3100
Labels service=$service,level=$level,environment=$environment
Label_keys service,level,environment
Line_Format json
Restart the collector with the updated config.
What you should see when this succeeds: no connection errors in collector logs, and new events appear in the backend within a few seconds.
Step 6: Run one incident-grade query now, before you need it
Use a query that depends on structured fields, not substring matches.
OpenSearch/Elasticsearch-compatible example:
curl -s http://logs.example.internal:9200/app-logs*/_search -H 'Content-Type: application/json' -d '{
"size": 5,
"sort": [{"timestamp": {"order": "desc"}}],
"query": {
"bool": {
"filter": [
{"term": {"service.keyword": "api"}},
{"term": {"level.keyword": "error"}},
{"range": {"timestamp": {"gte": "now-15m"}}}
]
}
}
}' | jq '.hits.hits[]._source'
Loki example:
curl -G -s 'http://logs.example.internal:3100/loki/api/v1/query_range' \
--data-urlencode 'query={service="api",level="error"} | json | trace_id="4bf92f3577b34da6a3ce929d0e0e4736"' \
--data-urlencode 'limit=5' | jq '.data.result'
What you should see when this succeeds: only matching error events, with fields like request_id and trace_id available without regex extraction.
Verify it works
Generate one known request, then query for it end to end.
Send a request with a known trace header:
TRACE_ID=4bf92f3577b34da6a3ce929d0e0e4736
curl -s -o /dev/null -w '%{http_code}\n' https://your-service.example.com/health -H "traceparent: 00-${TRACE_ID}-0123456789abcdef-01"
Expected output:
200
Now query the backend.
OpenSearch/Elasticsearch-compatible:
curl -s http://logs.example.internal:9200/app-logs*/_search -H 'Content-Type: application/json' -d '{
"size": 3,
"query": {"term": {"trace_id.keyword": "4bf92f3577b34da6a3ce929d0e0e4736"}}
}' | jq '.hits.hits[]._source | {timestamp,service,level,trace_id,message}'
Expected output shape:
{
"timestamp": "2026-10-01T12:00:00.123Z",
"service": "api",
"level": "info",
"trace_id": "4bf92f3577b34da6a3ce929d0e0e4736",
"message": "GET /health"
}
If you use Loki, the query should return at least one stream with your service label and a JSON line containing the same trace_id.
Common pitfalls
Logging JSON as a string instead of as the log line
Mistake: the app prints INFO {"level":"error"...} or wraps JSON inside a message field.
Symptom: the collector stores one big text field like log: "{\"level\":...}", and field queries return nothing.
Fix: change the logger output to write raw JSON lines to stdout/stderr and remove text prefixes.
Inconsistent field names across services
Mistake: one service uses requestId, another uses req_id, another uses correlationId.
Symptom: incident queries miss half the fleet unless you OR together multiple fields.
Fix: standardize on one schema now: timestamp, level, service, environment, request_id, trace_id, message, err.
Collector parses the wrong timestamp format
Mistake: parser expects %Y-%m-%dT%H:%M:%S.%L%z but logs emit Z or no milliseconds.
Symptom: events arrive with ingestion time instead of event time, so timelines are out of order.
Fix: align the app timestamp format and the collector Time_Format, or keep the original field and sort on it explicitly.
High-cardinality labels in Loki
Mistake: putting request_id or user_id into labels.
Symptom: ingestion gets expensive or slow, and queries degrade badly.
Fix: keep only low-cardinality labels like service, level, environment; leave request_id and trace_id in the JSON body.
Mapping explosions in OpenSearch/Elasticsearch-compatible backends
Mistake: logging arbitrary nested objects from request bodies or user payloads.
Symptom: indexing errors, rejected documents, or huge field counts.
Fix: log a bounded subset of fields and keep untrusted payloads in one string field like payload_raw if you truly need them.
Multiline stack traces break line-delimited JSON
Mistake: the app emits pretty-printed JSON or raw multiline exceptions.
Symptom: jq fails on validation, and the collector splits one event into many broken records.
Fix: emit compact single-line JSON and serialize stack traces into one escaped string field such as err.stack.
This article was written by an AI system and published pending human review. Verify anything you intend to act on.
Have a project in mind?
Get an instant AI price estimate for it, or talk directly to our team.
One email a month on what we learn building with AI