Retrieval Results Changed After Rebuilding the Index: How to Diagnose It
For developers running search or RAG pipelines, this runbook covers why retrieval output changes after an index rebuild and how to isolate the exact cause quickly. You’ll get concrete checks for chunking drift, embedding-model mismatch, ANN settings, document-order instability, and stale/partial data, plus remediation steps you can automate in CI.
TL;DR — If retrieval changed after an index rebuild, the most common causes are not "randomness" but a changed corpus, changed chunking, changed embedding model/version, or different ANN build/query parameters. Start by diffing document counts and chunk counts, then verify the embedding model identifier and index settings used before vs. after the rebuild. Reading time: ~6 min
The scenario
It’s Tuesday afternoon. You rebuild your retrieval index after a routine content sync or deploy, and suddenly the same query that used to return the right design doc now returns a changelog, a duplicate chunk, or nothing relevant in the top 5. The app is up, latency looks normal, and there are no obvious exceptions in the logs. But your eval set dropped, support is pasting screenshots into Slack, and you need to answer the uncomfortable question: what changed if the source docs supposedly did not?
Symptoms
- Same query string returns different top-k results before vs. after rebuild.
- Relevance metrics drop after rebuild:
Recall@10: 0.81 -> 0.62
MRR@10: 0.74 -> 0.49
nDCG@10: 0.78 -> 0.53
- Duplicate or near-duplicate chunks appear in top results.
- Expected documents disappear from results even though they still exist in storage.
- Result ordering changes between repeated runs with the same query.
- Logs show corpus/count drift:
INFO indexer: loaded documents=12483 chunks=98114
INFO indexer: previous snapshot documents=12483 chunks=75602
- Logs show model drift or fallback:
WARN embedder: requested model=text-embedding-3-large, using default model due to missing config
INFO embedder: model=all-MiniLM-L6-v2 dim=384
- Vector DB/index build logs show different metric or build params:
INFO ann: index_type=hnsw metric=cosine M=16 ef_construction=200
INFO ann: query ef_search=20
- Partial rebuild or failed ingest leaves missing segments:
ERROR worker-7: failed to fetch s3://docs-bucket/manuals/ops.pdf: 403 Forbidden
INFO indexer: completed with errors=17 skipped=241
Likely causes
| Cause | How common | Quick check |
|---|---|---|
| Corpus changed during rebuild (missing, extra, filtered, or newer docs) | Very common | ```bash |
| diff <(jq -r '.id' prebuild_manifest.json | sort) <(jq -r '.id' postbuild_manifest.json | sort) |
| Chunking/splitting changed (size, overlap, parser, normalization) | Very common | ```bash
jq '{chunk_size,chunk_overlap,parser,normalization}' index-build-metadata.json
``` |
| Embedding model/version changed, or fallback model was used | Very common | ```bash
jq '{embedding_model,embedding_dim,tokenizer}' index-build-metadata.json
``` |
| ANN index/query parameters changed (metric, HNSW/IVF params, exact vs approximate) | Common | ```bash
jq '{index_type,metric,build_params,query_params}' index-build-metadata.json
``` |
| Document/chunk IDs are unstable, causing duplicate rows or bad upserts | Common | ```bash
jq -r '.chunks[] | [.doc_id,.chunk_id,.content_hash] | @tsv' chunks-manifest.json | head
``` |
| Partial rebuild or mixed old/new index contents | Common | ```bash
grep -E 'ERROR|WARN|skipped|fallback' build.log | tail -n 50
``` |
| Tie-breaking/order instability on equal scores | Occasional | ```bash
curl -s http://localhost:8080/search?q='your query' | jq '.results[:10][] | {id,score}'
``` |
| Language-specific tokenization/OCR/parser output changed | Occasional | ```bash
jq -r '.parser_versions,.ocr_languages' index-build-metadata.json
``` |
## Step-by-step diagnosis
1. Compare corpus membership before vs. after rebuild.
```bash
diff <(jq -r '.documents[]?.id // .id' prebuild_manifest.json | sort) <(jq -r '.documents[]?.id // .id' postbuild_manifest.json | sort) | sed -n '1,80p'
This is your problem if you see added/removed IDs you did not expect, or counts differ.
< docs/runbooks/nginx-timeouts.md
> docs/runbooks/nginx-timeouts-v2.md
Jump to: ### Corpus changed during rebuild.
-
Compare document and chunk counts.
jq '{documents, chunks}' prebuild_stats.json postbuild_stats.jsonThis is your problem if document count is stable but chunk count moved materially without an intentional splitter change.
{ "documents": 12483, "chunks": 75602 } { "documents": 12483, "chunks": 98114 }Jump to:
### Chunking or parsing changed. -
Verify embedding model identity and vector dimension.
jq '{embedding_model, embedding_dim, tokenizer}' prebuild_metadata.json postbuild_metadata.jsonThis is your problem if model name, revision, dimension, or tokenizer differs, or if one build shows a fallback/default.
{ "embedding_model": "text-embedding-3-large", "embedding_dim": 3072, "tokenizer": "cl100k_base" } { "embedding_model": "all-MiniLM-L6-v2", "embedding_dim": 384, "tokenizer": "bert-base-uncased" }Jump to:
### Embedding model or version changed. -
Check ANN metric and build/query params.
jq '{index_type, metric, build_params, query_params}' prebuild_metadata.json postbuild_metadata.jsonThis is your problem if metric changed (
cosinevsdotvsl2), or HNSW/IVF params changed enough to affect recall. Jump to:### ANN index or query parameters changed. -
Check for unstable IDs and bad upserts.
jq -r '.chunks[] | [.doc_id,.chunk_id,.content_hash] | @tsv' chunks-manifest.json | sort | uniq -d | sed -n '1,20p'This is your problem if duplicate chunk IDs exist for different content, or if IDs are derived from array position/order that changed. Jump to:
### IDs are unstable or upserts are incorrect. -
Inspect build logs for partial failure or mixed contents.
grep -E 'ERROR|WARN|skipped|fallback|forbidden|timeout' build.log | tail -n 100This is your problem if the rebuild completed with skipped documents, auth failures, timeouts, or fallback parser/model messages. Jump to:
### Partial rebuild or mixed index contents. -
Test whether score ties are being returned in arbitrary order.
for i in 1 2 3; do curl -s 'http://localhost:8080/search?q=incident+response&k=10' | jq '.results[:5][] | [.id,.score]'; doneThis is your problem if scores are identical or nearly identical and ordering changes run-to-run. Jump to:
### Tie-breaking and ordering instability. -
Compare parser/OCR versions if PDFs, HTML cleanup, or multilingual docs are involved.
jq '{parser_versions, ocr_languages, html_sanitizer}' prebuild_metadata.json postbuild_metadata.jsonThis is your problem if extracted text changed because parser versions or OCR language packs changed. Jump to:
### Parser or OCR output changed.
Fixes
Corpus changed during rebuild
Rebuild from a pinned manifest instead of a live source that can change during indexing.
find docs -type f -print0 | sort -z | xargs -0 sha256sum > corpus-manifest.sha256
Use that manifest as the source of truth for the rebuild job, and fail if the live source differs:
sha256sum -c corpus-manifest.sha256
If your indexer applies filters, pin them in config instead of environment defaults:
{
"include_glob": ["docs/**/*.md", "manuals/**/*.pdf"],
"exclude_glob": ["archive/**", "drafts/**"],
"min_bytes": 200
}
Verify it worked:
diff <(jq -r '.documents[].id' prebuild_manifest.json | sort) <(jq -r '.documents[].id' postbuild_manifest.json | sort)
Chunking or parsing changed
Pin chunking and normalization settings in versioned config, not CLI defaults.
{
"chunk_size": 800,
"chunk_overlap": 120,
"split_on": ["heading", "paragraph"],
"normalization": {
"unicode_nf": "NFKC",
"collapse_whitespace": true,
"strip_boilerplate": true
},
"parser": "markdown+pdf-v2"
}
Rebuild with explicit flags:
./indexer build --config index-config.json --metadata-out index-build-metadata.json
Trade-off: smaller chunks often improve exact hit rate but can hurt context continuity; larger chunks can bury the relevant sentence. Keep evals for both retrieval and answer quality. Verify it worked:
jq '{chunk_size,chunk_overlap,parser,normalization}' index-build-metadata.json
Embedding model or version changed
Pin the exact model identifier and reject fallback behavior.
export EMBEDDING_MODEL='text-embedding-3-large'
export EMBEDDING_DIM='3072'
./indexer build --embedding-model "$EMBEDDING_MODEL" --fail-on-model-mismatch
If your vector store schema includes dimension, recreate the collection/index when dimension changes.
⚠️ Recreating the collection deletes the existing vectors unless you build into a new collection name and swap aliases afterward.
./indexer build --collection docs_v2026_10_01
./indexer alias swap --from docs_current --to docs_v2026_10_01
Verify it worked:
jq '{embedding_model,embedding_dim}' index-build-metadata.json
ANN index or query parameters changed
Pin metric and ANN params in config and keep them identical between build and query code paths.
{
"index_type": "hnsw",
"metric": "cosine",
"build_params": {"M": 16, "ef_construction": 200},
"query_params": {"ef_search": 80}
}
If recall dropped after a rebuild, raise query-time breadth first before rebuilding:
curl -s 'http://localhost:8080/search?q=runbook&k=10&ef_search=120' | jq '.results[:5]'
Trade-off: higher ef_search improves recall but increases latency. Metric mismatches are worse: dot vs cosine can completely reorder results if vectors are not normalized consistently.
Verify it worked:
jq '{index_type,metric,build_params,query_params}' index-build-metadata.json
IDs are unstable or upserts are incorrect
Generate IDs from stable content, not array position or ingestion order.
python3 - <<'PY'
import hashlib, json, sys
for line in sys.stdin:
obj=json.loads(line)
stable=f"{obj['doc_path']}|{obj['section_heading']}|{obj['content'].strip()}".encode()
obj['chunk_id']=hashlib.sha256(stable).hexdigest()[:24]
print(json.dumps(obj, ensure_ascii=False))
PY
Then upsert by chunk_id and delete tombstoned IDs from the previous manifest.
comm -23 <(sort old_chunk_ids.txt) <(sort new_chunk_ids.txt) > delete_ids.txt
./vectorctl delete --ids-file delete_ids.txt
Verify it worked:
jq -r '.chunks[].chunk_id' chunks-manifest.json | sort | uniq -d | wc -l
Partial rebuild or mixed index contents
Do not rebuild in place if the process can fail halfway. Build into a fresh collection/index, then atomically switch readers.
⚠️ In-place deletes plus failed ingest can leave you with a permanently incomplete index until the next full rebuild.
./indexer build --collection docs_v_next --log build.log
./healthcheck retrieval --collection docs_v_next --queries smoke-queries.txt
./indexer alias swap --from docs_current --to docs_v_next
Fix the underlying fetch/auth errors first:
grep -E '403|401|timeout|forbidden' build.log | sed -n '1,40p'
Verify it worked:
grep -E 'ERROR|skipped' build.log | tail -n 20
Tie-breaking and ordering instability
Add a deterministic secondary sort after score, usually by stable document ID or last-updated timestamp if your product expects freshness.
{
"sort": [
{"field": "score", "order": "desc"},
{"field": "doc_id", "order": "asc"}
]
}
If your scores are quantized or rounded, return more decimals internally before sorting. Verify it worked:
for i in 1 2 3; do curl -s 'http://localhost:8080/search?q=incident+response&k=5' | jq '.results[:5][] | [.id,.score]'; done
Parser or OCR output changed
Pin parser/container versions and OCR language packs in the build image.
docker inspect your-indexer-image:prod --format '{{.Id}}'
Example pinned build image snippet:
FROM python:3.12-slim
RUN pip install "pymupdf==1.25.3" "beautifulsoup4==4.12.3" "lxml==5.3.0"
RUN apt-get update && apt-get install -y tesseract-ocr tesseract-ocr-eng tesseract-ocr-deu
If extracted text changed, re-run extraction on a sample file and diff the output.
./extract-text docs/manuals/ops.pdf > new.txt
git show last-known-good:fixtures/ops.txt > old.txt
diff -u old.txt new.txt | sed -n '1,120p'
Verify it worked:
jq '{parser_versions,ocr_languages}' index-build-metadata.json
Prevention
- Store build metadata next to every index artifact and fail deploys on drift.
jq -n '{git_sha: env.GIT_SHA, embedding_model: env.EMBEDDING_MODEL, chunk_size: 800, chunk_overlap: 120, metric: "cosine"}' > index-build-metadata.json
- Add a CI check that compares retrieval on a fixed eval set before swapping aliases.
./eval-retrieval --queries eval/queries.jsonl --collection docs_v_next --min-recall-at-10 0.78 --min-mrr-at-10 0.70
- Pin the corpus with a manifest and fail the build if source files changed mid-run.
find docs -type f -print0 | sort -z | xargs -0 sha256sum > corpus-manifest.sha256 && sha256sum -c corpus-manifest.sha256
- Build into a new collection and use alias swap instead of in-place mutation.
./indexer build --collection docs_$(date +%Y%m%d%H%M%S) && ./indexer alias swap --from docs_current --to docs_$(date +%Y%m%d%H%M%S)
- Emit and alert on ingest completeness and duplicate-ID counts.
./index-stats --collection docs_current | jq '{documents,chunks,duplicate_chunk_ids,failed_fetches}'
- Keep parser/indexer in a pinned container image, and record the image digest in metadata.
docker image inspect your-indexer:prod --format '{{index .RepoDigests 0}}'
This article was written by an AI system and published pending human review. Verify anything you intend to act on.
Have a project in mind?
Get an instant AI price estimate for it, or talk directly to our team.
One email a month on what we learn building with AI