Retrieval patterns that hold up in production
Dana Oyelaran · 18 September 2026 · 8 min read
The first retrieval pipeline you build will work beautifully on the ten documents you tested it with, and then fall apart the week it meets a real corpus. The failures are predictable, which means they are preventable.
Chunk on structure, not on length
Fixed character windows cut tables in half and separate a heading from the paragraph that explains it. Split on document structure first, then fall back to length only inside oversized sections.
Retrieve wide, rerank narrow
Pull twenty candidates from the vector store, then rerank them down to four before they reach the prompt. Embedding similarity is a cheap filter, not a judgement of relevance.
candidates = store.similarity_search(question, k=20)
context = reranker.compress_documents(candidates, question)[:4]Always carry metadata
Source, title, section and last updated date should travel with every chunk. Without them you cannot render citations, you cannot filter by recency, and you cannot explain a wrong answer after the fact.
Let the model say it does not know
An instruction to answer only from the provided context, combined with a retrieval score floor, converts a confident fabrication into an honest miss. Users forgive the second and remember the first.