← All postsRetrieval

Retrieval patterns that hold up in production

Dana Oyelaran · 18 September 2026 · 8 min read


The first retrieval pipeline you build will work beautifully on the ten documents you tested it with, and then fall apart the week it meets a real corpus. The failures are predictable, which means they are preventable.

Chunk on structure, not on length

Fixed character windows cut tables in half and separate a heading from the paragraph that explains it. Split on document structure first, then fall back to length only inside oversized sections.

Retrieve wide, rerank narrow

Pull twenty candidates from the vector store, then rerank them down to four before they reach the prompt. Embedding similarity is a cheap filter, not a judgement of relevance.

python
candidates = store.similarity_search(question, k=20)
context = reranker.compress_documents(candidates, question)[:4]

Always carry metadata

Source, title, section and last updated date should travel with every chunk. Without them you cannot render citations, you cannot filter by recency, and you cannot explain a wrong answer after the fact.

Let the model say it does not know

An instruction to answer only from the provided context, combined with a retrieval score floor, converts a confident fabrication into an honest miss. Users forgive the second and remember the first.