Build ยท 9 min read

Retrieval augmented generation

Ground model answers in your own documents with a vector store.


Retrieval augmented generation gives the model the specific context it needs at answer time, instead of relying on what it memorised during training.

Index your documents

python
from langchain_text_splitters import RecursiveCharacterTextSplitter
from langchain_openai import OpenAIEmbeddings
from langchain_core.vectorstores import InMemoryVectorStore

splitter = RecursiveCharacterTextSplitter(chunk_size=1000, chunk_overlap=200)
chunks = splitter.split_documents(raw_documents)

store = InMemoryVectorStore.from_documents(chunks, OpenAIEmbeddings())
retriever = store.as_retriever(search_kwargs={"k": 4})

Answer with context

python
from langchain_core.runnables import RunnablePassthrough

template = ChatPromptTemplate.from_template(
    "Answer using only the context.\n\nContext:\n{context}\n\nQuestion: {question}"
)

rag = (
    {"context": retriever, "question": RunnablePassthrough()}
    | template
    | model
    | StrOutputParser()
)

Note

Chunk size is the single biggest lever on answer quality. Start near 1000 characters with generous overlap, then tune against a real evaluation set.

Common pitfalls

  • Retrieving too many chunks crowds out the question itself.
  • Missing metadata makes citations impossible to render.
  • Re-embedding on every deploy is slow and expensive; persist your index.