Build ยท 9 min read
Retrieval augmented generation
Ground model answers in your own documents with a vector store.
Retrieval augmented generation gives the model the specific context it needs at answer time, instead of relying on what it memorised during training.
Index your documents
python
from langchain_text_splitters import RecursiveCharacterTextSplitter
from langchain_openai import OpenAIEmbeddings
from langchain_core.vectorstores import InMemoryVectorStore
splitter = RecursiveCharacterTextSplitter(chunk_size=1000, chunk_overlap=200)
chunks = splitter.split_documents(raw_documents)
store = InMemoryVectorStore.from_documents(chunks, OpenAIEmbeddings())
retriever = store.as_retriever(search_kwargs={"k": 4})Answer with context
python
from langchain_core.runnables import RunnablePassthrough
template = ChatPromptTemplate.from_template(
"Answer using only the context.\n\nContext:\n{context}\n\nQuestion: {question}"
)
rag = (
{"context": retriever, "question": RunnablePassthrough()}
| template
| model
| StrOutputParser()
)Note
Chunk size is the single biggest lever on answer quality. Start near 1000 characters with generous overlap, then tune against a real evaluation set.
Common pitfalls
- Retrieving too many chunks crowds out the question itself.
- Missing metadata makes citations impossible to render.
- Re-embedding on every deploy is slow and expensive; persist your index.