Retrieval-augmented generation (RAG) is a technique where a program first looks up relevant passages from your own documents, then hands them to a language model along with the question, so the answer is written from your content instead of from the model's memory. The term comes from a 2020 paper by Lewis et al., Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. Nearly every "chat with your documents" product and AI chatbot for business is a RAG application underneath.
I wanted to see where RAG actually breaks, so I built the retrieval half from scratch in Python and pointed it at this blog: 28 posts, about 460 passages. No vector database, no framework. What follows is how RAG works, the code, what it returned, and three ways it went wrong.
How does RAG work?
Four steps, and only the last one involves the language model:
1. CHUNK split your documents into passages
2. INDEX make the passages searchable (keywords, embeddings, or both)
3. RETRIEVE for each question, fetch the few passages that match best
4. GENERATE put those passages and the question in a prompt; the model writes the answer
Most RAG quality problems are in steps 1 to 3. If the right passage never reaches the prompt, no model can write the right answer, which is why I only built the retrieval half. That is also the half you can test without calling a model at all.
A RAG retriever in about 40 lines of Python
This reads the post sources (point the POSTS environment variable at a folder of HTML files), cuts them into fixed 120-word chunks, and ranks them with BM25, the classic keyword-scoring formula. It uses only the standard library.
import re, math, glob, html, os, collections
ROOT = os.environ["POSTS"]
def strip(h):
h = re.sub(r"<(script|style|svg|pre|table)[\s\S]*?</\1>", " ", h)
h = re.sub(r"<[^>]+>", " ", h)
return re.sub(r"\s+", " ", html.unescape(h)).strip()
def title(raw):
m = re.search(r"^h2:\s*(.+)$", raw, re.M); return m.group(1) if m else "?"
chunks = []
for f in sorted(glob.glob(ROOT + "/*.html")):
raw = open(f).read()
body = raw.split("---", 2)[2]
body = re.sub(r"<!--\s*(ld|head):start\s*-->[\s\S]*?<!--\s*\1:end\s*-->", " ", body)
t = title(raw)
if t == "?":
m = re.search(r"^title:\s*(.+?)(?:\s+—\s+Blog)?$", raw, re.M); t = m.group(1) if m else "?"
words = strip(body).split()
for i in range(0, len(words), 120): # fixed 120-word chunks, no overlap
chunks.append((t, " ".join(words[i:i+120])))
tok = lambda s: re.findall(r"[a-z0-9]+", s.lower())
STOP = set("the a an of to and or in is are for on it that this with as be by at from your you i we how what do does".split())
docs = [[w for w in tok(c) if w not in STOP] for _, c in chunks]
N = len(docs); avg = sum(map(len, docs)) / N
df = collections.Counter(w for d in docs for w in set(d))
def bm25(q, d, k1=1.5, b=0.75):
tf = collections.Counter(d); s = 0
for w in q:
if w in tf:
idf = math.log(1 + (N - df[w] + .5) / (df[w] + .5))
s += idf * tf[w] * (k1 + 1) / (tf[w] + k1 * (1 - b + b * len(d) / avg))
return s
def search(q, k=3):
qt = [w for w in tok(q) if w not in STOP]
r = sorted(((bm25(qt, d), i) for i, d in enumerate(docs)), reverse=True)[:k]
return [(round(s, 2), chunks[i]) for s, i in r]
The generation step would be one more function: join the top chunks into a prompt such as "Answer using only the passages below; if they do not contain the answer, say so", then send it to a model. I did not run that part, so nothing below is about answer quality, only about retrieval.
What did it return?
I ran four questions against the cleaned index. The top hit for each:
| Question | Top result | BM25 score |
|---|---|---|
| what is an agent harness | What Is an AI Agent Harness? | 9.1 |
| how do I test a voice agent | The Essential Metrics for Voice AI Agents | 8.96 |
| why does my dead letter queue fill up | Using PostgreSQL as a Dead Letter Queue | 23.18 |
| which model is best for photography | What Is an AI Agent Harness? (wrong) | 7.03 |
Three of four are right. The fourth is the interesting one, and it was not the only problem.
Three ways my RAG failed
1. Markup leaked into the passages. My first run indexed the raw post files. Every post carries a JSON-LD metadata block in its header, and stripping those blocks removed 14 of the 474 chunks. The top hit for "what is an agent harness" was a chunk beginning { "keywords": [ "AI agent harness"..., which would have been pasted into a prompt as if it were prose. Stripping those blocks took the index to 460 chunks and put real sentences at the top. Cleaning your documents is unglamorous work, and it decides what the model gets to see.
2. Keywords cannot tell two meanings of "model". I have a post about learning photography with a camera, and it did not appear in the top three for "which model is best for photography". The query matched my AI posts on the word model. This is the case embeddings are for: they place text by meaning, so "camera" and "photography" sit near each other and an AI "model" does not. Production systems usually combine keyword and embedding search (hybrid search) rather than choosing one.
3. The score does not tell you when to refuse. I also asked "what is the capital of france", which my blog does not answer. The best chunk scored 5.47. The photography miss scored 7.03, and a correct answer on voice agents scored 8.96. BM25 scores are relative to the corpus, so there is no number that separates "found it" from "nothing here". A real application needs a separate check, such as a reranker or a model that judges whether the passage actually answers the question, and a prompt that is allowed to say "I don't know".
When is RAG the right tool?
RAG fits when the answers live in documents that change, such as support articles, policies, product docs or contracts, and when you need to point to a source. It is a poor fit for teaching a model a new skill or style, which is closer to fine-tuning, and unnecessary when everything fits comfortably in the prompt. Often the simplest design is to put the whole document in the context window and skip retrieval.
What a production RAG system adds
- Hybrid retrieval: keywords + embeddings, merged
- A reranker that re-scores the top 20 and keeps the best 3-5
- Chunking that respects headings and tables, not fixed word counts
- Citations: every claim links back to its passage
- A refusal path for questions the documents do not cover
- Evals: a fixed question set you re-run after every change
The last item is the one teams skip. A RAG system that is not measured degrades silently when documents change or someone edits a prompt. I wrote about how to set that up in A Dummy's Guide to Evals and the field guide to testing and observability.
Want one built for your documents?
Retrieval over your own data, with the evals to keep it honest, is one of the projects on my AI consulting page. Book a 30-minute call and bring a sample of the documents and five questions you would expect it to answer.
Method note: tested on 8 October 2026 with Python 3.10 and no third-party packages, against the 28 post sources in this site's repository. The code and results above come from that run. I used AI assistance to help write and edit this post and checked every number against the run output.
Comments
Have thoughts on this post? Join the discussion below! Comments are powered by Disqus.