New: The 5-Day Stoic Operator Challenge — Free. Start today →

Retrieval Augmented Generation: What RAG Is, When You Need It, and a 7-Day Build for Operators

Retrieval Augmented Generation: What RAG Is, When You Need It, and a 7-Day Build for Operators

A language model that answers from memory will guess, and retrieval augmented generation is the discipline of making it read first.

You have a body of knowledge the model has never seen. Your SOPs, your client intake notes, your offer docs, three years of call transcripts. Ask a general model about any of it and you get a confident paragraph built from averages.

Retrieval augmented generation, usually shortened to RAG, fixes that by pulling the relevant passages out of your own material and handing them to the model before it writes. The model stops answering from what it absorbed in training and starts answering from what you gave it.

This piece covers what RAG is, where the idea came from, how the pipeline works, when you do not need it at all, and a seven-day protocol for building a first version on your own documents.

What retrieval augmented generation actually is

The cleanest working definition comes from Anthropic's engineering write-up on retrieval: "RAG is a method that retrieves relevant information from a knowledge base and appends it to the user's prompt" (Anthropic, Introducing Contextual Retrieval). That is the whole mechanism. Search first, then generate with the search results in the prompt.

IBM Research frames the purpose in one line: "RAG is an AI framework for retrieving facts from an external knowledge base to ground large language models (LLMs)" (IBM Research). The same piece compares it to the difference between an open-book and a closed-book exam. A closed-book model recalls. An open-book model looks it up.

The operator's reason to care is simple. AWS lists the core weakness RAG addresses: "LLM training data is static and introduces a cut-off date on the knowledge it has" (AWS, What is RAG). AWS also names the other failure every operator has seen, the model "Presenting false information when it does not have the answer." Your pricing changed last month. Your onboarding sequence was rewritten in the spring. No model trained before those changes knows them. Retrieval is how it finds out.

Where the idea came from

The term comes from a 2020 paper, Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks, by Patrick Lewis and colleagues, accepted at NeurIPS 2020. The authors started from a specific problem: large pre-trained models store factual knowledge in their parameters but struggle to access and manipulate it precisely.

Their answer is stated in the abstract: "Pre-trained models with a differentiable access mechanism to explicit non-parametric memory can overcome this issue." In plain terms, give the model a separate, searchable memory instead of relying only on what is baked into its weights. Their system paired a generator with a dense vector index of Wikipedia, and the paper reports that it produced more specific, diverse and factual language than a parametric-only baseline.

Two terms from that paper are worth keeping. Parametric memory is what the model learned in training. Non-parametric memory is the external store you can read, edit and replace without retraining anything. Every RAG system since is a variation on that split.

How the pipeline works, step by step

A RAG system has two phases. An offline phase prepares your documents once. An online phase runs every time someone asks a question.

Offline: prepare the knowledge base

  1. Collect. Gather the documents the system is allowed to answer from. Nothing else goes in.
  2. Chunk. Split each document into passages small enough to retrieve precisely, usually a few paragraphs each.
  3. Embed. Convert each chunk into an embedding, a list of numbers that represents its meaning, so that similar passages sit close together in vector space.
  4. Store. Save the chunks and their embeddings in an index you can search, often a vector database.

Online: answer a question

  1. Embed the query. The question goes through the same embedding model.
  2. Retrieve. Find the chunks whose embeddings are nearest to the query.
  3. Assemble. Put those chunks into the prompt, with an instruction to answer only from them.
  4. Generate. The model writes the answer, ideally citing which chunk each claim came from.

One practical note for Claude users. As of October 2026, Anthropic's documentation states plainly: "Anthropic does not offer its own embedding model" (Claude Docs, Embeddings). The same page points to Voyage AI by MongoDB as one provider and lists models including voyage-4-large, voyage-4 and voyage-4-lite, with domain models such as voyage-law-2 and voyage-finance-2. The docs also tell you to set the input type to query or document for retrieval work, because the model prepends different instructions to each. Skip that and you weaken retrieval for no reason.

Where RAG breaks, and the fixes that are measured

The weak point is almost always retrieval, not generation. If the right chunk never reaches the prompt, the best model in the world answers from the wrong page.

Chunking causes most of it. A passage cut out of a long document loses its context. A chunk that says revenue grew three percent does not say which company or which quarter. Anthropic's contextual retrieval method addresses this by having a model write a short explanation of where each chunk sits in its source document, "usually 50-100 tokens", and attaching it to the chunk before embedding (Anthropic).

The reported results, from Anthropic's own tests:

  • "Contextual Embeddings reduced the top-20-chunk retrieval failure rate by 35%"
  • "Combining Contextual Embeddings and Contextual BM25 reduced the top-20-chunk retrieval failure rate by 49%"
  • "Reranked Contextual Embedding and Contextual BM25 reduced the top-20-chunk retrieval failure rate by 67%"

BM25 is a keyword ranking method. It catches exact terms such as product codes, names and error strings that embeddings can blur. A reranker is a second pass that rescores the retrieved candidates against the query before they reach the prompt. The same write-up puts the one-time cost of generating contextualized chunks at "$1.02 per million document tokens" when prompt caching is used, and reports that "Passing the top-20 chunks to the model is more effective than just the top-10 or top-5."

The principle underneath: retrieval quality compounds. Every point of retrieval accuracy shows up in every answer the system ever gives. Generation quality is borrowed from the model vendor. Retrieval quality is the part you build.

When you do not need RAG at all

The most useful sentence in the Anthropic write-up is a reason not to build anything: "If your knowledge base is smaller than 200,000 tokens (about 500 pages of material)", the post says, "you can just include the entire knowledge base in the prompt" (Anthropic).

Most solo operators sit under that line. A coaching business with a 40-page offer doc, a 60-page SOP binder and a few hundred FAQ answers does not need a vector database. It needs a well-organised document set loaded into the prompt, a reusable workspace such as a Claude Project set up for your business, and caching so you do not pay full price for the same context every time.

Use this decision rule:

  • Under roughly 500 pages, stable content: put it in the prompt. No retrieval layer.
  • Over that, or content that changes weekly: build retrieval, because you cannot fit it and you need to swap documents without rebuilding prompts.
  • Live systems such as a CRM or a calendar: this is less a RAG problem than a tool problem. Connect the model to the system directly, the way the Model Context Protocol connects Claude to your tools.

This is premeditatio malorum applied to architecture. Before you build the elaborate thing, rehearse whether the simple thing fails. Often it does not.

The seven-day RAG protocol for an operator

This builds a working first version on your own material in one week, one focused block per day. It assumes you are past the 500-page line or your content changes often enough to need it.

  1. Day 1, define the questions (45 minutes). Write the 25 questions the system must answer, taken from real client emails, DMs and team Slack. Write the correct answer and the source document beside each. This is your test set. Without it you are judging by feel.
  2. Day 2, curate the corpus (60 minutes). Collect only the documents that answer those questions. Delete drafts, duplicates and anything out of date. A stale SOP retrieved confidently is worse than no SOP. If the documents themselves are a mess, fix the source first with a standard SOP template.
  3. Day 3, chunk and embed (60 minutes). Split by heading, not by fixed character count, so each chunk is one idea. Embed with input type set to document. Store chunk text, source file and section heading together.
  4. Day 4, add context and keywords (60 minutes). Prepend a one- or two-sentence description of where each chunk sits in its document. Add a keyword index alongside the embeddings so exact names and codes are found.
  5. Day 5, write the answer prompt (45 minutes). Instruct the model to answer only from the supplied passages, to name the source of each claim, and to say the answer is not in the documents when it is not. On the Claude API, the citations feature returns the exact supporting passages, and the docs state that cited_text does not count toward output tokens (Claude Docs, Citations).
  6. Day 6, run the test set (60 minutes). Ask all 25 questions. Score each answer: correct, wrong, or correctly refused. For every wrong answer, check first whether the right chunk was retrieved. If it was not, the problem is retrieval. If it was, the problem is the prompt.
  7. Day 7, fix the largest failure and set the review (45 minutes). Fix the single biggest category of failure, rerun the test set, and put a monthly corpus review on your calendar. Documents rot. The test set tells you when.

Treat the test set the way you treat a training log. You would not judge a programme by how one session felt. Do not judge a retrieval system by one good answer.

Verification is still your job

RAG reduces invented answers. It does not remove them. A model can retrieve the right passage and still summarise it badly, or retrieve the wrong passage and summarise it faithfully. IBM's write-up names the real benefit as letting users check the model's sources, and that is the frame to keep.

The dichotomy of control applies cleanly here. You do not control what the model generates. You control what it is allowed to read, how that material is organised, and whether you check the answer against its source before acting. Put your effort there.

Build citations into every answer and spot-check them on anything that touches money, health or a client commitment. The habit is the same one covered in verifying delegated AI work before it reaches a decision. The model drafts. The operator signs.

Frequently asked questions

What is retrieval augmented generation in simple terms?

Retrieval augmented generation is a method where an AI system searches a set of documents for passages relevant to your question, adds those passages to the prompt, and then writes an answer from them. Instead of relying on what the model learned during training, it answers from material you supply, which keeps responses current and specific to your business.

Who invented RAG?

The term comes from the 2020 paper Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks by Patrick Lewis and co-authors, accepted at NeurIPS 2020. The paper combined a pre-trained text generator with a searchable dense vector index of Wikipedia and reported more specific and factual output than a model relying on its trained parameters alone.

Do I need a vector database for RAG?

Not always. Anthropic's guidance says that if your knowledge base is smaller than 200,000 tokens, about 500 pages, you can include the whole thing in the prompt and skip retrieval entirely. A vector database becomes useful when your material is larger than that or changes often enough that rebuilding a single large prompt is impractical.

Does Claude have its own embedding model?

No. As of October 2026, Anthropic's documentation states that Anthropic does not offer its own embedding model. The docs point to Voyage AI by MongoDB as one provider, list models such as voyage-4, voyage-4-large and voyage-4-lite, and advise assessing several embedding vendors to find the best fit for your use case.

Does RAG stop AI hallucinations?

It reduces them but does not eliminate them. Grounding the model in retrieved passages gives it real material to answer from, and citations let you check each claim. The model can still retrieve the wrong passage or misread the right one, so answers that affect money, health or client commitments should still be checked against the cited source.

What is the most common reason a RAG system gives wrong answers?

Poor retrieval. If the passage containing the answer never reaches the prompt, the model cannot use it. Chunks cut away from their context are a frequent cause. Anthropic reports that adding short context to each chunk, combining embeddings with keyword search, and reranking reduced top-20 retrieval failures by up to 67% in its tests.

Read before you write

Retrieval augmented generation is not a new kind of intelligence. It is an old discipline made mechanical: check the source before you speak. The operator who keeps a clean corpus, a written test set and a monthly review gets a system whose answers improve as the documents do.

The same principle runs through the body and the mind. Train from a programme, not from memory of what worked once. Decide from a principle you wrote down, not from the mood of the morning. If you want that discipline installed as a daily system, start the free 5-Day Stoic Operator Challenge.

AIai leverageapex life fitnesscompound performanceragretrieval augmented generation
TH

The Apex Desk

The editorial team behind Apex Life Fitness — operators writing about the systems where fitness, philosophy, and AI leverage intersect. Train. Think. Build.