retrieval · Level 3

RAG and knowledge bases, step by step

How to let a model answer from your own documents: chunking, embeddings, retrieval, citations, evaluation and security.

By Mickarle Wagstaff-Irons - Micky Irons

  • Level 3Building
  • 100 min
  • 6 chapters
  • Free PDF, no account
The Indexretrieval / 03

Start with the essentials

The short answer

Retrieval augmented generation, or RAG, lets a model answer from your own documents. You split the documents into chunks, turn them into embeddings, retrieve the passages most relevant to each question, put them in the prompt and ask the model to answer from them with citations. It brings freshness and private data, but only if retrieval is good.

What you will learn

  • You will be able to explain what RAG is, and when it is a better choice than fine-tuning or a long context.
  • You will be able to prepare a knowledge base: collecting, cleaning and chunking documents with sensible size and overlap.
  • You will understand embeddings, keyword and semantic search, hybrid search, top-k and reranking.
  • You will be able to assemble a prompt that answers from sources with citations, and check the citations.
  • You will be able to measure retrieval with hit rate and answers with a faithfulness check, and diagnose common failures.
  • You will know the main security risks: per-document access control and prompt injection inside documents.

Who it is for

Builders who want a chatbot or assistant that answers from their own documents. The exercises use pen and paper and any free chatbot. You need no programming to follow them, though the code example is in Python.

Before you start

  • Comfort with prompts, tokens and context windows. Having built a simple chatbot (see Build your first chatbot) helps.

Keep learning

The complete workbook

This workbook walks through retrieval augmented generation one stage at a time. You will collect and clean documents, choose chunk sizes, understand embeddings and hybrid search, assemble a prompt that cites its sources, and measure both retrieval and answers. You will build a tiny knowledge base by hand from five short documents, and learn the security and design trade-offs.

  1. 01
    What RAG is, and why you would use it

    A model knows only what it learned in training, and it learned nothing about your documents. Retrieval augmented generation fixes that without retraining anything.

    In the workbook · Reading
  2. 02
    Preparing the knowledge: collect, clean, chunk

    The quality of a RAG system is largely set before any model is called. Careful preparation usually beats clever prompting.

    In the workbook · 1 exercise
  3. 03
    Embed, store and retrieve

    With chunks ready, you need a way to find the right ones quickly for any question.

    In the workbook · 1 exercise
  4. 04
    Assemble the prompt and generate with citations

    Retrieval finds the evidence. The prompt decides whether the model uses it honestly.

    In the workbook · 1 exercise
  5. 05
    Evaluate: does it retrieve, and does it tell the truth?

    There are two separate questions to answer. Measure them separately, or you will not know what to fix.

    In the workbook · 1 exercise
  6. 06
    Security, and choosing between RAG, fine-tuning and long context

    A knowledge base can leak what it should not, and it can be turned against you. Then you must decide whether RAG is even the right tool.

    In the workbook · Reading

Also inside: a 9-point checklist, a glossary of 12 terms and 10 questions and answers to test yourself. 4 hands-on exercises, each with a worked answer at the back where the workbook gives one.

No login, no card, no account. Before the download we ask you to follow Mickai (two quick links). Free to download and use for personal learning, study groups and inside your own team. Please do not resell the workbooks or republish them as your own. Link people to trust-agent.ai instead.

Test yourself

Questions and answers

What is RAG?

Retrieval augmented generation. The system finds passages in your own documents that are relevant to a question, puts them in the prompt and asks the model to answer from them, ideally with citations. It gives a model access to fresh or private information without retraining it.

Does RAG stop hallucinations?

No, though it can reduce them. If retrieval misses the right passage, or the model ignores the sources, it can still produce a confident wrong answer. Good retrieval, firm rules to answer only from sources, citation checks and evaluation all matter.

How big should chunks be?

There is no single right size. Small chunks match precisely but lose context. Large chunks keep context but blur topics and use prompt space. Start with a few paragraphs, split at headings, add a small overlap, and test on your own questions.

What is an embedding?

It is a list of numbers, produced by an embedding model, that represents the meaning of a text. Texts with similar meanings get similar lists, so a search can find relevant chunks even when the wording differs from the question.

Why use hybrid search?

Keyword search is strong on exact names and codes but misses different wording. Semantic search handles meaning but can miss exact identifiers. Running both and merging the results covers each method's weaknesses, so it is a sensible default.

How do I measure whether retrieval works?

Write real questions with the chunk that answers each, then measure hit rate: the share of questions where a correct chunk appears in the top k. Include questions the documents cannot answer. Do this before changing the model or the prompt.

What is faithfulness?

It measures whether every claim in an answer is supported by the retrieved sources. Check a sample by hand, and confirm that citations point to passages that really say what is claimed. A second model can help judge, but spot-check it.

Should I use RAG or fine-tuning?

Use RAG when facts change, answers need citations or the data is private. Use fine-tuning to shape style, format or a narrow skill. They can be combined. Fine-tuning is a poor way to keep up with changing facts.

How do I stop users seeing documents they should not?

Attach an access label to every chunk and filter by the current user's rights before ranking, so restricted text never reaches the prompt. Do not depend on instructing the model to withhold it. Test with users who have different rights.

Can a document attack my RAG system?

Yes. A document can contain hostile instructions that the model may follow, called indirect prompt injection. Treat retrieved text as untrusted data, give the bot no powerful tools, review what enters the collection, and log which chunks each answer used.

When you have finished

Get your certificate of completion

Type your name and download a certificate for this workbook as a PDF, ready to print or to add to LinkedIn. It is made on your own device, so your name is never sent to us. It is a self-declared certificate, not an accredited qualification.

Learn the language

Key terms

RAG
Retrieval augmented generation: finding relevant passages in your documents and giving them to a model to answer from.
Chunk
A piece of a document, sized to be retrieved and placed in a prompt on its own.
Overlap
Text repeated at the boundary between two chunks so that ideas straddling the cut are not lost.
Metadata
Information about a document or chunk, such as title, source, date, owner and who may see it.
Embedding
A list of numbers representing the meaning of a piece of text, so that similar meanings can be found by similarity.
Semantic search
Searching by meaning using embeddings, so that different wording can still match.

6 of the workbook's 12 terms. The complete glossary is in the workbook.

Follow the evidence

Sources and checks

Facts last checked: .

These workbooks use AI assistance. See how the workbooks are made.

  1. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (Lewis et al., 2020)arXiv
  2. OWASP Top 10 for LLM Applications (includes prompt injection)OWASP Gen AI Security Project
  3. Prompt injection is not SQL injection (it may be worse)National Cyber Security Centre (NCSC)

Created by Mickarle Wagstaff-Irons - Micky Irons with the Mickai team. Published by Mickai LTD. Last updated 25 September 2026.

NextKeep going

Where to go next