retrieval · Level 3

Embeddings and vector search from first principles

What an embedding is, how closeness is measured and how nearest-neighbour search works, built and measured in Python with vectors you compute yourself.

By Mickarle Wagstaff-Irons - Micky Irons

  • Level 3Building
  • 95 min
  • 6 chapters
  • Free PDF, no account
The Indexretrieval / 03

Start with the essentials

The short answer

An embedding is a list of numbers standing for a piece of text, made so that texts with similar meanings sit close together. Vector search turns the question into the same kind of list, scores it against the stored vectors with a measure such as cosine similarity, and returns the nearest. It fails in predictable ways, so measure it with labelled questions using recall@k and mean reciprocal rank.

What you will learn

  • You will be able to explain what a vector and an embedding are, and build count and TF-IDF vectors from text.
  • You will be able to compute a dot product, cosine similarity and Euclidean distance by hand, and explain when they disagree and why normalising makes them agree.
  • You will have built and run an exact, brute-force nearest-neighbour search in Python.
  • You will be able to score a search with recall@k and mean reciprocal rank on a labelled question set.
  • You will be able to recognise the common failures and say which a neural embedding model would change.
  • You will understand the recall and speed trade-off of approximate search, and why vectors of personal text need protecting.

Who it is for

Builders who have met retrieval augmented generation and want to understand what an embedding is and how vector search works, by building one. You need to run a Python script from a terminal. The maths is multiplication, square roots and fractions.

Before you start

  • The RAG and knowledge bases workbook, or equivalent familiarity with chunks, embeddings and top-k retrieval. Enough Python to run a script from a terminal: the Python: from first script to a useful automation workbook covers this. Chapter 3 gives the two lines that install the packages into a virtual environment.

Keep learning

The complete workbook

This workbook builds vector search from first principles, using vectors you compute yourself rather than a downloaded model. You will compare vectors by hand, run an exact search over sixteen invented documents, score it with recall@k and mean reciprocal rank, and reproduce its typical failures. It ends with high dimensions, approximate indexes and privacy.

  1. 01
    Text as numbers

    Search compares lists of numbers, not sentences, so every vector search starts by turning text into one.

    In the workbook · 1 exercise
  2. 02
    Closeness: dot product, cosine and distance

    Search finds the stored vectors closest to the question's vector. Three measures are common, and they can disagree.

    In the workbook · 1 exercise
  3. 03
    Build a brute-force search

    Now search sixteen invented library help notes: one TF-IDF vector per note, every note scored, the best k returned.

    In the workbook · 1 exercise
  4. 04
    Measure it: recall@k and mean reciprocal rank

    Two questions prove nothing. Score the search on labelled questions, and keep that set fixed while you change things.

    In the workbook · 1 exercise
  5. 05
    Why it fails, and what a neural embedding changes

    Vector search fails in a few predictable ways. Reproduce them, then ask which a neural embedding model would change.

    In the workbook · 1 exercise
  6. 06
    High dimensions, faster search and privacy

    Brute force is exact but grows with the collection. See how high dimensions behave, what approximate indexes trade away, and why vectors need protecting.

    In the workbook · 1 exercise

Also inside: a 10-point checklist, a glossary of 12 terms and 10 questions and answers to test yourself. 6 hands-on exercises, each with a worked answer at the back where the workbook gives one.

No login, no card, no account. Before the download we ask you to follow Mickai (two quick links). Free to download and use for personal learning, study groups and inside your own team. Please do not resell the workbooks or republish them as your own. Link people to trust-agent.ai instead.

Test yourself

Questions and answers

What is an embedding, in plain words?

A list of numbers that stands for a piece of text. An embedding model is trained so that texts with similar meanings get similar lists, so you can find related text by comparing numbers instead of words. The list has a fixed length, and each position is a dimension.

Are TF-IDF vectors embeddings?

They are vectors of the same kind and are searched the same way, but they are not neural embeddings. Word-level TF-IDF only records which words a text contains, weighted by rarity, so 'bike' and 'bicycle' never match. A neural embedding model can place synonyms and paraphrases close together. That is why this workbook uses TF-IDF to teach the mechanics, not to show meaning.

Should I use cosine similarity, dot product or Euclidean distance?

For unit-length vectors it does not matter: all three rank in the same order. For other vectors they can disagree, because the dot product rewards length and distance punishes it. Cosine similarity ignores length. Use the measure your embedding tool documents, and normalise stored vectors and queries the same way.

What does normalising a vector do?

It divides the vector by its length, so the new length is 1 and only the direction remains. Afterwards the dot product equals the cosine, and squared distance equals 2 minus twice the cosine. scikit-learn's TfidfVectorizer normalises its rows by default.

Is brute-force search good enough?

Often. It is exact, simple and easy to test, and its cost grows in step with the number of vectors. Move to an approximate index only when brute force is measurably too slow, then check the index's recall@k against brute-force results on your own data.

What is recall@k, and how does it differ from hit rate?

Recall@k is the share of a question's gold documents that appear in the top k, averaged over questions. When each question has exactly one gold document, it equals the hit rate from the RAG workbook. When a question has several, finding some of them earns partial credit.

What is mean reciprocal rank?

For each question, take 1 divided by the rank of the first correct result, or 0 if none is found, then average across questions. A correct first result scores 1, a second-place one 0.5. It rewards putting the right answer at the top, which suits retrieval that feeds a model's prompt.

Why does vector search return results for questions it cannot answer?

Nearest-neighbour search always returns the k nearest vectors, even when none is relevant. In this workbook, a question about train tickets scored 0.47 against a note about lost books. Scores are not probabilities, so choose any cut-off from labelled data that includes unanswerable questions, and make the answer step able to say it does not know.

Do I need to rebuild the index when documents change?

Yes. Stored vectors describe the text as it was when it was embedded. With TF-IDF, one new document changes every word's weight, so refit and rebuild. With a neural model, re-embed the changed chunks, and re-embed everything if you change the model. Store a text hash and model version with each vector so you can detect stale entries.

Can someone recover my text from its embeddings?

Partly, and sometimes almost entirely. A TF-IDF vector gives its words straight back. Published research recovered 50–70% of input words from sentence embeddings, and 92% of 32-token texts exactly from others, plus full names from clinical notes in a separate test. If the text is personal data, protect its vectors in the same way.

When you have finished

Get your certificate of completion

Type your name and download a certificate for this workbook as a PDF, ready to print or to add to LinkedIn. It is made on your own device, so your name is never sent to us. It is a self-declared certificate, not an accredited qualification.

Learn the language

Key terms

Vector
An ordered list of numbers. Each position is a dimension.
Embedding
A vector made by a trained neural network so that texts with similar meanings get nearby vectors.
Bag of words
A vector with one dimension per vocabulary word, counting how often each appears. Word order is lost.
TF-IDF
Term frequency times inverse document frequency: a word weighting that makes words common across the collection count for less.
Dot product
Multiply matching positions of two vectors and add the results. It grows with both direction and length.
Cosine similarity
The dot product divided by both lengths: the cosine of the angle between two vectors, ignoring length.

6 of the workbook's 12 terms. The complete glossary is in the workbook.

Follow the evidence

Sources and checks

Facts last checked: .

Examples in this workbook were run on: Python 3.12.10 on Windows 11, in a fresh virtual environment with numpy 2.5.3 and scikit-learn 1.9.1 (which installed scipy 1.18.1, joblib 1.6.0, threadpoolctl 3.7.0, narwhals 2.26.0 and cloudpickle 3.1.2) using pip 25.0.1 (2026-09-26).

These workbooks use AI assistance. See how the workbooks are made.

  1. Introduction to Information Retrieval (Manning, Raghavan and Schutze): Term frequency and weighting (the bag of words model)Cambridge University Press (online edition, Stanford NLP Group)
  2. Introduction to Information Retrieval: Dot productsCambridge University Press (online edition, Stanford NLP Group)
  3. Introduction to Information Retrieval: Queries as vectorsCambridge University Press (online edition, Stanford NLP Group)
  4. Introduction to Information Retrieval: Inverse document frequencyCambridge University Press (online edition, Stanford NLP Group)
  5. Introduction to Information Retrieval: Evaluation of unranked retrieval sets (precision and recall)Cambridge University Press (online edition, Stanford NLP Group)
  6. Introduction to Information Retrieval: Evaluation of ranked retrieval resultsCambridge University Press (online edition, Stanford NLP Group)
  7. Introduction to Information Retrieval: Cluster pruningCambridge University Press (online edition, Stanford NLP Group)
  8. The TREC-8 Question Answering Track Report (Voorhees)National Institute of Standards and Technology (NIST)
  9. sklearn.feature_extraction.text.TfidfVectorizerscikit-learn 1.9.1 documentation
  10. Feature extraction: text feature extraction, tf-idf and the limits of bag of wordsscikit-learn 1.9.1 documentation
  11. numpy.linalg.normNumPy v2.5 Manual
  12. numpy.argsort (including the stable sort option)NumPy v2.5 Manual
  13. numpy 2.5.3 (Requires: Python >=3.12)Python Package Index (PyPI)
  14. Dense Passage Retrieval for Open-Domain Question Answering (Karpukhin et al., 2020)arXiv, 2004.04906
  15. When Is "Nearest Neighbor" Meaningful? (Beyer, Goldstein, Ramakrishnan and Shaft, 1999)Springer, Database Theory: ICDT'99
  16. Efficient and robust approximate nearest neighbor search using Hierarchical Navigable Small World graphs (Malkov and Yashunin)arXiv, 1603.09320
  17. Information Leakage in Embedding Models (Song and Raghunathan, 2020)arXiv, 2004.00053
  18. Text Embeddings Reveal (Almost) As Much As Text (Morris et al., 2023)arXiv, 2310.06816
  19. What is personal data?Information Commissioner's Office (ICO)

Created by Mickarle Wagstaff-Irons - Micky Irons with the Mickai team. Published by Mickai LTD. Last updated 26 September 2026.

NextKeep going

Where to go next