Skip to content

Module 2 of Building a RAG system. About 90 minutes.

Documents, chunking and embeddings

Get text out of the clinic's documents, including a real journal PDF, cut it into chunks, and turn each chunk into numbers that capture its meaning. Then search by meaning, and test whether "parvo" finds canine parvovirus and whether "HBC" finds hit by car.

Open the notebook in Colab View it on GitHub

Colab opens a fresh copy each time. Save your own with File, then Save a copy in Drive.

A simulation: the positions are placed by hand to show the idea. In the notebook you do it for real, with an embedding model that runs in Colab.

What you do in this module

  1. Pull the text out of a PDF and look at what happens to its tables.
  2. Write a chunking function, and compare two chunk sizes.
  3. Embed sentences and see that similar meanings land close together.
  4. Embed the whole corpus, and search it by meaning.

Nothing in this module calls Gemini, so it uses none of your free requests.

Before you start: will a search by meaning connect HBC to hit by car?

Not reliably. A general embedding model learned from general text, where "HBC" rarely means hit by car. Veterinary abbreviations and synonyms are where general models struggle. The notebook shows you, and Module 3 fixes it.

Where next

Building a RAG system: Module 3: Retrieval

A shared question channel is on the way. When it opens, each answer will be written once and shared with everyone taking the course.