Documents, chunking and embeddings
Get text out of the clinic's documents, including a real journal PDF, cut it into chunks, and turn each chunk into numbers that capture its meaning. Then search by meaning, and test whether "parvo" finds canine parvovirus and whether "HBC" finds hit by car.
Colab opens a fresh copy each time. Save your own with File, then Save a copy in Drive.
What you do in this module¶
- Pull the text out of a PDF and look at what happens to its tables.
- Write a chunking function, and compare two chunk sizes.
- Embed sentences and see that similar meanings land close together.
- Embed the whole corpus, and search it by meaning.
Nothing in this module calls Gemini, so it uses none of your free requests.
Before you start: will a search by meaning connect HBC to hit by car?
Not reliably. A general embedding model learned from general text, where "HBC" rarely means hit by car. Veterinary abbreviations and synonyms are where general models struggle. The notebook shows you, and Module 3 fixes it.
Where next
Building a RAG system: Module 3: Retrieval
A shared question channel is on the way. When it opens, each answer will be written once and shared with everyone taking the course.