Images
Models that read images and models that make them build on the same ideas as language models. An image becomes tokens, and a new image is made by removing noise one step at a time.
The same ideas carry over from text to pictures. A model that reads an image first turns it into tokens, as it does with words. A model that makes an image predicts too: it predicts noise and takes it away, one step at a time.
How a model sees¶
A vision model does not take in an image the way you do. It cuts the image into small square patches, and each patch becomes a token that sits beside the text tokens in the same model. Turn the paw print into tokens with big, medium and small patches.
Watch the token count. Smaller patches keep more detail but cost more tokens, and more computing for every image. Bigger patches are cheaper but blur fine detail, which is one reason a model can miss something small in an image.
How a model makes an image¶
An image generator starts from pure noise. At each step the model predicts the noise and removes a little of it, guided by the prompt. Pick a prompt, play it, and watch the picture emerge over 20 steps. Then switch the prompt and start again from the same noise: you get a different image.
Now switch to training mode. Noise is added to a real image, step by step, and the model learns to predict that noise. Training and generating are the same process, run in opposite directions.
What to take away¶
Images become tokens, and generated images are predictions too. A sharp, confident picture is not proof that the detail in it is right, so the same checking habits apply.
Where next
How AI works: Back to How AI works: all five modules
A shared question channel is on the way. When it opens, each answer will be written once and shared with everyone taking the course.