Multimodal models: text, images and sound together
See how one model can read images, hear audio and write text, and where it still stumbles
A taste of a lesson
I uploaded a photo of a bookshelf and asked how many books. It said 34, but there are 52. Why?
Counting many similar objects is a known weak spot. The model sees the image as a grid of patches turned into vectors, often after resizing, so thin book spines blur together and there is no step that counts items one by one. It produces a plausible number rather than a tally. Better approaches: crop the shelf into sections and count each, or ask it to list books it can read by title, then count the list. Try the section approach. How close does it get now?
Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.
What you will be able to do
- Explain how images and audio are converted into tokens a language model can use
- Distinguish multimodal understanding from image or audio generation
- Predict which visual and audio tasks are likely to be reliable or weak
- Use cropping, quoting and verification to improve multimodal results
Lesson plan
- 1 How images become tokens Understand the patch embedding approach to image input. Start
- 2 Sound and video See how audio and video are handled. Start
- 3 Understanding versus generating Tell apart reading media from producing it. Start
- 4 Strengths and weak spots Predict which tasks are reliable. Start
- 5 Better results and privacy Improve accuracy and protect people in images and recordings. Start
Try asking
About this tutor
For users and builders curious about models that take photos, screenshots, documents, audio and video as input, and sometimes produce images or speech. You learn in plain but accurate terms how images are turned into tokens through patch embeddings, how audio is encoded, how understanding and generation differ, and why some tasks are strong (describing scenes, reading clear text, answering questions about charts) while others remain weak (counting many objects, precise measurements, small or rotated text, spatial reasoning). You practise with realistic tasks and learn privacy points specific to images and voice.
Reviews
Students can review a tutor after a paid lesson. Nobody has yet.
About the teacher
I explain the kinds of AI models, what they cost to run and how to run one yourself
9 tutors 328 lessons taught Sample
I teach the practical side of modern models: reasoning models, multimodal models, open and closed weights, running a model on your own computer, and the money, energy and hardware behind every answer. I like starting with something you can see or measure, such as the memory a model needs or the number of tokens a task uses, then explaining the...
See Benedict's profile and tutorsMore like this
Other tutors on the same or nearby topics.