Images and Documents as Model Input
Send photos, screenshots, charts and PDFs to a model and get answers you can check and trust.
A taste of a lesson
The model reads totals from receipt photos but sometimes gets a digit wrong. How do I catch that?
Use the receipt against itself. Ask for line items and the total as separate structured fields, then check in code that the line items plus tax add up to the total. A misread digit usually breaks that sum. Also ask the model to return null when a value is unreadable, and send failed checks to a person. Cropping to the totals area and sending a sharper image helps too. How often does the sum check fail on your current sample?
Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.
What you will be able to do
- Send images and PDFs to multimodal models within size and cost limits
- Describe what current models read well and where they make mistakes
- Prompt for specific, structured extraction from images and documents
- Add checks that catch misreadings and route doubtful cases to people
- Handle images of people and personal documents responsibly
Lesson plan
- 1 Sending images and documents Include images and PDFs in requests and understand provider differences. Start
- 2 What models see well and badly Set realistic expectations for visual understanding. Start
- 3 Preparing inputs Crop, resize and order inputs for better results at lower cost. Start
- 4 Structured extraction Pull specific fields from images and documents into validated data. Start
- 5 Checks, review and alternatives Catch errors and know when another tool fits better. Start
- 6 Privacy and consent Handle images and documents containing personal data with care. Start
Try asking
About this tutor
For beginners building features that read images or documents: describing product photos, extracting fields from receipts, answering questions about a PDF report, reviewing screenshots. You learn how to send images and documents to models that accept them, how size and page count affect cost and limits, what current models do well and where they still make mistakes (small text, counting, precise numbers on charts, spatial details), and how to prompt for specific, structured extraction. The tutor also covers privacy for images of people and personal documents, and simple checks that catch misreadings.
Reviews
4.5
2 ratingsSample
- Kofi A.Sample
Resizing and cropping cut our image costs a lot without hurting accuracy. Good beginner level explanations.
- Isabel F.Sample
The line items must sum to the total check caught most of our misread receipts. Honest about what vision models still get wrong.
About the teacher
Structured output, tool calling and safe input handling for LLM applications that must behave predictably
9 tutors 308 lessons taught Sample
I teach the parts of LLM apps where free text has to meet real software: JSON that must parse, tools the model calls, images and documents coming in, and users who send things you did not plan for. I spent years writing integrations between messy systems, which taught me to treat every input as untrusted and every output as something...
See Greta's profile and tutorsMore like this
Other tutors on the same or nearby topics.