Skip to content
SamplePreview build: teacher profiles, ratings, reviews and lesson counts are sample data.
All tutors

Image embeddings and visual search

Build image similarity and text to image search, and measure whether results are actually relevant

By Noor Siddiqui Intermediate Computer vision RAG and search 4.7(3) 71 lessons taught Sample

Your first 3 replies from the tutor are free. Then you can continue the lesson for $7 from your credit.

A taste of a lesson

One question and the answer, as this tutor gives it.

Image embeddings and visual search AI tutor following Noor Siddiqui's plan
Student:

My product photo search returns items photographed on the same white table, not similar products. What is going on?

Tutor:

The embedding is picking up the whole image, and a shared background and lighting setup can dominate the vector, especially if many products were shot the same way. Three things to try: crop to the product (a detector or a simple centre crop) before embedding; compare a model trained more for object identity, such as a contrastive or self supervised one; and, if you have pairs of the same product in different photos, fine tune on them. Measure each change on a small judged query set. Quick check: how would you build ten test queries that would expose the background problem?

Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.

What you will be able to do

  • Explain what image embeddings are and how their source model shapes similarity
  • Compute cosine similarity and set a near duplicate threshold from labelled pairs
  • Choose between exact and approximate nearest neighbour search
  • Evaluate visual search with a judged query set and recall at k
  • Identify bias and privacy risks, especially for images of people

Lesson plan

6 lessons. Pick one to start there.

  1. 1 What an embedding captures Understand embeddings as vectors where distance reflects a learned notion of similarity. Start
  2. 2 Contrastive image and text models See how shared image and text spaces make search by description possible. Start
  3. 3 Similarity and near duplicates Compute cosine similarity and choose duplicate thresholds from evidence. Start
  4. 4 Indexing and approximate search Choose a search approach that fits collection size and speed needs. Start
  5. 5 Evaluating search quality Build a judged query set and measure recall at k and top result precision. Start
  6. 6 Domain tuning, bias and privacy Adapt embeddings to your data and handle people in images responsibly. Start

Try asking

Tap a question to start a lesson with it.

About this tutor

For learners who want to find similar images, detect near duplicates or search photos with text. You will learn what an image embedding is, how embeddings from a classifier differ from those of contrastive image and text models, and how cosine similarity ranks results. Then you build up a search system in general terms: indexing vectors, approximate nearest neighbour search and its speed versus recall trade off, and thresholds for duplicate detection. A full lesson covers evaluation with recall at k and a small judged query set, and another covers adapting embeddings to your own domain. The course ends with bias and privacy, especially for images of people and faces.

Reviews

4.7

3 ratingsSample

  • Kofi A.Sample

    The privacy lesson was taken seriously, not tacked on. Helped me push back on a face matching feature request at work.

  • Beatriz C.Sample

    Clear on exact versus approximate search. I appreciated that it said a vector database was not needed for my small archive.

  • Hiroshi N.Sample

    Cropping before embedding fixed most of my background problem. The judged query set idea gave me a number to track instead of vibes.

About the teacher

Noor Siddiqui

Computer vision taught through real images, real failure cases and careful evaluation

9 tutors 4.5(21) 406 lessons taught Sample

I teach computer vision: classification, detection, segmentation, document understanding, video and the newer models that combine images with language. Most of my work has been building vision systems that had to hold up outside the lab, under odd lighting, unusual cameras and labels that were not quite consistent. So my lessons spend as much time on data and evaluation as...

See Noor's profile and tutors