Skip to content
SamplePreview build: teacher profiles, ratings, reviews and lesson counts are sample data.
All tutors

Vision transformers in depth

Understand patch tokens, position embeddings, data needs and compute trade offs in vision transformers

Your first 3 replies from the tutor are free. Then you can continue the lesson for $12 from your credit.

A taste of a lesson

One question and the answer, as this tutor gives it.

Vision transformers in depth AI tutor following Noor Siddiqui's plan
Student:

If I fine tune a ViT at 384 instead of 224 with 16 pixel patches, what changes besides speed?

Tutor:

The patch grid grows from 14 by 14 to 24 by 24, so you go from 196 to 576 tokens. The learned position embeddings only exist for the 14 by 14 grid, so they must be interpolated in 2D to 24 by 24, and a short fine tune lets the model adjust. Attention compute grows roughly with the square of the token count, about (576/196) squared, so around 8.6 times, plus higher activation memory, which may force a smaller batch. The gain is finer detail per object. Quick check: what token count would 384 give with 32 pixel patches?

Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.

What you will be able to do

  • Compute patch counts, token shapes and parameter shapes for a vision transformer
  • Handle position embeddings when changing input resolution
  • Explain why vision transformers need more data or pretraining than CNNs
  • Compare plain, hierarchical and windowed designs for dense prediction tasks
  • Estimate how resolution and patch size change attention compute

Lesson plan

6 lessons. Pick one to start there.

  1. 1 Images as sequences of patches Compute patch counts and the patch embedding projection for a standard setup. Start
  2. 2 Position embeddings and resolution Give tokens spatial information and adapt it when the input size changes. Start
  3. 3 Inductive bias and data hunger Explain the data requirements of ViTs compared with CNNs. Start
  4. 4 Self supervised pretraining Understand how masked modelling and self distillation give strong ViT features. Start
  5. 5 Hierarchical and windowed designs See how local attention and patch merging adapt ViTs to dense prediction. Start
  6. 6 Compute, resolution and choosing a design Estimate costs and decide between ViT variants and CNNs for a task. Start

Try asking

Tap a question to start a lesson with it.

About this tutor

For practitioners who know transformers from text or CNNs from vision and want a solid understanding of vision transformers. You will work through patch embedding with real numbers, compare class tokens with pooling, and see what happens to position embeddings when image resolution changes. The course looks squarely at inductive bias: why vision transformers needed far more data or pretraining than CNNs, and how augmentation, distillation and self supervised pretraining such as masked image modelling closed the gap. Later lessons cover hierarchical and windowed variants built for detection and segmentation, and the compute cost of higher resolution. Expect shape tracking, cost estimates and careful reasoning about trade offs.

Reviews

Students can review a tutor after a paid lesson. Nobody has yet.

About the teacher

Noor Siddiqui

Computer vision taught through real images, real failure cases and careful evaluation

9 tutors 4.5(21) 406 lessons taught Sample

I teach computer vision: classification, detection, segmentation, document understanding, video and the newer models that combine images with language. Most of my work has been building vision systems that had to hold up outside the lab, under odd lighting, unusual cameras and labels that were not quite consistent. So my lessons spend as much time on data and evaluation as...

See Noor's profile and tutors