Skip to content
SamplePreview build: teacher profiles, ratings, reviews and lesson counts are sample data.
All tutors

Initialisation and normalisation layers

Understand how weight initialisation and normalisation keep deep networks trainable

Your first 3 replies from the tutor are free. Then you can continue the lesson for $11 from your credit.

A taste of a lesson

One question and the answer, as this tutor gives it.

Initialisation and normalisation layers AI tutor following Mira Okafor's plan
Student:

Why does He initialisation use 2 over fan in instead of 1 over fan in?

Tutor:

Start from the goal: keep each layer's output variance about equal to its input variance. For a linear unit, output variance is roughly fan_in times weight variance times input variance, so weight variance 1/fan_in keeps it steady. But ReLU sets every negative pre-activation to zero, which for zero mean inputs removes about half of the second moment. Doubling the weight variance to 2/fan_in compensates for that halving, so the signal neither shrinks nor grows through many ReLU layers. Your turn: for a leaky ReLU with a small negative slope, would the ideal factor be a bit above or a bit below 2?

Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.

What you will be able to do

  • Derive why activation variance should stay stable across layers
  • Choose Xavier, He or scaled initialisation for a given architecture
  • Compare batch, layer, RMS and group normalisation and their failure modes
  • Explain why pre norm transformers are easier to train at depth
  • Diagnose instability caused by initialisation or normalisation choices

Lesson plan

6 lessons. Pick one to start there.

  1. 1 Why initialisation matters Show how weight scale controls whether signals vanish or explode through depth. Start
  2. 2 Xavier, He and depth aware schemes Match initialisation variance to the activation function and architecture. Start
  3. 3 Batch normalisation in depth Understand what batch norm computes and where it causes problems. Start
  4. 4 Layer norm, RMSNorm and group norm Compare per example normalisation methods and pick one for a model. Start
  5. 5 Normalisation placement in transformers Explain the stability difference between pre norm and post norm designs. Start
  6. 6 Diagnosing instability Use activation and gradient statistics to locate initialisation and normalisation problems. Start

Try asking

Tap a question to start a lesson with it.

About this tutor

For practitioners who train deep networks and want to understand the settings that quietly decide whether training is stable. You will derive why activation variance must stay roughly constant across layers, compare zero, Xavier and He initialisation, and look at how residual networks and transformers scale their initial weights. The second half covers batch normalisation, layer normalisation, RMSNorm and group normalisation: what each normalises over, how they behave at training and inference time, and why pre norm transformers train more stably than post norm ones. Expect light derivations, small numerical experiments to reason about, and honest notes on which explanations are still debated.

Reviews

Students can review a tutor after a paid lesson. Nobody has yet.

About the teacher

Mira Okafor

I teach how neural networks learn, one small worked example at a time

9 tutors 4.5(21) 424 lessons taught Sample

I teach the core mechanics of deep learning: what a neuron computes, how a loss turns mistakes into numbers, and how gradients and optimisers change weights. My background is in building and training models for applied research teams, which mostly meant staring at loss curves that refused to go down. That shaped how I teach. I start every idea with...

See Mira's profile and tutors