Activation functions explained
Learn what ReLU, sigmoid, tanh, GELU and softmax do, and choose the right one for each layer
A taste of a lesson
My classifier outputs probabilities that all look almost equal. I used softmax and then cross entropy. What is wrong?
A very common cause: your cross entropy function already applies softmax internally, so you are applying it twice. Softmax of values that are already between 0 and 1 squashes them towards equal, which also flattens the gradients, so training stalls. Check the documentation for your loss: if it says it expects logits or raw scores, remove your own softmax layer and feed the raw outputs in. Keep softmax only when you want readable probabilities at prediction time. Quick check: what would softmax of (0.6, 0.3, 0.1) look like compared with the original numbers?
Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.
What you will be able to do
- Explain why neural networks need nonlinear activation functions
- Sketch sigmoid, tanh, ReLU and GELU and describe their slopes
- Choose the correct output activation for regression, binary, multi class and multi label tasks
- Avoid the double softmax bug and other activation and loss mismatches
Lesson plan
- 1 Why nonlinearity is needed Show that without activations a deep network is no more powerful than one linear layer. Start
- 2 Sigmoid and tanh Understand the S-shaped activations, their ranges and why they saturate. Start
- 3 ReLU and its relatives Learn why ReLU became the default hidden activation and what can go wrong with it. Start
- 4 Softmax and output probabilities Turn raw scores into class probabilities and read them sensibly. Start
- 5 Matching outputs to tasks and losses Pick output activation and loss together for each common task type. Start
Try asking
About this tutor
For beginners who have seen activation functions in diagrams or code and want to know why they exist and how to pick one. You will plot each common activation, compute its value and slope at a few points, and see how the slope affects learning in deep networks. The lessons cover hidden layer choices (ReLU and its relatives, smooth variants used in transformers) and output layer choices (none, sigmoid, softmax) matched to regression, binary, multi class and multi label tasks. You also learn the practical bugs that come from mismatching activations and loss functions. Light maths only, with every formula explained in words.
Reviews
4.5
2 ratingsSample
- Kenji T.Sample
I had exactly the double softmax bug in a class project. Found it in lesson four. The task to activation table is now taped to my monitor.
- Sofia R.Sample
Good explanations of saturation. I wanted a bit more on why GELU is used in transformers but the tutor was honest that the difference is mostly empirical.
About the teacher
I teach how neural networks learn, one small worked example at a time
9 tutors 424 lessons taught Sample
I teach the core mechanics of deep learning: what a neuron computes, how a loss turns mistakes into numbers, and how gradients and optimisers change weights. My background is in building and training models for applied research teams, which mostly meant staring at loss curves that refused to go down. That shaped how I teach. I start every idea with...
See Mira's profile and tutorsMore like this
Other tutors on the same or nearby topics.