Activation functions and their derivatives
Overview
Section titled “Overview”| Module | M01.3 · build · Python · Pass 2 · 3 to 4 h |
| You build | python/tinyllm/num/activations.py: sigmoid, tanh, relu, softplus, gelu_tanh, gelu_erf, silu, each with its derivative d<name> |
| Contract | course/contracts/py/tinyllm/num/activations.pyi |
| Tests | course/tests/M01.3/test_activations.py, with the PyTorch golden values in course/fixtures/M01.3/activations_torch.npz (what they check: section 4) |
| Needs | M02.1 erf_series, which gelu_erf calls (or --ref-deps). Reading: M01.1 (derivatives), M00.1 (exp and log) |
| Used by | M08.1 checks its dual-number derivatives against these closed forms · later L0.2 wraps them as autograd ops, L3.2 uses sigmoid and tanh for LSTM gates, L7.2 uses SiLU in SwiGLU, L9.6 ports SiLU to C |
| Milestone | MS-P2 (the Pass 2 gate) |
| Optional depth | OpenStax, Calculus Volume 1 (free), sections 3.3 to 3.6 (derivative rules, the chain rule) and 3.9 (exponential and logarithm); Hendrycks and Gimpel, “Gaussian Error Linear Units” (2016); Ramachandran, Zoph, and Le, “Searching for Activation Functions” (2017, SiLU/swish) |
Key Takeaways
Section titled “Key Takeaways”- An activation is a nonlinear function applied elementwise between linear layers; without one, any stack of layers is a single matrix. Training needs each one’s derivative, which the chain rule multiplies into every gradient (
test_hand_example_at_one,test_derivatives_pass_gradcheck). - The derivatives come from three rules: , the quotient or chain rule, and the product rule. , , , , (
test_matches_torch_golden). - Overflow is avoided, not silenced: never overflows, so sigmoid and softplus are written in terms of it, and every function is finite and warning-free at in float32 and float64 (
test_no_overflow_for_any_finite_input). - Writing $\sigma’ $ as instead of keeps full relative accuracy in the tails (
test_derivative_tails_keep_relative_accuracy). - The two GELUs are different functions (they differ in the fourth digit at ), and by convention: match the checkpoint you load (
test_hand_example_gelus,test_relu_derivative_at_zero_is_zero).
How to work this chapter
Section titled “How to work this chapter”ol start M01.3 # stubs python/tinyllm/num/activations.py into your repool tests M01.3 # read the test catalog first: rung R0, you write no tests hereol check M01.3 # exit code is the verdictol check M01.3 --ref-deps # only if your M02.1 is not passing yetol diff M01.3 # after passing: your code against the reference1. Why now
Section titled “1. Why now”The tracer’s bigram is one table lookup: no layers, nothing between them. Pass 2 replaces it with models that have hidden layers (the MLP of L0.4, later recurrent cells and transformer blocks), and a hidden layer is a matrix product followed by an activation function. Your autograd engine (L0.2) needs both halves of every activation: the function for the forward pass and its derivative for the backward pass. Get a derivative wrong and the model still trains, slowly and to a worse loss, with no error anywhere; get the forward pass numerically careless and one large logit after a bad optimizer step turns the whole batch into nan. This module derives each activation and its derivative from the rules of calculus, writes them so that no input can overflow, and checks them against PyTorch at 1000 points from to .
2. Principles
Section titled “2. Principles”| Symbol | Meaning | Type / shape |
|---|---|---|
| the input, one number per element | float32 or float64 array | |
| the logistic sigmoid | same shape as | |
| hyperbolic tangent | same | |
| same | ||
| same | ||
| standard normal CDF, | same | |
| standard normal density, ; note | same | |
, computed by M02.1’s erf_series | same | |
or df | the derivative of with respect to | same |
2.1 Why a nonlinearity
Section titled “2.1 Why a nonlinearity”A linear layer computes . Two in a row compute : again one linear layer, so depth would add nothing. Inserting a nonlinear between them, , breaks the collapse, and with enough hidden units such networks can approximate any continuous function on a bounded set. Every below is applied elementwise: output depends only on input .
2.2 Sigmoid and tanh
Section titled “2.2 Sigmoid and tanh” squashes the line into : , as , as , and . It turns a score into a probability (LSTM gates in L3.2, logistic regression in M07.7). Its derivative, by the chain rule on :
is a rescaled sigmoid, , with range . By the quotient rule, .
Evaluating without overflow. For , overflows to inf. The value happens to be right, but the overflow raises a floating-point warning (an error under np.errstate(over="raise")), and the same pattern elsewhere produces $\infty / \infty = $ nan. The fix uses only , which can never overflow:
(the second line is the first with numerator and denominator multiplied by ).
Relative accuracy in the tails. . Computed as with , the subtraction leaves only the last few bits of , and the result is : three correct digits. never subtracts and keeps all sixteen. Tiny derivatives matter because the chain rule multiplies them into other factors.
2.3 ReLU and softplus
Section titled “2.3 ReLU and softplus”is the identity for positive inputs and zero otherwise; its derivative is 1 for and 0 for . At the graph has a corner and no derivative exists. Frameworks pick a value: PyTorch uses 0, and so does this course, because a gradient that differs from the framework’s exactly at zeros makes your gradients disagree with any checkpoint’s training history there.
is a smooth relu: about for large , about for very negative . Its derivative is . Naively, overflows, and gives exactly 0 because rounds to 1. The stable form splits off the large part and uses log1p, which computes accurately for tiny :
2.4 GELU, exact and approximate
Section titled “2.4 GELU, exact and approximate”The Gaussian error linear unit weights its input by the probability that a standard normal variable is below it:
by the product rule, since . Your gelu_erf calls erf_series from M02.1. Before erf was fast on GPUs, GPT-2 and BERT used a tanh approximation:
Its derivative needs the product rule and the chain rule: , with . The two GELUs differ by up to (near ); a checkpoint trained with one must be served with the same one. For , is already exactly, and would overflow for ( in float32), so is clamped to inside . Likewise underflows to 0 beyond , and inside it is clamped there.
2.5 SiLU
Section titled “2.5 SiLU” (also called swish) is the gate inside SwiGLU, the MLP of Llama-family models (L7.2). Product rule:
Forgetting the product rule and writing alone is the classic bug: it is right at and wrong everywhere else.
2.6 Identities the tests use
Section titled “2.6 Identities the tests use”; ; and for in , (each is times something that adds to 1 at , or for softplus ). These hold for every , so test_identities checks them at 500 seeded points.
2.7 Dtypes
Section titled “2.7 Dtypes”A float32 tensor stays float32 and a float64 tensor stays float64: promoting silently doubles memory and breaks the bit-for-bit comparisons with the C kernels of Pass 6. Integer input is computed in float64. The GELU with erf is computed in float64 internally and cast back.
3. Worked example by hand
Section titled “3. Worked example by hand”At , with :
| Function | Computation | Value | Derivative | Value |
|---|---|---|---|---|
| sigmoid | 0.7310586 | 0.1966119 | ||
| tanh | 0.7615942 | 0.4199743 | ||
| relu | 1 | 1 | ||
| softplus | 1.3132617 | 0.7310586 | ||
| silu | 0.7310586 | 0.9276705 | ||
| gelu (erf) | 0.8413447 | 1.0833155 | ||
| gelu (tanh) | , | 0.8411920 | (section 2.4) |
The exact and approximate GELU differ by here. These numbers are the first two tests, test_hand_example_at_one and test_hand_example_gelus.
4. The interface
Section titled “4. The interface”def sigmoid(x: ArrayLike) -> NDArray: ... # and dsigmoiddef tanh(x: ArrayLike) -> NDArray: ... # and dtanhdef relu(x: ArrayLike) -> NDArray: ... # and drelu (0 at x = 0)def softplus(x: ArrayLike) -> NDArray: ... # and dsoftplus (= sigmoid)def gelu_tanh(x: ArrayLike) -> NDArray: ... # and dgelu_tanhdef gelu_erf(x: ArrayLike) -> NDArray: ... # and dgelu_erf, through erf_series (M02.1)def silu(x: ArrayLike) -> NDArray: ... # and dsiluEvery function keeps the shape and the float32 or float64 dtype of its input, returns finite values for finite inputs, and emits no overflow, invalid, or divide warning.
What the tests check
Section titled “What the tests check”| Test | KIND | Checks | Why it matters downstream |
|---|---|---|---|
test_hand_example_at_one | unit, smoke | section 3’s table for sigmoid, tanh, relu, softplus, silu | you and the tests agree on every definition |
test_hand_example_gelus | unit, smoke | both GELUs and the exact derivative at 1 | the two GELUs are not interchangeable |
test_matches_torch_golden | golden | each function and derivative against PyTorch autograd at 1000 points in , float64 tolerances | L0.2 is checked against PyTorch forward and backward |
test_derivatives_pass_gradcheck | gradcheck | each smooth derivative against the frozen central-difference gradcheck | the derivative belongs to your function, not just to a formula |
test_relu_gradcheck_away_from_zero | gradcheck | relu’s derivative away from the kink | the same, for relu |
test_identities | property | the identities of section 2.6 at 500 seeded points | sign and symmetry slips |
test_derivative_tails_keep_relative_accuracy | boundary | and to 14 digits | tiny gradients still multiply other factors |
test_no_overflow_for_any_finite_input | boundary | and the largest float, float32 and float64, with warnings raised as errors | one bad logit must not become nan |
test_dtype_and_shape_preserved | unit | float32 stays float32, shapes kept, ints become float64 | memory and C-kernel comparisons |
test_relu_derivative_at_zero_is_zero | boundary | PyTorch’s convention |
5. Pitfalls
Section titled “5. Pitfalls”| Pitfall | Symptom | Caught by |
|---|---|---|
| 1. exponentials that can overflow: , , an unclamped | overflow warnings, inf, or nan at | test_no_overflow_for_any_finite_input (mutants s10, s11, s12) |
| 2. subtracting from 1 in the tails: , | three correct digits at ; | test_derivative_tails_keep_relative_accuracy (mutants s03, s09) |
| 3. | gradients differ from PyTorch’s exactly at zeros | test_relu_derivative_at_zero_is_zero, test_matches_torch_golden (mutant s04) |
| 4. using the tanh approximation for the exact GELU (or the reverse) | off at ; a loaded checkpoint drifts | test_hand_example_gelus (mutant s05) |
| 5. forgetting the product rule: , without the term | right at 0, wrong everywhere else | test_hand_example_at_one (mutant s02), test_derivatives_pass_gradcheck (mutant s06) |
| 6. promoting float32 to float64 | a float64 array where the kernel expects float32 | test_dtype_and_shape_preserved (mutant s13) |
| dropping the square in | instead of 0.420 | test_hand_example_at_one (mutant s01) |
| a wrong GELU constant (0.0447 for 0.044715) | the 6th digit of every GELU | test_hand_example_gelus (mutant s07) |
| subtracting the log1p term in softplus | test_hand_example_at_one (mutant s08) |
6. Where it’s used next
Section titled “6. Where it’s used next”| Direction | Module | How it uses this |
|---|---|---|
| Back | M02.1 | gelu_erf and dgelu_erf call erf_series(x / sqrt(2), 160) |
| Back | M01.1 | derivatives as limits; the frozen gradcheck in the tests is M01.1’s central difference (reading) |
| Back | M00.1 | and (reading) |
| Forward | M08.1 | dual numbers compute the same derivatives by forward-mode autodiff; its tests compare them with yours to |
| Forward | L0.2 | the elementwise ops of tinyllm.F (forward and backward) |
| Forward | L3.2 | LSTM gates: sigmoid for input, forget, output; tanh for the cell |
| Forward | L7.2 | SwiGLU |
| Forward | L9.6 | tl_silu_mul_f32 in C, checked against your Python |
L0.2, L3.2, L7.2, and L9.6 join used_by when they are authored (course/DEVIATIONS.md row B31-02).
Going further
Section titled “Going further”| Your piece | Production equivalent | What it adds | Where to look |
|---|---|---|---|
sigmoid, silu, gelu_* | PyTorch torch.nn.functional | fused CUDA kernels, autocast to bf16, in-place variants | aten/src/ATen/native/Activation.cpp, cuda/ActivationGeluKernel.cu |
stable softplus | PyTorch F.softplus(beta, threshold) | returns above a threshold (default 20) instead of computing the log, trading for speed | aten/src/ATen/native/cpu/Activation.cpp |
gelu_erf via a series | erf in libm, erff on GPUs | minimax rational approximations, a few ulps everywhere | glibc sysdeps/ieee754/dbl-64/s_erf.c |
| the activation zoo | llama.cpp ggml_silu, ggml_gelu | lookup tables for f16 GELU, vectorized SiLU | ggml/src/ggml-cpu/vec.h |