Skip to content

Activation functions and their derivatives

ModuleM01.3 · build · Python · Pass 2 · 3 to 4 h
You buildpython/tinyllm/num/activations.py: sigmoid, tanh, relu, softplus, gelu_tanh, gelu_erf, silu, each with its derivative d<name>
Contractcourse/contracts/py/tinyllm/num/activations.pyi
Testscourse/tests/M01.3/test_activations.py, with the PyTorch golden values in course/fixtures/M01.3/activations_torch.npz (what they check: section 4)
NeedsM02.1 erf_series, which gelu_erf calls (or --ref-deps). Reading: M01.1 (derivatives), M00.1 (exp and log)
Used byM08.1 checks its dual-number derivatives against these closed forms · later L0.2 wraps them as autograd ops, L3.2 uses sigmoid and tanh for LSTM gates, L7.2 uses SiLU in SwiGLU, L9.6 ports SiLU to C
MilestoneMS-P2 (the Pass 2 gate)
Optional depthOpenStax, Calculus Volume 1 (free), sections 3.3 to 3.6 (derivative rules, the chain rule) and 3.9 (exponential and logarithm); Hendrycks and Gimpel, “Gaussian Error Linear Units” (2016); Ramachandran, Zoph, and Le, “Searching for Activation Functions” (2017, SiLU/swish)
  • An activation is a nonlinear function applied elementwise between linear layers; without one, any stack of layers is a single matrix. Training needs each one’s derivative, which the chain rule multiplies into every gradient (test_hand_example_at_one, test_derivatives_pass_gradcheck).
  • The derivatives come from three rules: (ex)′=ex(e^x)' = e^x, the quotient or chain rule, and the product rule. σ′=σ(x)σ(−x)\sigma' = \sigma(x)\sigma(-x), tanh⁡′=1−tanh⁡2\tanh' = 1 - \tanh^2, softplus′=σ\mathrm{softplus}' = \sigma, silu′=σ(x)(1+xσ(−x))\mathrm{silu}' = \sigma(x)(1 + x\sigma(-x)), gelu′=Φ+xφ\mathrm{gelu}' = \Phi + x\varphi (test_matches_torch_golden).
  • Overflow is avoided, not silenced: e−∣x∣e^{-|x|} never overflows, so sigmoid and softplus are written in terms of it, and every function is finite and warning-free at x=±1030x = \pm 10^{30} in float32 and float64 (test_no_overflow_for_any_finite_input).
  • Writing $\sigma’ $ as σ(x)σ(−x)\sigma(x)\sigma(-x) instead of σ(1−σ)\sigma(1 - \sigma) keeps full relative accuracy in the tails (test_derivative_tails_keep_relative_accuracy).
  • The two GELUs are different functions (they differ in the fourth digit at x=1x = 1), and relu′(0)=0\mathrm{relu}'(0) = 0 by convention: match the checkpoint you load (test_hand_example_gelus, test_relu_derivative_at_zero_is_zero).
Terminal window
ol start M01.3 # stubs python/tinyllm/num/activations.py into your repo
ol tests M01.3 # read the test catalog first: rung R0, you write no tests here
ol check M01.3 # exit code is the verdict
ol check M01.3 --ref-deps # only if your M02.1 is not passing yet
ol diff M01.3 # after passing: your code against the reference

The tracer’s bigram is one table lookup: no layers, nothing between them. Pass 2 replaces it with models that have hidden layers (the MLP of L0.4, later recurrent cells and transformer blocks), and a hidden layer is a matrix product followed by an activation function. Your autograd engine (L0.2) needs both halves of every activation: the function for the forward pass and its derivative for the backward pass. Get a derivative wrong and the model still trains, slowly and to a worse loss, with no error anywhere; get the forward pass numerically careless and one large logit after a bad optimizer step turns the whole batch into nan. This module derives each activation and its derivative from the rules of calculus, writes them so that no input can overflow, and checks them against PyTorch at 1000 points from −40-40 to 4040.

SymbolMeaningType / shape
xxthe input, one number per elementfloat32 or float64 array
σ(x)\sigma(x)the logistic sigmoid 11+e−x\frac{1}{1 + e^{-x}}same shape as xx
tanh⁡x\tanh xhyperbolic tangent ex−e−xex+e−x\frac{e^x - e^{-x}}{e^x + e^{-x}}same
relu(x)\mathrm{relu}(x)max⁡(x,0)\max(x, 0)same
softplus(x)\mathrm{softplus}(x)ln⁡(1+ex)\ln(1 + e^x)same
Φ(x)\Phi(x)standard normal CDF, 12(1+erf⁡(x/2))\frac{1}{2}\left(1 + \operatorname{erf}(x/\sqrt 2)\right)same
φ(x)\varphi(x)standard normal density, 12πe−x2/2\frac{1}{\sqrt{2\pi}} e^{-x^2/2}; note Φ′=φ\Phi' = \varphisame
erf⁡(z)\operatorname{erf}(z)2π∫0ze−t2 dt\frac{2}{\sqrt\pi}\int_0^z e^{-t^2}\,dt, computed by M02.1’s erf_seriessame
f′f' or dfthe derivative of ff with respect to xxsame

A linear layer computes y=Wx+by = Wx + b. Two in a row compute W2(W1x+b1)+b2=(W2W1)x+(W2b1+b2)W_2(W_1 x + b_1) + b_2 = (W_2 W_1)x + (W_2 b_1 + b_2): again one linear layer, so depth would add nothing. Inserting a nonlinear ff between them, W2f(W1x+b1)+b2W_2 f(W_1 x + b_1) + b_2, breaks the collapse, and with enough hidden units such networks can approximate any continuous function on a bounded set. Every ff below is applied elementwise: output ii depends only on input ii.

σ(x)=11+e−x\sigma(x) = \frac{1}{1 + e^{-x}} squashes the line into (0,1)(0, 1): σ(0)=12\sigma(0) = \frac12, σ(x)→1\sigma(x) \to 1 as x→∞x \to \infty, →0\to 0 as x→−∞x \to -\infty, and σ(−x)=1−σ(x)\sigma(-x) = 1 - \sigma(x). It turns a score into a probability (LSTM gates in L3.2, logistic regression in M07.7). Its derivative, by the chain rule on (1+e−x)−1(1 + e^{-x})^{-1}:

σ′(x)=e−x(1+e−x)2=11+e−x⋅e−x1+e−x=σ(x) σ(−x)=σ(x)(1−σ(x)).\sigma'(x) = \frac{e^{-x}}{(1 + e^{-x})^2} = \frac{1}{1 + e^{-x}} \cdot \frac{e^{-x}}{1 + e^{-x}} = \sigma(x)\,\sigma(-x) = \sigma(x)(1 - \sigma(x)) .

tanh⁡\tanh is a rescaled sigmoid, tanh⁡x=2σ(2x)−1\tanh x = 2\sigma(2x) - 1, with range (−1,1)(-1, 1). By the quotient rule, tanh⁡′x=1−tanh⁡2x\tanh' x = 1 - \tanh^2 x.

Evaluating without overflow. For x=−1000x = -1000, e−x=e1000e^{-x} = e^{1000} overflows to inf. The value 1/(1+∞)=01/(1 + \infty) = 0 happens to be right, but the overflow raises a floating-point warning (an error under np.errstate(over="raise")), and the same pattern elsewhere produces $\infty / \infty = $ nan. The fix uses only z=e−∣x∣∈(0,1]z = e^{-|x|} \in (0, 1], which can never overflow:

σ(x)={11+zx≥0z1+zx<0\sigma(x) = \begin{cases} \frac{1}{1 + z} & x \ge 0 \\ \frac{z}{1 + z} & x < 0 \end{cases}

(the second line is the first with numerator and denominator multiplied by exe^{x}).

Relative accuracy in the tails. σ′(30)=9.3576×10−14\sigma'(30) = 9.3576 \times 10^{-14}. Computed as s(1−s)s(1 - s) with s=σ(30)=1−9.36×10−14s = \sigma(30) = 1 - 9.36 \times 10^{-14}, the subtraction 1−s1 - s leaves only the last few bits of ss, and the result is 9.348×10−149.348 \times 10^{-14}: three correct digits. σ(30)σ(−30)\sigma(30)\sigma(-30) never subtracts and keeps all sixteen. Tiny derivatives matter because the chain rule multiplies them into other factors.

relu(x)=max⁡(x,0)\mathrm{relu}(x) = \max(x, 0) is the identity for positive inputs and zero otherwise; its derivative is 1 for x>0x > 0 and 0 for x<0x < 0. At x=0x = 0 the graph has a corner and no derivative exists. Frameworks pick a value: PyTorch uses 0, and so does this course, because a gradient that differs from the framework’s exactly at zeros makes your gradients disagree with any checkpoint’s training history there.

softplus(x)=ln⁡(1+ex)\mathrm{softplus}(x) = \ln(1 + e^x) is a smooth relu: about xx for large xx, about exe^x for very negative xx. Its derivative is ex1+ex=σ(x)\frac{e^x}{1 + e^x} = \sigma(x). Naively, ln⁡(1+e1000)\ln(1 + e^{1000}) overflows, and ln⁡(1+e−40)\ln(1 + e^{-40}) gives exactly 0 because 1+4.2×10−181 + 4.2 \times 10^{-18} rounds to 1. The stable form splits off the large part and uses log1p, which computes ln⁡(1+u)\ln(1 + u) accurately for tiny uu:

softplus(x)=max⁡(x,0)+log1p⁡ ⁣(e−∣x∣).\mathrm{softplus}(x) = \max(x, 0) + \operatorname{log1p}\!\left(e^{-|x|}\right) .

The Gaussian error linear unit weights its input by the probability that a standard normal variable is below it:

gelu(x)=x Φ(x)=x2(1+erf⁡x2),gelu′(x)=Φ(x)+x φ(x)\mathrm{gelu}(x) = x\,\Phi(x) = \frac{x}{2}\left(1 + \operatorname{erf}\frac{x}{\sqrt 2}\right), \qquad \mathrm{gelu}'(x) = \Phi(x) + x\,\varphi(x)

by the product rule, since Φ′=φ\Phi' = \varphi. Your gelu_erf calls erf_series from M02.1. Before erf was fast on GPUs, GPT-2 and BERT used a tanh approximation:

gelutanh⁡(x)=x2(1+tanh⁡u),u=2/π (x+0.044715 x3).\mathrm{gelu}_{\tanh}(x) = \frac{x}{2}\left(1 + \tanh u\right), \quad u = \sqrt{2/\pi}\,(x + 0.044715\, x^3) .

Its derivative needs the product rule and the chain rule: 12(1+tanh⁡u)+x2(1−tanh⁡2u) u′\frac12(1 + \tanh u) + \frac{x}{2}(1 - \tanh^2 u)\,u', with u′=2/π(1+3⋅0.044715 x2)u' = \sqrt{2/\pi}(1 + 3 \cdot 0.044715\, x^2). The two GELUs differ by up to 4.7×10−44.7 \times 10^{-4} (near x=2.7x = 2.7); a checkpoint trained with one must be served with the same one. For ∣x∣>10|x| > 10, tanh⁡u\tanh u is already ±1\pm 1 exactly, and x3x^3 would overflow for ∣x∣>10102|x| > 10^{102} (7×10127 \times 10^{12} in float32), so xx is clamped to [−10,10][-10, 10] inside uu. Likewise φ\varphi underflows to 0 beyond ∣x∣=40|x| = 40, and x2x^2 inside it is clamped there.

silu(x)=x σ(x)\mathrm{silu}(x) = x\,\sigma(x) (also called swish) is the gate inside SwiGLU, the MLP of Llama-family models (L7.2). Product rule:

silu′(x)=σ(x)+x σ(x)(1−σ(x))=σ(x)(1+x σ(−x)).\mathrm{silu}'(x) = \sigma(x) + x\,\sigma(x)(1 - \sigma(x)) = \sigma(x)\left(1 + x\,\sigma(-x)\right) .

Forgetting the product rule and writing σ(x)\sigma(x) alone is the classic bug: it is right at x=0x = 0 and wrong everywhere else.

σ(x)+σ(−x)=1\sigma(x) + \sigma(-x) = 1; tanh⁡x=2σ(2x)−1\tanh x = 2\sigma(2x) - 1; and for ff in {softplus,silu,gelu,gelutanh⁡}\{\mathrm{softplus}, \mathrm{silu}, \mathrm{gelu}, \mathrm{gelu}_{\tanh}\}, f(x)−f(−x)=xf(x) - f(-x) = x (each is xx times something that adds to 1 at ±x\pm x, or for softplus ln⁡1+ex1+e−x=ln⁡ex\ln\frac{1 + e^x}{1 + e^{-x}} = \ln e^x). These hold for every xx, so test_identities checks them at 500 seeded points.

A float32 tensor stays float32 and a float64 tensor stays float64: promoting silently doubles memory and breaks the bit-for-bit comparisons with the C kernels of Pass 6. Integer input is computed in float64. The GELU with erf is computed in float64 internally and cast back.

At x=1x = 1, with e−1=0.3678794e^{-1} = 0.3678794:

FunctionComputationValueDerivativeValue
sigmoid1/1.36787941 / 1.36787940.73105860.7310586×0.26894140.7310586 \times 0.26894140.1966119
tanh2σ(2)−12\sigma(2) - 10.76159421−0.761594221 - 0.7615942^20.4199743
relumax⁡(1,0)\max(1, 0)1x>0x > 01
softplusln⁡(1+2.7182818)\ln(1 + 2.7182818)1.3132617σ(1)\sigma(1)0.7310586
silu1×0.73105861 \times 0.73105860.73105860.7310586×(1+0.2689414)0.7310586 \times (1 + 0.2689414)0.9276705
gelu (erf)1×Φ(1)1 \times \Phi(1)0.8413447Φ(1)+φ(1)=0.8413447+0.2419707\Phi(1) + \varphi(1) = 0.8413447 + 0.24197071.0833155
gelu (tanh)u=0.7978846×1.044715=0.8335620u = 0.7978846 \times 1.044715 = 0.8335620, 12(1+tanh⁡u)=12(1+0.6823840)\frac12(1 + \tanh u) = \frac12(1 + 0.6823840)0.8411920(section 2.4)

The exact and approximate GELU differ by 1.5×10−41.5 \times 10^{-4} here. These numbers are the first two tests, test_hand_example_at_one and test_hand_example_gelus.

def sigmoid(x: ArrayLike) -> NDArray: ... # and dsigmoid
def tanh(x: ArrayLike) -> NDArray: ... # and dtanh
def relu(x: ArrayLike) -> NDArray: ... # and drelu (0 at x = 0)
def softplus(x: ArrayLike) -> NDArray: ... # and dsoftplus (= sigmoid)
def gelu_tanh(x: ArrayLike) -> NDArray: ... # and dgelu_tanh
def gelu_erf(x: ArrayLike) -> NDArray: ... # and dgelu_erf, through erf_series (M02.1)
def silu(x: ArrayLike) -> NDArray: ... # and dsilu

Every function keeps the shape and the float32 or float64 dtype of its input, returns finite values for finite inputs, and emits no overflow, invalid, or divide warning.

TestKINDChecksWhy it matters downstream
test_hand_example_at_oneunit, smokesection 3’s table for sigmoid, tanh, relu, softplus, siluyou and the tests agree on every definition
test_hand_example_gelusunit, smokeboth GELUs and the exact derivative at 1the two GELUs are not interchangeable
test_matches_torch_goldengoldeneach function and derivative against PyTorch autograd at 1000 points in [−40,40][-40, 40], float64 tolerancesL0.2 is checked against PyTorch forward and backward
test_derivatives_pass_gradcheckgradcheckeach smooth derivative against the frozen central-difference gradcheckthe derivative belongs to your function, not just to a formula
test_relu_gradcheck_away_from_zerogradcheckrelu’s derivative away from the kinkthe same, for relu
test_identitiespropertythe identities of section 2.6 at 500 seeded pointssign and symmetry slips
test_derivative_tails_keep_relative_accuracyboundaryσ′(±30)\sigma'(\pm 30) and softplus(−40)\mathrm{softplus}(-40) to 14 digitstiny gradients still multiply other factors
test_no_overflow_for_any_finite_inputboundary±1030\pm 10^{30} and the largest float, float32 and float64, with warnings raised as errorsone bad logit must not become nan
test_dtype_and_shape_preservedunitfloat32 stays float32, shapes kept, ints become float64memory and C-kernel comparisons
test_relu_derivative_at_zero_is_zeroboundaryrelu′(0)=0\mathrm{relu}'(0) = 0PyTorch’s convention
PitfallSymptomCaught by
1. exponentials that can overflow: 1/(1+e−x)1/(1 + e^{-x}), ln⁡(1+ex)\ln(1 + e^x), an unclamped x3x^3overflow warnings, inf, or nan at x=±1000x = \pm 1000test_no_overflow_for_any_finite_input (mutants s10, s11, s12)
2. subtracting from 1 in the tails: σ(1−σ)\sigma(1 - \sigma), ln⁡(1+e−40)\ln(1 + e^{-40})three correct digits at x=30x = 30; softplus(−40)=0\mathrm{softplus}(-40) = 0test_derivative_tails_keep_relative_accuracy (mutants s03, s09)
3. relu′(0)=1\mathrm{relu}'(0) = 1gradients differ from PyTorch’s exactly at zerostest_relu_derivative_at_zero_is_zero, test_matches_torch_golden (mutant s04)
4. using the tanh approximation for the exact GELU (or the reverse)1.5×10−41.5 \times 10^{-4} off at x=1x = 1; a loaded checkpoint driftstest_hand_example_gelus (mutant s05)
5. forgetting the product rule: silu′=σ\mathrm{silu}' = \sigma, gelutanh⁡′\mathrm{gelu}_{\tanh}' without the tanh⁡′\tanh' termright at 0, wrong everywhere elsetest_hand_example_at_one (mutant s02), test_derivatives_pass_gradcheck (mutant s06)
6. promoting float32 to float64a float64 array where the kernel expects float32test_dtype_and_shape_preserved (mutant s13)
dropping the square in 1−tanh⁡21 - \tanh^2tanh⁡′(1)=0.238\tanh'(1) = 0.238 instead of 0.420test_hand_example_at_one (mutant s01)
a wrong GELU constant (0.0447 for 0.044715)the 6th digit of every GELUtest_hand_example_gelus (mutant s07)
subtracting the log1p term in softplussoftplus(1)=0.687\mathrm{softplus}(1) = 0.687test_hand_example_at_one (mutant s08)
DirectionModuleHow it uses this
BackM02.1gelu_erf and dgelu_erf call erf_series(x / sqrt(2), 160)
BackM01.1derivatives as limits; the frozen gradcheck in the tests is M01.1’s central difference (reading)
BackM00.1exe^x and ln⁡\ln (reading)
ForwardM08.1dual numbers compute the same derivatives by forward-mode autodiff; its tests compare them with yours to 10−1510^{-15}
ForwardL0.2the elementwise ops of tinyllm.F (forward and backward)
ForwardL3.2LSTM gates: sigmoid for input, forget, output; tanh for the cell
ForwardL7.2SwiGLU =silu(xW)⊙xV= \mathrm{silu}(xW) \odot xV
ForwardL9.6tl_silu_mul_f32 in C, checked against your Python

L0.2, L3.2, L7.2, and L9.6 join used_by when they are authored (course/DEVIATIONS.md row B31-02).

Your pieceProduction equivalentWhat it addsWhere to look
sigmoid, silu, gelu_*PyTorch torch.nn.functionalfused CUDA kernels, autocast to bf16, in-place variantsaten/src/ATen/native/Activation.cpp, cuda/ActivationGeluKernel.cu
stable softplusPyTorch F.softplus(beta, threshold)returns xx above a threshold (default 20) instead of computing the log, trading 2×10−92 \times 10^{-9} for speedaten/src/ATen/native/cpu/Activation.cpp
gelu_erf via a serieserf in libm, erff on GPUsminimax rational approximations, a few ulps everywhereglibc sysdeps/ieee754/dbl-64/s_erf.c
the activation zoollama.cpp ggml_silu, ggml_gelulookup tables for f16 GELU, vectorized SiLUggml/src/ggml-cpu/vec.h