IEEE 754 anatomy, ulp, round to nearest even, and bf16 and fp16 emulation
Overview
Section titled “Overview”| Module | M09.1 · build · Python · Pass 2 · 2 to 3 h |
| You build | python/tinyllm/num/fp.py: decompose_f32, compose_f32, ulp, f32_to_bf16_bits, bf16_bits_to_f32, round_to_bf16, round_to_fp16 |
| Contract | course/contracts/py/tinyllm/num/fp.pyi |
| Tests | course/tests/M09.1/test_fp.py (what they check: section 4) |
| Needs | nothing to build first. Reading: M00.1 exponents and logarithms (powers of two, ) |
| Used by | L0.6 reads and writes BF16 safetensors through f32_to_bf16_bits and bf16_bits_to_f32 · later L7.9 loads bf16 checkpoints through bf16_bits_to_f32, L11.1 rounds weights and activations with round_to_bf16, M09.4 builds fp8 and MX formats on the same bit arithmetic |
| Milestone | MS-P2 (the foundations gate) |
| Optional depth | Goldberg, “What Every Computer Scientist Should Know About Floating-Point Arithmetic” (ACM Computing Surveys, 1991), sections 1 and 2; Muller et al., Handbook of Floating-Point Arithmetic (2nd ed.), ch. 2 and 3; IEEE Std 754-2019, section 3 |
Key Takeaways
Section titled “Key Takeaways”- A float32 is three integers: sign , biased exponent , and mantissa ; a normal value is , and and hold the zeros, subnormals, infinities, and NaNs (
test_decompose_hand_example,test_decompose_every_class). - The gap between neighbouring floats, one ulp, is : it doubles at every power of two and stops shrinking in the subnormal range (
test_ulp_matches_numpy_spacing). - bfloat16 is the top half of a float32, so rounding to it is integer arithmetic on the bits: add plus the lowest kept bit, then shift. Ties go to the even neighbour, and a carry walks into the exponent exactly as it should (
test_bf16_matches_exact_rational_golden). - The trick is wrong for NaN, which can turn into infinity; every NaN must map to the quiet code
0x7FC0(test_bf16_nan_stays_nan). - Codes are decoded with a bit view, never a value cast:
astypeturns the code0x3F80into instead of (test_bf16_all_codes_roundtrip).
How to work this chapter
Section titled “How to work this chapter”ol start M09.1 # stubs python/tinyllm/num/fp.py into your repool tests M09.1 # read the test catalog first: rung R0, you write no tests hereol check M09.1 # exit code is the verdictol diff M09.1 # after passing: your code against the reference1. Why now
Section titled “1. Why now”Everything your tracer stores is float32: L0.0 writes the bigram’s weights as F32 in safetensors, and M03.1 multiplies them in C. Pass 2 starts to lean on what float32 actually is. The stable softmax of M09.2 exists because float32 overflows at , and that number comes from the exponent field. Later passes go below 32 bits. The SmolLM2 checkpoint L7.9 loads is stored as BF16, two bytes per weight, and the obvious loader, np.frombuffer(raw, np.uint16).astype(np.float32), gives you a model whose weight 1.0 reads as 16256.0. Its first forward pass returns NaN and nothing crashes. Mixed-precision training (L11.1) rounds every weight to bf16, and the obvious rounding (keep the top 16 bits) is biased: it always rounds toward zero, and over thousands of steps the weights drift. This module takes the float32 format apart bit by bit and rebuilds the two 16-bit formats from it, so the loader and the rounding are right the first time.
2. Principles
Section titled “2. Principles”| Symbol | Meaning | Type |
|---|---|---|
| sign bit: 0 for positive, 1 for negative | int, 0 or 1 | |
| biased exponent field: 8 bits in float32 | int, 0 to 255 | |
| mantissa (fraction) field: 23 bits in float32 | int, 0 to | |
| the unbiased exponent of a normal float32 | int, -126 to 127 | |
| precision: significand bits including the implicit leading 1 | 24 (f32), 11 (f16), 8 (bf16) | |
| smallest normal exponent | -126 (f32, bf16), -14 (f16) | |
| the gap from to the next larger magnitude in the format | float64 | |
| unit roundoff: the largest relative error of one rounding | float | |
| rounded to the nearest representable value, ties to even | a value of the format | |
the 32 bits of a float32 read as an unsigned integer (a uint32 view) | int |
2.1 The binary32 encoding
Section titled “2.1 The binary32 encoding”A float32 occupies 32 bits: bit 31 is the sign , bits 30 to 23 the biased exponent , bits 22 to 0 the mantissa . Reading the same 32 bits as an unsigned integer gives , so the fields are b >> 31, (b >> 23) & 0xFF, and b & 0x7FFFFF. The value depends on :
| class | value | |
|---|---|---|
| 1 to 254 | normal | |
| 0 | zero () or subnormal () | |
| 255 | infinity () or NaN () | , or not a number |
A normal number has significant bits: the 23 stored ones plus a leading 1 that is implied, not stored. The bias 127 lets the exponent field be an unsigned integer, which has a useful consequence: for non-negative floats, ordering the bit patterns as integers orders the values. The largest finite float32 is , , which is . The smallest normal is . Below it, drops the implicit 1 and keeps the exponent at , so the subnormals fill the gap down to in equal steps. There are two zeros, and , which compare equal. A NaN is any pattern with and ; there are about of them.
2.2 Spacing and the ulp
Section titled “2.2 Spacing and the ulp”Between and (one binade) a format with precision has equally spaced values, so their spacing is
The spacing doubles at every power of two and is constant across the subnormal range, which is why is clamped at . At : in float32, in float16, in bfloat16. At the gap to the next value is the smallest subnormal, . Infinity and NaN have no next value, so their ulp is NaN.
Rounding to nearest moves a value by at most half an ulp, and within a binade the ulp is at most , so one rounding has a relative error of at most : with , for any in the normal range. That one fact is the model of every error bound in this track (M09.2, M09.3).
Computing needs care. np.log2(x) returns a rounded float, and for a hair below the true value rounds up to exactly , giving the next binade’s ulp. The exponent is already stored exactly: np.frexp(x) returns with and , so with no rounding at all.
2.3 Round to nearest even, on the bits, for bfloat16
Section titled “2.3 Round to nearest even, on the bits, for bfloat16”bfloat16 has the float32 layout cut in half: 1 sign bit, the same 8 exponent bits with the same bias, and 7 mantissa bits (). Every bf16 code is the top 16 bits of a float32, and decoding is a shift: bf16_bits_to_f32(c) is the float32 whose bits are c << 16. Because the range is float32’s, nothing a float32 model produces overflows on the way down; only precision is lost.
Encoding has to round. IEEE 754’s default is round to nearest, ties to even: take the nearest representable value, and when the value is exactly halfway between two, take the one whose last kept bit is 0. Ties to even is unbiased: half of all ties round up and half round down, where “ties away from zero” pushes every tie the same way.
Split the float32 bits into the half we keep and the half we drop: and . The halfway point of the dropped half is . The rule:
| dropped half | result |
|---|---|
below 0x8000 | (round down) |
above 0x8000 | (round up) |
exactly 0x8000 | if is even, if odd |
All three rows are one integer expression:
When , adding at most keeps the low half below , so no carry reaches . When , adding already carries. When , the sum is plus the parity bit, so it carries exactly when is odd.
The carry needs no special case. If the 7 kept mantissa bits are all ones, adding 1 to clears them and increments the exponent field: the next code is the first value of the next binade, which is the correct rounded result. At the top, the largest finite float32 0x7F7FFFFF rounds to 0x7F80, which is : exactly IEEE’s rule, since is nearer to than to the largest bf16 value. Infinity itself has and stays infinity. Work in uint64 (or check for it) so that the addition cannot wrap at 0xFFFFFFFF.
NaN breaks it. A NaN may have its only set mantissa bits in the dropped half: 0x7F800001 has and , and so does every NaN whose payload is small. Truncated, it becomes 0x7F80, infinity. Rounded, 0x7F800001 + 0x7FFF = 0x7F808000, still infinity. A NaN loss that turns into infinity no longer trips np.isnan. So NaNs are mapped first, to the single quiet NaN code 0x7FC0, matching the C conversions in tinyllm/numerics.h.
2.4 binary16: a different exponent range
Section titled “2.4 binary16: a different exponent range”float16 (IEEE binary16) has 1 sign bit, 5 exponent bits with bias 15, and 10 mantissa bits (, ). Its range is tiny: the largest finite value is , and the smallest subnormal is . So a float32 can overflow it, land in its subnormal range, or underflow it entirely, and the exponent field has to be re-biased. The top-half trick does not apply.
Write a finite with an integer significand ( and for a normal float32, and for a subnormal). binary16 spaces its values apart with . Dividing by that spacing gives with : shift right by and round the remainder to nearest even, exactly as for bf16 but with a variable cut. Then
For a normal result () this equals the field form ; for a subnormal (, ) it is just ; and a that rounded up to carries into the exponent on its own, as in bf16. Any code at or above 0x7C00 is infinity (overflow), and NaN gets 0x7E00. numpy’s float32.astype(float16) implements the same IEEE rounding, which gives the tests an independent oracle.
2.5 Bit views versus value casts
Section titled “2.5 Bit views versus value casts”np.asarray(x, np.float32).view(np.uint32) reinterprets the same bytes as integers and changes no bit; .astype(np.uint32) converts the value (1.0 becomes 1). Every function here moves between the two worlds with views. Decoding bf16 is the case people get wrong: the codes arrive as uint16, and the value of code is the float32 with bits , not the number .
3. Worked example by hand
Section titled “3. Worked example by hand”Anatomy of . . So , , and the mantissa is the fraction after the leading 1, , padded to 23 bits: . All 32 bits: 1 10000001 10010000000000000000000 = 0xC0C80000. decompose_f32(-6.25) == (1, 129, 0x480000). Check: .
ulp. lies in the binade , , so in float32, in float16, and in bfloat16. At 1.0 the three are , , (test_ulp_hand_values).
to bfloat16. has no finite binary expansion. float32(0.1) has bits 0x3DCCCCCD and the value . The kept half is , the dropped half :
| step | value |
|---|---|
vs 0x8000 | 0xCCCD > 0x8000: round up |
0 (0x3DCC is even) | |
0x3DCCCCCD + 0x7FFF = 0x3DCD4CCC | |
0x3DCD |
Decode 0x3DCD = 0 01111011 1001101: , , mantissa , value . The error, about , is below half the bf16 ulp at , .
Ties. has bits 0x3F808000: (that is 1.0), , exactly halfway to . is even, so the sum 0x3F808000 + 0x7FFF = 0x3F80FFFF does not carry: the result is 0x3F80, . Next, has bits 0x3F818000: halfway between 0x3F81 (odd) and 0x3F82 (even). Now the parity bit is 1, 0x3F818000 + 0x7FFF + 1 = 0x3F820000, and the result is 0x3F82 . One tie went down and one went up, both to the even code (test_bf16_hand_examples).
to float16. , so the result is normal and the cut is at . The float32 mantissa 0x4CCCCD is ; the top 10 bits are and the 13 dropped bits are , below the halfway point 0x1000, so it rounds down. The code is , the value . float16 lands closer to than bf16 did: 3 more mantissa bits.
Edges of float16. is below the midpoint between and the would-be next value , so it rounds to ; is the tie, and ties to even picks , which does not exist, so the result is . At the bottom, is half the smallest subnormal : a tie between and , and the even choice is (test_fp16_edges).
4. The interface
Section titled “4. The interface”def decompose_f32(x: float) -> tuple[int, int, int]: ... # (sign, biased exponent, mantissa)def compose_f32(sign: int, exponent: int, mantissa: int) -> float: ...def ulp(x: ArrayLike, dtype: Literal["f32", "f16", "bf16"]) -> NDArray: ... # float64def f32_to_bf16_bits(x: ArrayLike) -> NDArray: ... # uint16, RNE, NaN -> 0x7FC0def bf16_bits_to_f32(u16: ArrayLike) -> NDArray: ... # float32, exactdef round_to_bf16(x: ArrayLike) -> NDArray: ... # float32 in, float32 outdef round_to_fp16(x: ArrayLike) -> NDArray: ... # float32, same bits as numpy's castArray functions accept anything numpy converts, convert it to float32 first, and keep the input’s shape (a scalar gives a 0-d array). L7.9 reads a BF16 tensor as bf16_bits_to_f32(np.frombuffer(raw, "<u2").reshape(shape)); L11.1 keeps float32 master weights and computes with round_to_bf16(w).
What the tests check
Section titled “What the tests check”| Test | KIND | Checks | Why it matters downstream |
|---|---|---|---|
test_decompose_hand_example | unit, smoke | section 3’s , plus 1.0 and | you and the test agree on the fields |
test_decompose_every_class | boundary | , the smallest subnormal and normal, the largest finite, , NaN | a checkpoint can hold any of them |
test_compose_roundtrip | property | compose inverts decompose on 3000 floats; out-of-range fields raise | the three fields describe the number completely |
test_ulp_hand_values | unit, smoke | in each format, , NaN for and NaN, unknown dtype | the relative precision of each format |
test_ulp_matches_numpy_spacing | differential | float32 and float16 ulps equal np.spacing, subnormals included | the spacing M09.3 builds tolerances from |
test_ulp_just_below_a_power_of_two | boundary | nextafter(1024, 0) and nextafter(2**100, 0) stay in the lower binade | np.log2 rounds up there |
test_bf16_hand_examples | unit, smoke | section 3: 0.1 -> 0x3DCD, the two ties | the rounding rule itself |
test_bf16_matches_exact_rational_golden | golden | 10 156 float32 patterns against rounding done with exact fractions | every binade, tie, and overflow |
test_bf16_nan_stays_nan | boundary | four NaNs, including 0x7F800001, give 0x7FC0 | a NaN loss stays visible in L11.1 |
test_bf16_all_codes_roundtrip | property | all 65536 codes decode exactly and encode back | L7.9 loads every weight exactly |
test_rounding_is_idempotent_and_monotone | property | rounding twice changes nothing; sorted inputs stay sorted | clipping and comparisons after rounding |
test_fp16_matches_numpy_bitwise | differential, golden | the golden inputs and random patterns equal numpy’s cast bit for bit | an independent IEEE implementation agrees |
test_fp16_edges | boundary | , , , subnormals, | why fp16 training needs loss scaling and bf16 does not |
test_shapes_and_scalars | boundary | matrices keep their shape, a scalar gives a 0-d array, the sign survives | weight matrices go through unchanged |
The golden file course/fixtures/M09.1/round_f32.npz was written by course/oracle/M09.1/lowp_golden.py, which rounds each float32 value as an exact fractions.Fraction and checks its float16 column against numpy.
5. Pitfalls
Section titled “5. Pitfalls”| Pitfall | Symptom | Caught by |
|---|---|---|
| keeping the top 16 bits (truncation) | every value rounds toward zero; 0.1 gives 0x3DCC | test_bf16_hand_examples (mutant s01) |
adding 0x8000 (ties away from zero) | 1 + 2**-8 rounds to 0x3F81, a biased tie | test_bf16_hand_examples (mutant s02) |
| no special case for NaN | 0x7F800001 becomes 0x7F80, infinity | test_bf16_nan_stays_nan (mutant s03) |
through np.log2 | the ulp just below a power of two is twice too large | test_ulp_just_below_a_power_of_two (mutant s04) |
| no clamp at | subnormal ulps keep shrinking past the real spacing | test_ulp_matches_numpy_spacing (mutant s05) |
| flushing float16 subnormals to zero | values below vanish | test_fp16_edges (mutant s06) |
no clamp at 0x7C00 | values past 65504 wrap into NaN codes | test_fp16_edges (mutant s07) |
| float16 ties rounded up | half of all ties off by one ulp | test_fp16_matches_numpy_bitwise (mutant s08) |
astype instead of a bit view | code 0x3F80 reads as | test_bf16_all_codes_roundtrip (mutant s09) |
exponent not masked with 0xFF | the sign bit leaks into the exponent of negative numbers | test_decompose_hand_example (mutant s10) |
6. Where it’s used next
Section titled “6. Where it’s used next”| Direction | Module | How it uses this |
|---|---|---|
| Back | M00.1 | powers of two and (reading) |
| Forward | M09.2 | where float32 overflows and underflows, and the unit roundoff behind its summation bounds (reading) |
| Forward | L0.6 | safetensors for every dtype: BF16 tensors are encoded and decoded with this module’s bit converters |
| Forward | M09.4 | fp8 E4M3 and E5M2 and MX block formats: the same bit-level rounding with fewer bits, in Python and C |
| Forward | L7.9 | loads the SmolLM2 BF16 safetensors through bf16_bits_to_f32 |
| Forward | L11.1 | mixed precision: float32 master weights, round_to_bf16 for the compute copy |
If you skip this module, ol check L0.6 stops with needs M09.1: build it, or pass --ref-deps.
Going further
Section titled “Going further”| Your piece | Production equivalent | What it adds | Where to look |
|---|---|---|---|
f32_to_bf16_bits | PyTorch c10::BFloat16 | the same add-0x7FFF-plus-parity rounding, in C++ and in CUDA for every kernel | c10/util/BFloat16.h, round_to_nearest_even |
round_to_fp16 | the F16C instructions (vcvtps2ph), CUDA __float2half_rn | the conversion in one hardware instruction, with the rounding mode as an operand | Intel SDM, VCVTPS2PH; CUDA Math API, half precision intrinsics |
round_to_bf16 in training | mixed-precision training with loss scaling | float32 master weights, low-precision compute, dynamic loss scaling for fp16’s narrow range | Micikevicius et al., “Mixed Precision Training” (ICLR 2018) |
| bf16 and fp16 | OCP FP8 and MX formats | 8-bit and block-scaled formats built on the same encoding rules | OCP Microscaling Formats (MX) v1.0 spec; M09.4 |