Skip to content

FP8 E4M3/E5M2, MXFP4/MXFP8 with E8M0 scales

ModuleM09.4 · build · Python · Pass 6
You buildpython/tinyllm/num/lowp.py: FP8 E4M3/E5M2 conversion, FP4 E2M1 values, block scales, and MX quantize/dequantize
Contractcourse/contracts/py/tinyllm/num/lowp.pyi; expected bytes are in course/fixtures/M09.4/lowp_golden.npz
Testscourse/tests/M09.4/test_lowp.py: exact code tables, tie behavior, saturation, signed zero, block layouts, and round trips
NeedsM09.1 for IEEE-754 bit layouts; S-M09a for rounding by hand
Used byL8.5 quantizes weights; L10.1 converts Rust model weights; craft.13 stores fp8 KV data
MilestoneMS-L8

Float32 weights and activations move four bytes per value. FP8 stores each value in one byte; MX formats group values into blocks and share a scale, while FP4 stores four bits per value. These formats reduce memory traffic, but every conversion has edge cases: ties, tiny magnitudes, signed zero, infinities, NaNs, and values beyond the finite range. This module writes the conversions from the format definitions and tests the resulting bit patterns against checked-in vectors.

An optional standalone C conversion exercise lives in M09.7. It consumes the same fixture files and does not load this Python code.

An IEEE-like binary format uses a sign bit, an exponent field, and a fraction field. For a normalized value with fraction width m and bias b, the decoded value is

[ (-1)^s (1 + f/2^m) 2^{e-b}. ]

Subnormals omit the implicit leading one and use the minimum exponent. E4M3 uses four exponent bits and three fraction bits, giving more precision near one than E5M2. E5M2 uses two fraction bits and therefore covers a wider range. The project defines saturation and special-value behavior explicitly, so callers do not depend on a library default.

Rounding is round-to-nearest, ties-to-even. A tie at 2.5 selects 2; a tie at 3.5 selects 4. FP8 encode/decode uses exact powers of two to map between values and codes. Microscaling formats apply one E8M0 power-of-two scale to a fixed-size block, then encode each scaled value in FP8 or E2M1.

In E4M3, 1.0625 lies halfway between adjacent representable values 1.0 and 1.125. Their final significand bits are even and odd respectively, so ties-to-even encodes 1.0. A value just above the midpoint encodes 1.125. The test test_ties_go_to_even checks both sides of these boundaries as well as the exact midpoint.

For an MX block, first find the largest absolute element. Choose the shared power-of-two scale that moves that maximum into the element format’s top range, then divide each value by the scale and round once into the element format. Decoding multiplies each value by the same scale. The fixture’s block example lets the tests check the scale code and every packed element independently.

TestWhy it matters
test_hand_example and test_hand_example_mxPin scalar and block calculations to hand-worked values.
test_decode_every_code_golden and test_encode_goldenCheck the full 8-bit code space against stable expected values.
test_saturation_and_specialsDefines behavior for overflow, infinities, and NaNs.
test_tiny_values_and_signed_zeroProtects underflow and sign preservation at zero.
test_encode_inverts_decodeEvery finite code remains stable through decode then encode.
test_rounding_is_nearest_and_monotone and test_ties_go_to_evenCatches incorrect midpoint handling and non-monotone encoders.
test_scaled_quantization_roundtrip and test_quantize_rounds_onceEnsures scaling does not introduce an extra rounding step.
test_fp4_values_and_ties, test_e8m0_scale_code, test_mx_golden, test_mx_block_law, test_mx_shapes_and_errorsCheck element codes, scale choices, block layout, and input validation.

Start with ol start M09.4, inspect the catalog with ol tests M09.4, then run ol check M09.4. The file fixture makes results reproducible without optional dependencies at runtime.

PitfallSymptomCaught by
Rounding half away from zeroExact midpoints choose the wrong code.test_ties_go_to_even
Treating subnormals like normalized valuesTiny values jump to zero or the minimum normal.test_tiny_values_and_signed_zero
Producing infinity on finite overflowCodes disagree with the course’s saturation rule.test_saturation_and_specials
Applying the block scale after roundingQuantized values differ by one code near boundaries.test_quantize_rounds_once
Choosing scale from an average instead of the maximum magnitudeA large block member clips unexpectedly.test_mx_block_law
Accepting malformed block shapesValues are silently dropped or grouped incorrectly.test_mx_shapes_and_errors
DirectionModuleHow it uses this
BackM09.1Provides the float bit layouts and exact powers-of-two reasoning.
ForwardL8.5Quantizes model weights for lower-bandwidth inference.
ForwardL10.1Converts quantized model weights in the Rust engine.
Forwardcraft.13Stores fp8 KV values and relies on the defined saturation and decoding rules.
Optional parallelM09.7Implements selected conversions in standalone C against the shared golden files.
Your pieceProduction equivalentWhat it addsWhere to look
FP8 E4M3/E5M2accelerator matrix formatstrades exponent range against significand precisionIEEE 754 and vendor format guides
microscalingblock-scaled MX formatscompact values share a power-of-two scale per blockcourse/fixtures/M09.4/lowp_golden.npz