Exponents, logs, change of base, units of information
Overview
Section titled “Overview”| Module | M00.1 · build · Python · Pass 2 · 2 to 3 h |
| You build | python/tinyllm/num/units.py: nats_to_bits, bits_to_nats, log_base, bits_per_byte |
| Contract | course/contracts/py/tinyllm/num/units.pyi |
| Tests | course/tests/M00.1/test_units.py (what they check: section 4) |
| Needs | nothing to call. Reading: lang.01 (numpy arrays), L0.0 (the bigram’s negative log-likelihood, used as the worked example) |
| Used by | M11.2 perplexity and the bits-per-byte accumulator calls bits_per_byte · M12.6 decibels and Whisper’s log-mel call log_base(x, 10) · later L1.6 tokenizer metrics · L6.7 the model-zoo table · C1 the capstone report |
| Milestone | MS-P2 (the Pass 2 gate: every math module of the pass passes ol check) |
| Optional depth | OpenStax, Precalculus 2e (free), ch. 6 (exponential and logarithmic functions); MacKay, Information Theory, Inference, and Learning Algorithms (free), ch. 2.4 (the information content of an outcome) |
Key Takeaways
Section titled “Key Takeaways”- A logarithm undoes an exponential: exactly when . It turns products into sums, which is why every loss in the course adds log-probabilities instead of multiplying probabilities (
test_log_laws). - Change of base is one division: . Inverting it, or mixing with , gives numbers that look plausible and are wrong (
test_change_of_base_hand_values,test_log_base_golden). - The base of the log is the unit of information: base gives nats, base 2 gives bits, and one bit is nats (
test_one_bit_is_ln2_nats). - Bits per byte divides a text’s total cost in nats by . It does not depend on how the text was split into tokens, so it is the one number that compares a byte bigram with a BPE transformer (
test_hand_example_bits_per_byte,test_bits_per_byte_is_per_byte). - A model that knows nothing about bytes pays exactly 8 bits per byte; anything trained must do better (
test_uniform_byte_model_is_eight_bits).
How to work this chapter
Section titled “How to work this chapter”ol start M00.1 # stubs python/tinyllm/num/units.py into your repool tests M00.1 # read the test catalog first: rung R0, you write no tests hereol check M00.1 # exit code is the verdictol diff M00.1 # after passing: your code against the reference1. Why now
Section titled “1. Why now”Your bigram model from L0.0 reports nll: the mean negative log-likelihood, in nats per predicted byte, because numpy’s log is the natural logarithm. That number is meaningful only next to another model scored on the same units of text. In Pass 3 your tokenizers (L1.*) split text into multi-byte tokens, and a model over BPE tokens reports nats per token: a smaller vocabulary of longer pieces makes every per-token number larger, even for a better model. Milestones MS-L2 and MS-L3 compare an n-gram model, a neural model, and an LSTM trained on different token streams, and the model zoo of L6.7 ranks every architecture in one table. All of them convert to bits per byte through the function you write here. This module also fixes the vocabulary the whole math track uses: exponent, logarithm, base, nat, bit.
2. Principles
Section titled “2. Principles”| Symbol | Meaning | Type / shape |
|---|---|---|
| a base: a real number with and | float | |
| whole numbers (integers) | int | |
| raised to the power | float | |
| Euler’s number, | math.e | |
| the natural exponential function | ||
| the logarithm of in base : the with | float | |
the natural logarithm, (numpy’s np.log) | float | |
| a probability, | float | |
| the information content (surprisal) of an outcome of probability : | float | |
| the total negative log-likelihood of a text, in nats | float | |
| the length of the text in UTF-8 bytes | int | |
| bits per byte, | float |
2.1 Exponents
Section titled “2.1 Exponents”For a whole number , is multiplied by itself times: . Counting factors gives the three laws of exponents:
The first law forces the meaning of the other exponents. , so . , so . , so , the positive number whose -th power is (this is why : is not a real number). Fractions follow: . Filling in the gaps between the fractions continuously gives for every real , a function that is always positive, increasing when and decreasing when .
One base is special. is the base whose exponential grows, at every , at a rate equal to its own value; M01.1 proves this from the derivative. is the exponential the rest of the course uses (softmax, the sigmoid, every probability a model outputs).
2.2 Logarithms undo exponentials
Section titled “2.2 Logarithms undo exponentials”Because is strictly increasing (or strictly decreasing) and takes every positive value exactly once, it has an inverse: the logarithm.
So , , and in every base. Each law of exponents becomes a law of logarithms (take of both sides):
The first is the reason models are trained on log-probabilities. The probability of a text is a product of one probability per token, thousands of numbers below 1 multiplied together, which underflows to 0 in floating point after a few hundred tokens. Its logarithm is a sum, which stays representable and is what L0.0 already computes.
As approaches 0, falls without bound when : . An outcome a model calls impossible () is infinitely surprising. That is a value, not an error, and your log_base returns it.
2.3 Change of base
Section titled “2.3 Change of base”Write and take the natural log of both sides: (power law). Divide:
Every logarithm is the natural log divided by a constant. Example: , which checks out because . Two consequences: when , so base 1 is not a base (the formula divides by zero); and , because .
In floating point the division rounds, so change of base is accurate to about one unit in the last place, not exact: math.log(2**29) / math.log(2) is 29.000000000000004. When you need an exact power of two (counting bits of a block size), use integer arithmetic or np.log2; the tests compare with float64 tolerances, not equality.
2.4 Units of information
Section titled “2.4 Units of information”An outcome with probability carries units of information, also called its surprisal: a certain outcome () carries none, and rarer outcomes carry more. The base names the unit:
| Base | Unit | One fair coin flip () | One uniform byte () |
|---|---|---|---|
| 2 | bit | 1 bit | 8 bits |
| nat | nats | nats |
By change of base, : to turn nats into bits, divide by , and multiply by to go back. One nat is bits.
A language model assigns each next token a probability, and its total negative log-likelihood on a text is the sum of the surprisals, nats. Dividing by the number of tokens gives nats per token, which depends on the tokenizer. Dividing by the number of bytes does not: the same text has the same bytes however it is cut. So the course compares models by
A model that gives all 256 byte values probability scores exactly ; English text under a good model scores around 1. M11.2 adds perplexity, , and explains why bits per byte and cross-entropy are the same idea.
3. Worked example by hand
Section titled “3. Worked example by hand”The bigram of L0.0, priced in bits per byte. The L0.0 chapter fits an add-one bigram to the text “abbacab” (ids 0, 1, 1, 0, 2, 0, 1 over the alphabet a, b, c) and gets these next-symbol probabilities:
| context | P(a) | P(b) | P(c) |
|---|---|---|---|
| a | 1/6 | 3/6 | 2/6 |
| b | 2/5 | 2/5 | 1/5 |
| c | 2/4 | 1/4 | 1/4 |
The first symbol is never predicted, so the model makes 6 predictions, one per remaining byte:
| step | pair | nats, | bits, | |
|---|---|---|---|---|
| 1 | a to b | 1/2 | 0.693147 | 1.000000 |
| 2 | b to b | 2/5 | 0.916291 | 1.321928 |
| 3 | b to a | 2/5 | 0.916291 | 1.321928 |
| 4 | a to c | 1/3 | 1.098612 | 1.584963 |
| 5 | c to a | 1/2 | 0.693147 | 1.000000 |
| 6 | a to b | 1/2 | 0.693147 | 1.000000 |
| total | 7.228819 |
Then . The bits column gives the same answer directly: . Per byte in nats it is , the nll that L0.0’s test test_hand_example_nll checks, and again. Against the 8 bits per byte of knowing nothing, the bigram has learned something about this (tiny) text. These numbers are test_hand_example_bits_per_byte.
Change of base. (section 2.3), , and because : test_change_of_base_hand_values.
4. The interface
Section titled “4. The interface”LN2 = math.log(2.0)def nats_to_bits(x: float) -> float: ... # x / ln 2def bits_to_nats(x: float) -> float: ... # x * ln 2def log_base(x: ArrayLike, b: float) -> NDArray: ... # ln x / ln b, float64, x's shapedef bits_per_byte(nll_nats_sum: float, n_bytes: int) -> float: ...log_base returns at (for ) without a warning, and raises ValueError for a base that is not positive, finite, and different from 1, and for negative or NaN inputs. bits_per_byte raises ValueError unless n_bytes is a positive integer and the sum is a non-negative number. The full rules are in the contract.
What the tests check
Section titled “What the tests check”| Test | KIND | Checks | Why it matters downstream |
|---|---|---|---|
test_hand_example_bits_per_byte | unit, smoke | section 3: “abbacab” costs 1.204803 bits per byte | you and the tests agree on the definition |
test_uniform_byte_model_is_eight_bits | unit | nats over bytes is 8 bits per byte | the baseline every model in the zoo must beat |
test_one_bit_is_ln2_nats | unit, smoke | nats is one bit, 1 nat is 1.4427 bits, and back | the factor every later conversion uses |
test_nats_to_bits_golden | golden | six values against mpmath at 60 digits | full float64 precision, not an approximation |
test_nats_bits_roundtrip | property | bits_to_nats(nats_to_bits(x)) == x from to | the two conversions are inverses |
test_change_of_base_hand_values | unit | , , | the section 2.3 derivation |
test_log_base_golden | golden | 66 pairs, from to , bases 0.5 to 256 | float64 all the way through |
test_log_laws | property | product, power, and reciprocal-base laws on random inputs | the identities later chapters rewrite losses with |
test_log_base_keeps_shape_and_dtype | boundary | arrays keep their shape, results are float64, scalars stay scalars | L1.6 and M11.2 pass whole arrays |
test_log_of_zero_is_infinite | boundary | , , no warning | a zero-probability token is a value, not a crash |
test_bad_base_rejected | boundary | bases 1, 0, , , NaN raise ValueError | base 1 silently divides by zero |
test_negative_or_nan_x_rejected | boundary | negative or NaN raises ValueError | a NaN surprisal poisons every later sum |
test_bits_per_byte_golden | golden | four totals against mpmath | the conversion MS-L2 and MS-L3 compare models with |
test_bits_per_byte_is_per_byte | property | doubling text and cost keeps bpb; doubling cost alone doubles it | the byte count really divides |
test_bits_per_byte_rejects_bad_input | boundary | zero, negative, or fractional byte counts and negative or NaN totals raise ValueError | an empty text has no bits per byte |
5. Pitfalls
Section titled “5. Pitfalls”| Pitfall | Symptom | Caught by |
|---|---|---|
| 1. multiplying by to get bits (or dividing to get nats) | every number off by a factor of ; a 1.2 bpb model reports 0.58 | test_one_bit_is_ln2_nats (mutants s01, s08); test_hand_example_bits_per_byte (mutant s04, bpb left in nats) |
| 2. inverting change of base () or mixing logs () | plausible numbers, wrong by a constant factor or a reciprocal | test_change_of_base_hand_values (mutants s02, s03) |
| 3. accepting base 1 | : the division returns or NaN without an error | test_bad_base_rejected (mutant s05) |
| 4. treating as an error | evaluating a model that assigns probability 0 to some byte crashes instead of reporting | test_log_of_zero_is_infinite (mutant s11) |
| 5. dividing by tokens, or not dividing at all | a BPE model looks worse than a byte model that is worse; bpb grows with text length | test_bits_per_byte_is_per_byte (mutant s12) |
| 6. computing in float32 | underflows to 0 and the log becomes ; 7 digits instead of 16 | test_log_base_golden (mutant s10) |
6. Where it’s used next
Section titled “6. Where it’s used next”| Direction | Module | How it uses this |
|---|---|---|
| Back | L0.0 | its nll (nats per predicted byte) is section 3’s worked example (reading) |
| Back | lang.01 | numpy arrays and dtypes (reading) |
| Forward | M11.2 | the perplexity and NLL accumulator reports bpb through bits_per_byte |
| Forward | M12.6 | power_to_db and log_mel_whisper take base-10 logs with log_base |
| Forward | L1.6 | tokenizer metrics: bytes per token and the bits-per-byte of each tokenizer’s model |
| Forward | L6.7 | the model-zoo table ranks every architecture by bits per byte |
| Forward | C1 | the capstone report’s headline number |
| Forward | M00.2, M00.3, M00.4 | use exponentials and logs in their derivations (reading) |
Going further
Section titled “Going further”| Your piece | Production equivalent | What it adds | Where to look |
|---|---|---|---|
bits_per_byte | EleutherAI lm-evaluation-harness, perplexity tasks | reports word_perplexity, byte_perplexity, and bits_per_byte for every model on the same text | lm_eval/api/metrics.py (bits_per_byte) |
| bits per byte as the cross-tokenizer unit | The Pile (Gao et al., 2020) | reports BPB so that models with different tokenizers land on one scale | the paper’s evaluation of GPT-2 and GPT-3 on Pile components |
log_base | numpy log2, log10, log1p, logaddexp | exact-for-powers-of-two log2, accuracy near 1 (log1p, used in M00.3), and stable sums of exponentials (M09.2) | numpy reference, “Exponents and logarithms” |