Sampling, Nyquist, windows, and the STFT
Overview
Section titled “Overview”| Module | M12.5 · build · Python · Pass 12 · 3 h |
| You build | python/tinyllm/sig/stft.py: hann (periodic and symmetric), frame_count, stft (torch.stft semantics: center reflect padding, explicit window), istft (weighted overlap-add), and power |
| Contract | course/contracts/py/tinyllm/sig/stft.pyi |
| Tests | course/tests/M12.5/test_stft.py (what they check: section 4), golden values from course/oracle/M12.5/stft_torch.py (torch.stft, torch.hann_window) in course/fixtures/M12.5/stft_torch.npz |
| Needs | M12.4 rfft and irfft (or --ref-deps) |
| Used by | later L14.1 the Whisper frontend (power spectrogram of 400-sample frames at hop 160), with B14’s L14 group |
| Milestone | MS-P12 (the multimodal gate) |
| Optional depth | Oppenheim and Schafer, Discrete-Time Signal Processing, ch. 4 and 10; Harris, “On the use of windows for harmonic analysis with the discrete Fourier transform” (Proc. IEEE 1978); Smith, Spectral Audio Signal Processing (free), the overlap-add chapters |
Key Takeaways
Section titled “Key Takeaways”- Sampling at rate can only hold frequencies below ; a tone above folds to (
test_tone_above_nyquist_folds). - A window tapers each frame so its edges do not smear energy across every bin; the periodic Hann window is three cosines, so a tone on a bin lands in exactly three bins (
test_power_and_periodic_hann_leakage,test_hann_matches_torch). center=Truepads reflected samples on each side, so there are frames: 3001 for 30 s of Whisper audio, which then drops the last (test_frame_count_table,test_hand_example_stft).- When the squared window overlap-adds to a constant (periodic Hann at hop ), weighted overlap-add inverts the STFT exactly (
test_istft_reconstructs_under_cola).
How to work this chapter
Section titled “How to work this chapter”ol start M12.5 # stubs python/tinyllm/sig/stft.py into your repool tests M12.5 # read the test catalog first: rung R0, you write no tests hereol check M12.5 # exit code is the verdictol check M12.5 --ref-deps # only if your M12.4 is not passing yetol diff M12.5 # after passing: your code against the reference1. Why now
Section titled “1. Why now”M12.4 transforms one block of samples. Speech is not one block: its spectrum changes every few tens of milliseconds, and Whisper’s encoder wants a picture of that change, 100 spectra per second. L14.1 computes exactly torch.stft(audio, 400, 160, window=hann_window(400), center=True, pad_mode='reflect') and squares the magnitudes; Whisper was trained on those numbers. Using a symmetric window, zero padding, or a frame count off by one gives features that look like a spectrogram and shift the transcription. The same module explains why M12.3 must filter before it downsamples: aliasing is a sampling fact, and here you can see it in one bin.
2. Principles
Section titled “2. Principles”| Symbol | Meaning | Type / shape |
|---|---|---|
| sample rate | Hz | |
| Nyquist frequency | Hz | |
| frame length (even) | int | |
| samples between frame starts | int | |
| analysis window | float64[n_fft] | |
| the signal after center padding | float64[n + 2 (n_fft // 2)] | |
| STFT bin of frame | complex128[n_fft//2 + 1, T] | |
| frequency of bin | Hz |
2.1 Sampling and aliasing
Section titled “2.1 Sampling and aliasing”Sampling at gives , and since has period in its argument, the frequencies and give identical samples, as do and . So only is distinguishable: the Nyquist limit. A 5 kHz tone sampled at 8 kHz produces exactly the samples of a 3 kHz tone, and no later processing can tell them apart. Anything above must be removed before sampling or resampling, which is why M12.3 low-passes (antialiasing) and why Whisper’s 16 kHz input cannot represent anything above 8 kHz.
2.2 Windows and leakage
Section titled “2.2 Windows and leakage”The DFT of a frame treats the frame as one period of a periodic signal. A sinusoid that does not complete a whole number of periods in the frame has a jump where the period wraps, and the jump spreads energy into every bin: leakage. A window that falls to 0 at the ends removes the jump. The Hann window is
with (periodic) or (symmetric). The periodic version is one period of a cosine of the DFT size, so its own DFT is three nonzero bins, at : multiplying a tone by it convolves the tone’s single bin with those three. The symmetric version is the classic filter-design window and leaks a little everywhere. Spectral analysis uses periodic; torch.hann_window defaults to it, and so does Whisper.
2.3 The STFT and its inverse
Section titled “2.3 The STFT and its inverse”Frame starts at sample of the padded signal; its spectrum is the rfft of the windowed frame: To invert, take the irfft of each column, multiply by the window again, and add the frames back at their offsets (overlap-add). Sample then holds , so dividing by the overlap-added squared window recovers wherever that sum is nonzero (the NOLA condition). For the periodic Hann window at the sum is the constant (constant overlap-add, COLA), so the reconstruction is exact to rounding. Removing the padding at the start returns the original signal.
2.4 Center padding and frame counts
Section titled “2.4 Center padding and frame counts”With center=True the signal is padded by samples on each side, so frame is centered on sample of the original. Reflect padding mirrors without repeating the edge sample, , which keeps the signal continuous at the boundary where zeros would create a click. The frame count is
and without padding . Whisper pads every clip to 30 s, , gets frames, and drops the last one so the encoder sees 3000. Reflect padding needs more than samples; torch.stft raises otherwise, and so does stft.
2.5 Power
Section titled “2.5 Power”The spectrogram Whisper uses is the power , computed without a square root (no need to take a root you square again). M12.6 pools it into mel bands and takes a log.
3. Worked example by hand
Section titled “3. Worked example by hand”, , , periodic Hann (the symmetric one would be ). Reflect padding by 2 gives , and frames:
| frame | windowed | ||||
|---|---|---|---|---|---|
| 0 | 3 | ||||
| 1 | 6 | 0 | |||
| 2 | 9 | 1 |
So , bins down, frames across. This is test_hand_example_stft.
4. The interface
Section titled “4. The interface”def hann(n, periodic=True) -> NDArray: ...def frame_count(n_samples, n_fft, hop, center=True) -> int: ...def stft(x, n_fft, hop, window, center=True, pad_mode='reflect') -> NDArray: ... # [n_fft//2 + 1, T]def istft(S, n_fft, hop, window, length) -> NDArray: ...def power(S) -> NDArray: ...What the tests check
Section titled “What the tests check”| Test | KIND | Checks | Why it matters downstream |
|---|---|---|---|
test_hand_example_stft | unit, smoke | section 3’s windows and spectrogram | padding, windowing, and orientation |
test_frame_count_table | unit, smoke | Whisper’s 3001, centered and uncentered counts, bad arguments | L14.1’s 3000 frames |
test_hann_matches_torch | golden | torch.hann_window for , both flags | Whisper’s window |
test_stft_matches_torch | golden | Whisper-sized, small, zero-padded symmetric, and short-hop cases within | L14.1 parity |
test_istft_reconstructs_under_cola | property | round trip error below at hop | nothing is lost in the frames |
test_tone_above_nyquist_folds | property | 5 kHz at 8 kHz peaks at the 3 kHz bin, 96 | aliasing is a sampling fact |
test_power_and_periodic_hann_leakage | property | power is ; a bin-centered tone fills bins 19 to 21 in ratio | why periodic Hann |
test_rejects_bad_frames | boundary | wrong window length, too short to reflect, hop 0, bad pad mode, short uncentered signal, wrong spectrogram height; input untouched | misconfigured frontends fail early |
5. Pitfalls
Section titled “5. Pitfalls”| Pitfall | Symptom | Caught by |
|---|---|---|
| 1. the symmetric Hann window | leakage into every bin; features drift from what Whisper was trained on | test_hann_matches_torch, test_power_and_periodic_hann_leakage (mutant s01) |
| 2. zero padding instead of reflect | a click in the first and last frames | test_hand_example_stft, test_stft_matches_torch (mutant s02) |
| 3. the centered frame count without the | 3000 frames before Whisper drops one: 2999 | test_frame_count_table (mutant s03) |
| 4. forgetting to window the frames | leakage everywhere; the round trip breaks | test_hand_example_stft, test_istft_reconstructs_under_cola (mutant s04) |
| 5. normalizing overlap-add by , not | the reconstruction scaled by about 0.75 | test_istft_reconstructs_under_cola (mutant s05) |
| 6. returning frames by bins | every consumer transposes or breaks | test_stft_matches_torch, test_tone_above_nyquist_folds (mutant s06) |
| 7. power as the magnitude | log-mel features halved | test_power_and_periodic_hann_leakage (mutant s07) |
| reflect-padding a signal of exactly samples | numpy repeats the reflection where torch raises | test_rejects_bad_frames (mutant m01) |
| accepting hop 0 | a division by zero instead of a clear error | test_rejects_bad_frames (mutant m02) |
| accepting a signal one sample short of a frame | zero frames instead of an error | test_frame_count_table (mutant m03) |
6. Where it’s used next
Section titled “6. Where it’s used next”| Direction | Module | How it uses this |
|---|---|---|
| Back | M12.4 | stft runs rfft on every frame; istft runs irfft |
| Forward | M12.6 | its filterbank pools this power spectrogram into mel bands (reading) |
| Forward | L14.1 | the Whisper frontend: stft(x, 400, 160, hann(400)), drop the last frame, power (with B14’s L14 group) |
| Forward | S-M12 | Nyquist, aliasing, windows, COLA, and frame counts by hand |
Going further
Section titled “Going further”| Your piece | Production equivalent | What it adds | Where to look |
|---|---|---|---|
stft | torch.stft | batched frames as a strided view, cuFFT on GPU, onesided and normalized flags | aten/src/ATen/native/SpectralOps.cpp |
istft | librosa.istft | window-sum caching, Griffin-Lim phase recovery on top | librosa/core/spectrum.py |
hann | scipy.signal.get_window | dozens of windows; sym=False for the periodic ones | scipy/signal/windows/_windows.py |
| Whisper’s frontend in C | optional L14.6 | STFT power with center reflect padding and drop-last in C | c/src/kernels/audio.c |