Mel scale and filterbanks
Overview
Section titled “Overview”| Module | M12.6 · build · Python · Pass 12 · 2 to 3 h |
| You build | python/tinyllm/sig/mel.py: hz_to_mel and mel_to_hz (Slaney and HTK), mel_filterbank (librosa’s triangles with Slaney area normalization), power_to_db, and log_mel_whisper (log10, floor, 8-unit range, rescale) |
| Contract | course/contracts/py/tinyllm/sig/mel.pyi |
| Tests | course/tests/M12.6/test_mel.py (what they check: section 4), golden values from course/oracle/M12.6/mel_librosa.py (librosa 1.0) in course/fixtures/M12.6/mel_librosa.npz |
| Needs | M00.1 log_base (or --ref-deps). Reading: M12.5 (the power spectrogram the filters pool) |
| Used by | later L14.1 (Whisper’s log-mel) · data.11 (mel dHash for audio near-duplicates); they join the registry with B14’s L14 and data groups |
| Milestone | MS-P12 (the multimodal gate) |
| Optional depth | Stevens, Volkmann, Newman, “A scale for the measurement of the psychological magnitude pitch” (1937); Slaney, Auditory Toolbox (1998), MakeERBFilters and mfcc; Radford et al., “Robust speech recognition via large-scale weak supervision” (2022), section 2.2 |
Key Takeaways
Section titled “Key Takeaways”- The Slaney mel scale is linear below 1000 Hz (15 mel, Hz per mel) and logarithmic above (27 mel per factor 6.4); HTK’s is (
test_hand_example_mel_scale). - A mel filterbank is triangles with edges equally spaced in mel, sampled at the rfft bin frequencies (
test_triangles_have_the_declared_edges,test_whisper_filterbanks_match_librosa). - Slaney normalization gives every triangle unit area in Hz, so wide high bands do not collect more energy for being wide (
test_slaney_rows_have_unit_area). - Power becomes decibels as with a floor; Whisper uses , clips 8 units (80 dB) below the maximum, and maps by (
test_hand_example_decibels,test_whisper_log_mel_scaling).
How to work this chapter
Section titled “How to work this chapter”ol start M12.6 # stubs python/tinyllm/sig/mel.py into your repool tests M12.6 # read the test catalog first: rung R0, you write no tests hereol check M12.6 # exit code is the verdictol check M12.6 --ref-deps # only if your M00.1 is not passing yetol diff M12.6 # after passing: your code against the reference1. Why now
Section titled “1. Why now”M12.5 turns 25 ms of audio into 201 power values per frame, evenly spaced from 0 to 8 kHz. Whisper’s encoder does not take those 201 numbers: it takes 80 (or 128 for large-v3) mel bands, each a weighted sum of bins, then a clipped logarithm. The weights are librosa’s mel filters, shipped inside Whisper as mel_filters.npz; if your filters differ in the fifth decimal, every feature L14.1 hands the encoder is wrong by a little, and parity with HF’s WhisperFeatureExtractor (within ) fails. The audio near-duplicate detector in data.11 hashes the same log-mel picture. This module builds the filters and the logs exactly as the reference libraries do.
2. Principles
Section titled “2. Principles”| Symbol | Meaning | Type / shape |
|---|---|---|
| frequency | Hz | |
| mel value of | mel | |
| filter edges, equally spaced in mel | Hz | |
| triangle at frequency | float | |
| center of rfft bin | Hz | |
| power spectrogram | float64[n_fft//2 + 1, T] | |
| floor before a log | float | |
| dynamic range kept below the maximum | dB |
2.1 The mel scale
Section titled “2.1 The mel scale”Listeners judge equal pitch steps as roughly equal frequency steps below about 1 kHz and equal frequency ratios above. The Slaney scale (librosa’s default, used by Whisper) models this piecewise: so 1000 Hz is 15 mel and every factor of 6.4 in frequency adds 27 mel. The HTK scale is one smooth formula, . The two give different numbers and different filters; Whisper uses Slaney.
2.2 Triangular filters and their normalization
Section titled “2.2 Triangular filters and their normalization”Place points equally spaced in mel between and (default ) and map them back to Hz: . Filter rises linearly from 0 at to 1 at and falls back to 0 at : sampled at the bin centers . Neighbouring triangles overlap by half, and high triangles are much wider in Hz than low ones. With Slaney normalization each row is multiplied by ; a triangle of height 1 and base has area , so every normalized filter has unit area and measures power density rather than total power. The filterbank is the matrix , and the mel spectrogram is .
2.3 Decibels
Section titled “2.3 Decibels”Loudness is perceived on a log scale too. A power ratio in decibels is : a factor of 10 in power is 10 dB. (Amplitudes use , because power is amplitude squared; using 20 on a power doubles every number.) power_to_db floors both and ref at amin so never happens, and with top_db raises everything to at least the maximum minus top_db, so silence does not stretch the range to . The base-10 logarithm is M00.1’s log_base(x, 10): a change of base, .
2.4 Whisper’s log-mel
Section titled “2.4 Whisper’s log-mel”Whisper’s features skip decibels and use directly: The floor is dB; the clip keeps 8 units, 80 dB, below the loudest bin of the clip; the affine map puts typical speech near . Silence floors everywhere at , which maps to . The maximum is taken over the whole clip, so the features of a frame depend on the loudest moment anywhere in the 30 s window.
3. Worked example by hand
Section titled “3. Worked example by hand”Mel values on the Slaney scale: Hz is mel; Hz is 15 mel; Hz is mel. On the HTK scale Hz is mel (test_hand_example_mel_scale).
Decibels of the powers with ref 1 and : , , and is floored to , giving . With the floor is , so the result is . With ref 10 everything drops by 10 dB: when not clipped. This is test_hand_example_decibels.
Whisper’s scaling of the identity filterbank applied to : with the floor gives ; clipping at gives ; and (in test_whisper_log_mel_scaling).
4. The interface
Section titled “4. The interface”def hz_to_mel(f, scale='slaney') -> NDArray: ...def mel_to_hz(m, scale='slaney') -> NDArray: ...def mel_filterbank(sr, n_fft, n_mels, fmin=0.0, fmax=None, scale='slaney', norm='slaney') -> NDArray: ...def power_to_db(p, ref=1.0, amin=1e-10, top_db=80.0) -> NDArray: ...def log_mel_whisper(power_spec, filters) -> NDArray: ...What the tests check
Section titled “What the tests check”| Test | KIND | Checks | Why it matters downstream |
|---|---|---|---|
test_hand_example_mel_scale | unit, smoke | section 3’s mel values on both scales | the scale itself |
test_hand_example_decibels | unit, smoke | section 3’s decibels with and without top_db and with a ref | the dB conventions |
test_whisper_filterbanks_match_librosa | golden | 80 and 128 bands at 16 kHz, , within of float64 librosa and of the float32 filters Whisper ships | L14.1’s features |
test_other_filterbanks_and_scales_match_librosa | golden | an HTK bank without normalization, a band-limited bank, both scale conversions, the default fmax | edges and scales |
test_power_to_db_matches_librosa | golden | power_to_db with defaults and with ref, amin, no top_db | decibel parity |
test_triangles_have_the_declared_edges | property | every row is the triangle formula on the mel-spaced edges; zero outside | the definition |
test_slaney_rows_have_unit_area | property | on a fine grid each normalized row integrates to 1 | why the normalization exists |
test_whisper_log_mel_scaling | unit | section 3’s matrix, silence at , range at most 2 units | Whisper’s feature scaling |
test_rejects_bad_arguments | boundary | amin 0, negative top_db, complex input, empty band, unknown norm or scale, mismatched shapes; input untouched | caller bugs fail early |
5. Pitfalls
Section titled “5. Pitfalls”| Pitfall | Symptom | Caught by |
|---|---|---|
| 1. the Slaney scale linear everywhere | high bands far too narrow | test_hand_example_mel_scale, test_whisper_filterbanks_match_librosa (mutant s01) |
| 2. skipping the Slaney normalization | high bands dominate every feature | test_slaney_rows_have_unit_area, test_whisper_filterbanks_match_librosa (mutant s02) |
| 3. edges equally spaced in Hz | a linear filterbank, not a mel one | test_triangles_have_the_declared_edges, test_whisper_filterbanks_match_librosa (mutant s03) |
| 4. on a power | every dB value doubled | test_hand_example_decibels, test_power_to_db_matches_librosa (mutant s04) |
5. clipping top_db against 0 dB instead of the maximum | quiet recordings clipped flat | test_hand_example_decibels, test_power_to_db_matches_librosa (mutant s05) |
| 6. the natural log in Whisper’s features | features scaled by | test_whisper_log_mel_scaling (mutant s06) |
| 7. Whisper’s range as 80 instead of 8 | the clip never happens | test_whisper_log_mel_scaling (mutant s07) |
| accepting amin = 0 | in the output | test_rejects_bad_arguments (mutant m01) |
| accepting fmin = fmax | a filterbank of NaNs | test_rejects_bad_arguments (mutant m02) |
| accepting a negative top_db | every value clipped above the maximum | test_rejects_bad_arguments (mutant m03) |
6. Where it’s used next
Section titled “6. Where it’s used next”| Direction | Module | How it uses this |
|---|---|---|
| Back | M00.1 | power_to_db and log_mel_whisper take base-10 logs with log_base(x, 10) |
| Back | M12.5 | the filters pool its power spectrogram (reading) |
| Forward | L14.1 | log_mel_whisper(power(stft(...))[:, :-1], mel_filterbank(16000, 400, 80)) (with B14’s L14 group) |
| Forward | data.11 | mel dHash windows for audio near-duplicates |
| Forward | S-M12 | mel, logs, decibels, and Whisper’s scaling by hand |
Going further
Section titled “Going further”| Your piece | Production equivalent | What it adds | Where to look |
|---|---|---|---|
mel_filterbank | librosa.filters.mel | the same filters, cached; norm as any norm | librosa/filters.py |
log_mel_whisper | HF WhisperFeatureExtractor | the full frontend in numpy or torch, batched, with padding to 30 s | transformers/models/whisper/feature_extraction_whisper.py |
power_to_db | torchaudio.transforms.AmplitudeToDB | the same floor and top_db on tensors | torchaudio/functional/functional.py |
| filters in Rust | the engine’s tl-engine/src/mm/audio.rs (L15.1) | log-mel from the committed filter file, proven against this module’s fixtures | L15.1 chapter |