Skip to content

Mel scale and filterbanks

ModuleM12.6 · build · Python · Pass 12 · 2 to 3 h
You buildpython/tinyllm/sig/mel.py: hz_to_mel and mel_to_hz (Slaney and HTK), mel_filterbank (librosa’s triangles with Slaney area normalization), power_to_db, and log_mel_whisper (log10, floor, 8-unit range, rescale)
Contractcourse/contracts/py/tinyllm/sig/mel.pyi
Testscourse/tests/M12.6/test_mel.py (what they check: section 4), golden values from course/oracle/M12.6/mel_librosa.py (librosa 1.0) in course/fixtures/M12.6/mel_librosa.npz
NeedsM00.1 log_base (or --ref-deps). Reading: M12.5 (the power spectrogram the filters pool)
Used bylater L14.1 (Whisper’s log-mel) · data.11 (mel dHash for audio near-duplicates); they join the registry with B14’s L14 and data groups
MilestoneMS-P12 (the multimodal gate)
Optional depthStevens, Volkmann, Newman, “A scale for the measurement of the psychological magnitude pitch” (1937); Slaney, Auditory Toolbox (1998), MakeERBFilters and mfcc; Radford et al., “Robust speech recognition via large-scale weak supervision” (2022), section 2.2
  • The Slaney mel scale is linear below 1000 Hz (15 mel, 200/3200/3 Hz per mel) and logarithmic above (27 mel per factor 6.4); HTK’s is 2595log⁡10(1+f/700)2595 \log_{10}(1 + f/700) (test_hand_example_mel_scale).
  • A mel filterbank is nmelsn_\text{mels} triangles with edges equally spaced in mel, sampled at the rfft bin frequencies (test_triangles_have_the_declared_edges, test_whisper_filterbanks_match_librosa).
  • Slaney normalization gives every triangle unit area in Hz, so wide high bands do not collect more energy for being wide (test_slaney_rows_have_unit_area).
  • Power becomes decibels as 10log⁡1010 \log_{10} with a floor; Whisper uses log⁡10\log_{10}, clips 8 units (80 dB) below the maximum, and maps by (x+4)/4(x + 4)/4 (test_hand_example_decibels, test_whisper_log_mel_scaling).
Terminal window
ol start M12.6 # stubs python/tinyllm/sig/mel.py into your repo
ol tests M12.6 # read the test catalog first: rung R0, you write no tests here
ol check M12.6 # exit code is the verdict
ol check M12.6 --ref-deps # only if your M00.1 is not passing yet
ol diff M12.6 # after passing: your code against the reference

M12.5 turns 25 ms of audio into 201 power values per frame, evenly spaced from 0 to 8 kHz. Whisper’s encoder does not take those 201 numbers: it takes 80 (or 128 for large-v3) mel bands, each a weighted sum of bins, then a clipped logarithm. The weights are librosa’s mel filters, shipped inside Whisper as mel_filters.npz; if your filters differ in the fifth decimal, every feature L14.1 hands the encoder is wrong by a little, and parity with HF’s WhisperFeatureExtractor (within 10−410^{-4}) fails. The audio near-duplicate detector in data.11 hashes the same log-mel picture. This module builds the filters and the logs exactly as the reference libraries do.

SymbolMeaningType / shape
fffrequencyHz
m(f)m(f)mel value of ffmel
e0,…,enmels+1e_0, \dots, e_{n_\text{mels}+1}filter edges, equally spaced in melHz
Fi(f)F_i(f)triangle ii at frequency fffloat
fk=k sr/nfftf_k = k\,\text{sr}/n_\text{fft}center of rfft bin kkHz
PPpower spectrogramfloat64[n_fft//2 + 1, T]
amin\text{amin}floor before a logfloat
top_db\text{top\_db}dynamic range kept below the maximumdB

Listeners judge equal pitch steps as roughly equal frequency steps below about 1 kHz and equal frequency ratios above. The Slaney scale (librosa’s default, used by Whisper) models this piecewise: m(f)={f/(200/3)f<100015+ln⁡(f/1000)ln⁡(6.4)/27f≥1000m(f) = \begin{cases} f / (200/3) & f < 1000 \\ 15 + \dfrac{\ln(f / 1000)}{\ln(6.4)/27} & f \ge 1000 \end{cases} so 1000 Hz is 15 mel and every factor of 6.4 in frequency adds 27 mel. The HTK scale is one smooth formula, m(f)=2595log⁡10(1+f/700)m(f) = 2595 \log_{10}(1 + f/700). The two give different numbers and different filters; Whisper uses Slaney.

2.2 Triangular filters and their normalization

Section titled “2.2 Triangular filters and their normalization”

Place nmels+2n_\text{mels} + 2 points equally spaced in mel between m(fmin)m(f_\text{min}) and m(fmax)m(f_\text{max}) (default sr/2\text{sr}/2) and map them back to Hz: e0<e1<⋯<enmels+1e_0 < e_1 < \dots < e_{n_\text{mels}+1}. Filter ii rises linearly from 0 at eie_i to 1 at ei+1e_{i+1} and falls back to 0 at ei+2e_{i+2}: Fi(f)=max⁡(0,  min⁡(f−eiei+1−ei,  ei+2−fei+2−ei+1)),F_i(f) = \max\left(0,\; \min\left(\frac{f - e_i}{e_{i+1} - e_i},\; \frac{e_{i+2} - f}{e_{i+2} - e_{i+1}}\right)\right), sampled at the bin centers fkf_k. Neighbouring triangles overlap by half, and high triangles are much wider in Hz than low ones. With Slaney normalization each row is multiplied by 2/(ei+2−ei)2/(e_{i+2} - e_i); a triangle of height 1 and base ei+2−eie_{i+2} - e_i has area (ei+2−ei)/2(e_{i+2} - e_i)/2, so every normalized filter has unit area and measures power density rather than total power. The filterbank is the matrix M∈Rnmels×(nfft/2+1)M \in \mathbb{R}^{n_\text{mels} \times (n_\text{fft}/2+1)}, and the mel spectrogram is MPM P.

Loudness is perceived on a log scale too. A power ratio in decibels is 10log⁡10(p/ref)10 \log_{10}(p / \text{ref}): a factor of 10 in power is 10 dB. (Amplitudes use 20log⁡1020 \log_{10}, because power is amplitude squared; using 20 on a power doubles every number.) power_to_db floors both pp and ref at amin so log⁡0\log 0 never happens, and with top_db raises everything to at least the maximum minus top_db, so silence does not stretch the range to −∞-\infty. The base-10 logarithm is M00.1’s log_base(x, 10): a change of base, ln⁡x/ln⁡10\ln x / \ln 10.

Whisper’s features skip decibels and use log⁡10\log_{10} directly: L=log⁡10(max⁡(MP,10−10)),L←max⁡(L,max⁡L−8),features=(L+4)/4.L = \log_{10}(\max(M P, 10^{-10})), \qquad L \leftarrow \max(L, \max L - 8), \qquad \text{features} = (L + 4)/4 . The floor 10−1010^{-10} is −100-100 dB; the clip keeps 8 units, 80 dB, below the loudest bin of the clip; the affine map puts typical speech near [−1,1][-1, 1]. Silence floors everywhere at log⁡1010−10=−10\log_{10} 10^{-10} = -10, which maps to −1.5-1.5. The maximum is taken over the whole clip, so the features of a frame depend on the loudest moment anywhere in the 30 s window.

Mel values on the Slaney scale: 500500 Hz is 500⋅3/200=7.5500 \cdot 3/200 = 7.5 mel; 10001000 Hz is 15 mel; 64006400 Hz is 15+27ln⁡(6.4)/ln⁡(6.4)=4215 + 27 \ln(6.4)/\ln(6.4) = 42 mel. On the HTK scale 700700 Hz is 2595log⁡102=781.172595 \log_{10} 2 = 781.17 mel (test_hand_example_mel_scale).

Decibels of the powers (1,10,10−12)(1, 10, 10^{-12}) with ref 1 and amin=10−10\text{amin} = 10^{-10}: 10log⁡101=010 \log_{10} 1 = 0, 10log⁡1010=1010 \log_{10} 10 = 10, and 10−1210^{-12} is floored to 10−1010^{-10}, giving −100-100. With top_db=80\text{top\_db} = 80 the floor is 10−80=−7010 - 80 = -70, so the result is (0,10,−70)(0, 10, -70). With ref 10 everything drops by 10 dB: (−10,0,−110)(-10, 0, -110) when not clipped. This is test_hand_example_decibels.

Whisper’s scaling of the identity filterbank applied to P=(10−121001104)P = \begin{pmatrix} 10^{-12} & 100 \\ 1 & 10^4 \end{pmatrix}: log⁡10\log_{10} with the floor gives (−10204)\begin{pmatrix} -10 & 2 \\ 0 & 4 \end{pmatrix}; clipping at 4−8=−44 - 8 = -4 gives (−4204)\begin{pmatrix} -4 & 2 \\ 0 & 4 \end{pmatrix}; and (L+4)/4=(01.512)(L + 4)/4 = \begin{pmatrix} 0 & 1.5 \\ 1 & 2 \end{pmatrix} (in test_whisper_log_mel_scaling).

def hz_to_mel(f, scale='slaney') -> NDArray: ...
def mel_to_hz(m, scale='slaney') -> NDArray: ...
def mel_filterbank(sr, n_fft, n_mels, fmin=0.0, fmax=None, scale='slaney', norm='slaney') -> NDArray: ...
def power_to_db(p, ref=1.0, amin=1e-10, top_db=80.0) -> NDArray: ...
def log_mel_whisper(power_spec, filters) -> NDArray: ...
TestKINDChecksWhy it matters downstream
test_hand_example_mel_scaleunit, smokesection 3’s mel values on both scalesthe scale itself
test_hand_example_decibelsunit, smokesection 3’s decibels with and without top_db and with a refthe dB conventions
test_whisper_filterbanks_match_librosagolden80 and 128 bands at 16 kHz, nfft=400n_\text{fft} = 400, within 10−1210^{-12} of float64 librosa and 10−710^{-7} of the float32 filters Whisper shipsL14.1’s features
test_other_filterbanks_and_scales_match_librosagoldenan HTK bank without normalization, a band-limited bank, both scale conversions, the default fmaxedges and scales
test_power_to_db_matches_librosagoldenpower_to_db with defaults and with ref, amin, no top_dbdecibel parity
test_triangles_have_the_declared_edgespropertyevery row is the triangle formula on the mel-spaced edges; zero outsidethe definition
test_slaney_rows_have_unit_areapropertyon a fine grid each normalized row integrates to 1why the normalization exists
test_whisper_log_mel_scalingunitsection 3’s matrix, silence at −1.5-1.5, range at most 2 unitsWhisper’s feature scaling
test_rejects_bad_argumentsboundaryamin 0, negative top_db, complex input, empty band, unknown norm or scale, mismatched shapes; input untouchedcaller bugs fail early
PitfallSymptomCaught by
1. the Slaney scale linear everywherehigh bands far too narrowtest_hand_example_mel_scale, test_whisper_filterbanks_match_librosa (mutant s01)
2. skipping the Slaney normalizationhigh bands dominate every featuretest_slaney_rows_have_unit_area, test_whisper_filterbanks_match_librosa (mutant s02)
3. edges equally spaced in Hza linear filterbank, not a mel onetest_triangles_have_the_declared_edges, test_whisper_filterbanks_match_librosa (mutant s03)
4. 20log⁡1020 \log_{10} on a powerevery dB value doubledtest_hand_example_decibels, test_power_to_db_matches_librosa (mutant s04)
5. clipping top_db against 0 dB instead of the maximumquiet recordings clipped flattest_hand_example_decibels, test_power_to_db_matches_librosa (mutant s05)
6. the natural log in Whisper’s featuresfeatures scaled by ln⁡10\ln 10test_whisper_log_mel_scaling (mutant s06)
7. Whisper’s range as 80 instead of 8the clip never happenstest_whisper_log_mel_scaling (mutant s07)
accepting amin = 0log⁡0=−∞\log 0 = -\infty in the outputtest_rejects_bad_arguments (mutant m01)
accepting fmin = fmaxa filterbank of NaNstest_rejects_bad_arguments (mutant m02)
accepting a negative top_dbevery value clipped above the maximumtest_rejects_bad_arguments (mutant m03)
DirectionModuleHow it uses this
BackM00.1power_to_db and log_mel_whisper take base-10 logs with log_base(x, 10)
BackM12.5the filters pool its power spectrogram (reading)
ForwardL14.1log_mel_whisper(power(stft(...))[:, :-1], mel_filterbank(16000, 400, 80)) (with B14’s L14 group)
Forwarddata.11mel dHash windows for audio near-duplicates
ForwardS-M12mel, logs, decibels, and Whisper’s scaling by hand
Your pieceProduction equivalentWhat it addsWhere to look
mel_filterbanklibrosa.filters.melthe same filters, cached; norm as any LpL^p normlibrosa/filters.py
log_mel_whisperHF WhisperFeatureExtractorthe full frontend in numpy or torch, batched, with padding to 30 stransformers/models/whisper/feature_extraction_whisper.py
power_to_dbtorchaudio.transforms.AmplitudeToDBthe same floor and top_db on tensorstorchaudio/functional/functional.py
filters in Rustthe engine’s tl-engine/src/mm/audio.rs (L15.1)log-mel from the committed filter file, proven against this module’s fixturesL15.1 chapter