Skip to content

Signals and Images

Convolution, resampling, the Fourier transform, and the mel scale: the numerical front end of every vision and speech model, built as numpy reference code the multimodal passes call.

  • Primary textbook: Oppenheim and Schafer, Discrete-Time Signal Processing (3rd ed.); free alternative: Smith, Mathematics of the Discrete Fourier Transform and Spectral Audio Signal Processing (free online, ccrma.stanford.edu)
  • Supplementary: Dumoulin and Visin, “A guide to convolution arithmetic for deep learning” (free); Goodfellow, Bengio, Courville, Deep Learning, ch. 9 (free); Szeliski, Computer Vision: Algorithms and Applications, ch. 3 (free)
  • Prerequisites: Precalculus (logs, Euler’s formula), Linear Algebra, Matrix Calculus & Autodiff, Numerical Methods & Floating Point
  • Estimated time: 2 weeks at 8-10 hrs/week (the six course modules plus the problem set)
  • A convolution layer is a linear map, a Toeplitz matrix; im2col turns it into one matmul and col2im (its adjoint) gives the backward pass
  • A layout is a memory order: the same logical NCHW tensor can store its channels first or last, and the strides say which
  • Sampling keeps only frequencies below half the rate; resizing and resampling must low-pass first or they alias
  • The DFT is a unitary change of basis, and an FFT computes it in O(nlog⁡n)O(n \log n) for every size whose factors it supports, including Whisper’s 400
  • Speech features are a windowed STFT, pooled into mel bands, on a log scale; each step has a convention a checkpoint was trained on
  • Work each chapter’s worked example by hand before reading its code; the first course test is that example
  • Reproduce the reference library’s numbers, not just its idea: Pillow, PyTorch, torchaudio, numpy, and librosa each make small choices (pass order, rounding, padding, window symmetry) that matter for parity
  • Finish with the problem set: it is where the arithmetic of L13 and L14 becomes second nature

Images and audio are samples of continuous signals. Every operation in a multimodal front end (convolve, resize, transform, filter, log) is a linear map followed at most by a pointwise nonlinearity, and getting its conventions exactly right is what makes a reimplementation agree with the model’s training pipeline.

Key ideas:

  • Deep learning’s convolution is cross-correlation; true convolution flips the kernel
  • Output size ⌊(n+2p−d(k−1)−1)/s⌋+1\lfloor (n + 2p - d(k-1) - 1)/s \rfloor + 1; groups split channels into independent convs
  • im2col and col2im: forward as one matmul, backward as the adjoint copy

Key ideas:

  • NCHW versus NHWC strides; a transpose is a view, a layout change is a copy
  • Rescale and normalize to float32 CHW; Pillow’s fixed-point grayscale
  • Patchify as a reshape: a ViT patch embedding is a stride-pp conv

Key ideas:

  • Nearest, bilinear, bicubic (Keys, a=−0.5a = -0.5 or −0.75-0.75), Lanczos3
  • Antialiasing stretches the kernel by the scale factor
  • Separable fixed-point passes reproduce Pillow; windowed-sinc filter banks reproduce torchaudio

Key ideas:

  • Unitarity and Parseval; conjugate symmetry of real signals
  • The convolution theorem and zero padding
  • Mixed-radix Cooley-Tukey; real FFTs by packing; the DCT-II via an FFT

Key ideas:

  • Nyquist and aliasing; windows and spectral leakage
  • Center reflect padding and frame counts; COLA and weighted overlap-add

Key ideas:

  • Slaney and HTK mel scales; triangular filterbanks with area normalization
  • Decibels, top_db, and Whisper’s log10 with an 8-unit range

ConceptKey FormulaApplication
Conv output size⌊(n+2p−d(k−1)−1)/s⌋+1\lfloor (n + 2p - d(k-1) - 1)/s \rfloor + 1sizing every conv, pool, and patch layer
im2col matmulY=Wflat⋅im2col⁡(x)Y = W_\text{flat} \cdot \operatorname{im2col}(x)fast conv forward; col2im for backward
Antialiased resizekernel stretched by max⁡(scale,1)\max(\text{scale}, 1)image preprocessing without moire
DFTX[k]=∑jx[j]e−2πijk/nX[k] = \sum_j x[j] e^{-2\pi i jk/n}spectra, convolution, DCT
STFT frames1+⌊n/hop⌋1 + \lfloor n/\text{hop} \rfloor (centered)Whisper’s 3001 frames
Slaney mel15+27log⁡6.4(f/1000)15 + 27 \log_{6.4}(f/1000) above 1 kHzspeech features

This topic is new in Pass 12 (the multimodal pass): no earlier topic owns sampling, the DFT, or interpolation. Each build module is numpy reference code with a call site in the vision (L13) or speech (L14) parts; optional C kernels for the same math are L13.8, L13.9, and L14.6.

ModuleTopicKindPass
M12.1Convolution as a linear mapbuild12
M12.2Image tensors and layoutsbuild12
M12.3Resampling and interpolationbuild12
M12.4DFT and FFTbuild12
M12.5Sampling, Nyquist, windows, and the STFTbuild12
M12.6Mel scale and filterbanksbuild12
S-M12Signals and images problem setsolve12
#ModuleChapterKindPass
1M12.1Convolution as a linear mapbuild12
2M12.2Image tensors and layoutsbuild12
3M12.3Resampling and interpolationbuild12
4M12.4DFT and FFTbuild12
5M12.5Sampling, Nyquist, windows, and the STFTbuild12
6M12.6Mel scale and filterbanksbuild12
7S-M12Signals and images problem set: convolution, sampling, Fourier, mel, contrastive losses, and WERsolve12
Signals ConceptConnected TrackApplication
Convolution, im2colNeural ArchitecturesCNN layers and their C kernels
FFT, DCTInformation Theorytransform coding and rate-distortion
Aliasing and resamplingNumerical Methods & Floating Pointerror bounds and stable evaluation
STFT and melFoundation ModelsWhisper and audio LLM front ends
CompanyHow This AppearsDifficulty
OpenAIWhisper’s log-mel frontend; image preprocessing for vision modelsAdvanced
GoogleSigLIP and ViT preprocessing; audio front endsAdvanced
MetaImage and video pipelines at scale; on-device ASRAdvanced
NVIDIAcuDNN convolution algorithms; DALI preprocessingAdvanced
AppleOn-device vision and speech with fixed-point kernelsAdvanced