Skip to content

RoPE (half and interleaved layouts, partial rotary)

ModuleL7.3 · build · Python · Pass 5 · 3 h, plus your graded tests (rung R5)
You buildpython/tinyllm/modern/rope.py: rope_cos_sin, apply_rope, RopeSpec; and your own oracle tests in python/tests/l7-3-rope/
Contractcourse/contracts/py/tinyllm/modern/rope.pyi
Testscourse/tests/L7.3/test_rope.py (what they check: section 4), golden values from transformers 5.19.0 (LlamaRotaryEmbedding, apply_rotary_pos_emb, GPT-NeoX partial rotary) and Meta’s complex-number apply_rotary_emb in course/fixtures/L7.3/rope_hf.npz (course/oracle/L7.3/rope_hf.py); your tests are graded by mutation, threshold 0.80 with every pitfall fault required
NeedsL0.2 F.concat, F.stack, F.reshape · L0.1 Tensor and its indexing · M00.2 rotate_pairs (the tests compare against it) · reading: M00.3 (the frequency ladder) (or --ref-deps)
Used byL7.5 rotates queries and keys · L7.6 the decoupled rope key of MLA · L7.7 windowed attention · L7.9 every Llama attention layer · L8.2 decoding at the right positions · later: L9.6 the C tl_rope_f32
MilestoneMS-L7 (SmolLM2-135M logits match Hugging Face)
Optional depthSu et al., “RoFormer: Enhanced Transformer with Rotary Position Embedding” (2021), sections 3.2 to 3.4; EleutherAI blog, “Rotary Embeddings: A Relative Revolution” (2021)
  • RoPE rotates pair ii of a query or key at position pp by the angle p θip\,\theta_i; rotations keep lengths and add angles, so a score depends only on the offset between positions (test_scores_depend_only_on_relative_position).
  • Two layouts choose which entries form a pair: HF’s half layout pairs xix_i with xi+r/2x_{i + r/2}, Meta’s interleaved layout pairs x2ix_{2i} with x2i+1x_{2i+1}; they are one rotation on a permuted vector (test_layouts_are_a_permutation_of_each_other).
  • Partial rotary rotates only the first rr entries and passes the rest through (test_partial_rotary_golden).
  • The angle is one float32 product, as in Hugging Face; lower precision destroys long positions (test_cos_sin_table_golden).
  • The backward is the inverse rotation (test_gradient_is_the_inverse_rotation).
Terminal window
ol start L7.3 # stubs rope.py; prints your test path and rung (R5)
ol tests L7.3 # the course tests
# write your oracle tests in python/tests/l7-3-rope/, then:
ol check L7.3 # course tests and the mutation grade of your tests
ol diff L7.3 # after passing: your code against the reference

Your 2017 transformer adds a sinusoidal or learned vector to each token embedding (L5.4), so position enters once, at the bottom, mixed into the content. SmolLM2 has no position embedding at all: its checkpoint has no wpe table, and a model that ignores position cannot tell “dog bites man” from “man bites dog”. The position goes in inside every attention layer, by rotating the queries and keys. Load SmolLM2 without it and the logits are garbage; rotate the wrong pairs and they are subtly wrong. This module builds the rotation exactly as Hugging Face does, in both layouts real checkpoints use.

SymbolMeaningType / shape
q,k∈Rdhq, k \in \mathbb{R}^{d_h}one head’s query and keyfloat32[d_h]
pp, mm, nntoken positionsint ≥0\ge 0
rrrotary_dim: how many entries rotate, even, r≤dhr \le d_hint
θi\theta_iinverse frequency of pair ii, i=0..r/2−1i = 0 .. r/2 - 1 (inv_freq)float[r/2]
R(α)R(\alpha)the 2D rotation (cos⁡α−sin⁡αsin⁡αcos⁡α)\begin{pmatrix} \cos\alpha & -\sin\alpha \\ \sin\alpha & \cos\alpha \end{pmatrix}2×22 \times 2
RpR_protate pair ii by p θip\,\theta_i for every iidh×dhd_h \times d_h
ssattention_scaling (YaRN, L7.4)float

M00.2 showed that R(α)R(\alpha) keeps lengths, R(α)⊤=R(−α)R(\alpha)^\top = R(-\alpha), and R(α)R(β)=R(α+β)R(\alpha) R(\beta) = R(\alpha + \beta). RoPE splits qq into r/2r/2 pairs and rotates pair ii of the token at position pp by p θip\,\theta_i. For a query at mm and a key at nn, pair by pair:

⟨R(mθi)qi,R(nθi)ki⟩=qi⊤R(mθi)⊤R(nθi) ki=qi⊤R((n−m)θi) ki.\langle R(m\theta_i) q_i, R(n\theta_i) k_i \rangle = q_i^\top R(m\theta_i)^\top R(n\theta_i)\, k_i = q_i^\top R((n - m)\theta_i)\, k_i .

The score depends on n−mn - m only. Attention gets relative position for free, inside the dot product, and at no parameter cost. The frequencies θi=base−2i/r\theta_i = \text{base}^{-2i/r} form M00.3’s geometric ladder: pair 0 turns once per 2π2\pi positions, the last pair barely moves over the whole context, so fast pairs resolve nearby order and slow pairs distinguish far positions.

Which entries form a pair is a convention fixed by how the weights were trained:

  • half (rotate_half, HF Llama, Mistral, Qwen, SmolLM2, gpt-oss): pair ii is (xi,xi+r/2)(x_i, x_{i + r/2}).
  • interleaved (the RoFormer paper, Meta’s original Llama code): pair ii is (x2i,x2i+1)(x_{2i}, x_{2i+1}), read as one complex number x2i+i x2i+1x_{2i} + i\,x_{2i+1} multiplied by eipθie^{i p \theta_i}.

Rotating pair (a,b)(a, b) by α\alpha gives (acos⁡α−bsin⁡α, asin⁡α+bcos⁡α)(a\cos\alpha - b\sin\alpha,\ a\sin\alpha + b\cos\alpha) in either layout. Reorder the entries (a0,b0,a1,b1,… )(a_0, b_0, a_1, b_1, \dots) into (a0,a1,…,b0,b1,… )(a_0, a_1, \dots, b_0, b_1, \dots) and the interleaved rotation becomes the half rotation. Hugging Face converts Meta’s checkpoints by permuting the rows of q_proj and k_proj exactly this way, so their projections produce half-layout vectors directly.

rope_cos_sin(positions, inv_freq, s) returns one cosine and one sine per pair, shape positions.shape + (r/2,). The angle p θip\,\theta_i is computed as Hugging Face does: one float32 product of float32(p) and float32(θ_i). At position 65535 the float32 spacing is about 0.004, so the angle itself is known only to a few thousandths of a radian; computing it in float16 (spacing 32 at that position, and 65535 overflows) destroys it. Both tables are multiplied by ss, so the query and the key each grow by ss and every score by s2s^2: YaRN’s temperature (L7.4).

GPT-NeoX and Phi rotate only the first r<dhr < d_h entries of each head and leave the other dh−rd_h - r unchanged: those entries carry content with no position at all. apply_rope(x, cos, sin, layout, rotary_dim=r) rotates x[..., :r] and concatenates x[..., r:] back unchanged.

RpR_p is linear and orthogonal, so the gradient with respect to xx is Rp⊤g=R−p gR_p^\top g = R_{-p}\, g: the upstream gradient rotated back, which is apply_rope(g, cos, -sin). Building apply_rope from indexing, products with constant tables, F.concat, and F.stack gives exactly that through autograd.

dh=r=4d_h = r = 4, θ=(1,0.01)\theta = (1, 0.01), position p=1p = 1, x=(1,0,0,1)x = (1, 0, 0, 1). Angles: 11 and 0.010.01 radians. cos⁡1=0.540302\cos 1 = 0.540302, sin⁡1=0.841471\sin 1 = 0.841471, cos⁡0.01=0.999950\cos 0.01 = 0.999950, sin⁡0.01=0.010000\sin 0.01 = 0.010000.

Half layout. Pair 0 is (x0,x2)=(1,0)(x_0, x_2) = (1, 0), rotated by 1: (0.540302,0.841471)(0.540302, 0.841471). Pair 1 is (x1,x3)=(0,1)(x_1, x_3) = (0, 1), rotated by 0.01: (−0.010000,0.999950)(-0.010000, 0.999950). Writing pair 0 back to entries 0 and 2 and pair 1 to entries 1 and 3: (0.540302,−0.010000,0.841471,0.999950)(0.540302, -0.010000, 0.841471, 0.999950).

Interleaved layout. Pair 0 is (x0,x1)=(1,0)(x_0, x_1) = (1, 0): (0.540302,0.841471)(0.540302, 0.841471). Pair 1 is (x2,x3)=(0,1)(x_2, x_3) = (0, 1): (−0.010000,0.999950)(-0.010000, 0.999950). Output (0.540302,0.841471,−0.010000,0.999950)(0.540302, 0.841471, -0.010000, 0.999950).

Same two rotations, different entries. This is test_hand_example.

def rope_cos_sin(positions, inv_freq, attention_scaling=1.0) -> tuple[NDArray, NDArray]: ... # float32 [..., r/2]
def apply_rope(x: Tensor, cos, sin, layout="half", rotary_dim=None) -> Tensor: ... # [..., d_h]
@dataclass
class RopeSpec: inv_freq: NDArray; attention_scaling: float; layout: str; rotary_dim: int
TestKINDChecksWhy it matters downstream
test_hand_exampleunitsection 3 in both layoutsyou and the test agree on the pairs
test_cos_sin_table_goldengoldenHF’s tables up to position 65535long contexts in L7.4
test_cos_sin_attention_scaling_goldengoldenYaRN’s scaled tablesrope_scaling configs in L7.9
test_half_layout_goldengoldenHF’s apply_rotary_pos_emb, per-row positionsSmolLM2 in MS-L7
test_interleaved_layout_goldengoldenMeta’s complex-number rotationMeta-layout checkpoints
test_partial_rotary_goldengoldenGPT-NeoX with rotary_dim 8 of 16partial-rotary models
test_scores_depend_only_on_relative_positionpropertyshifted pairs give equal scoresthe reason RoPE exists
test_rotation_preserves_length_and_position_zero_is_identitypropertynorms kept; position 0 unchangedno scale drift with position
test_layouts_are_a_permutation_of_each_otherpropertythe HF conversion permutationloading Meta weights
test_interleaved_matches_rotate_pairsdifferentialagrees with M00.2one definition of a rotation
test_gradcheck_both_layouts_and_partialgradcheckfloat64 central differencesq_proj, k_proj learn through RoPE
test_gradient_is_the_inverse_rotationpropertybackward = rotation by −pθ-p\thetathe C backward later
test_rope_spec_fieldsunitthe spec drives apply_ropeL7.5, L7.6 take a RopeSpec
test_validationboundaryodd or oversized rotary_dim, bad tables, bad positionswiring bugs fail loudly

Your oracle is complex multiplication in numpy float64: form a+iba + ib from each pair (by the layout’s rule), multiply by eipθe^{i p \theta}, and split back. Compare apply_rope for both layouts and partial rotary, check the hand example, that attention_scaling multiplies both tables, that large positions stay precise, that the gradient is the inverse rotation, and that bad input raises. Import only tinyllm.modern.rope, tinyllm.autograd.tensor, and tinyllm.autograd.functional. ol check L7.3 requires 0.80 with every pitfall fault killed.

PitfallSymptomCaught by
1. adjacent pairs where the checkpoint expects half pairsthe model loads and every attention pattern is wrongtest_half_layout_golden (mutant s01)
2. rotating by −pθ-p\theta (a sign slip in sin)relative scores still look fine; the weights disagreetest_hand_example, test_cos_sin_table_golden (mutant s02)
3. partial rotary on the wrong entriesGPT-NeoX-style models degrade quietlytest_partial_rotary_golden (mutant s03)
4. angles in float16positions past 2048 round to even numbers; 65535 overflowstest_cos_sin_table_golden, test_scores_depend_only_on_relative_position (mutant s04)
scaling only the cosineYaRN scores are not s2s^2 times largertest_cos_sin_attention_scaling_golden (mutant s05)
reading half of each pair as a constantk_proj learns from half its gradienttest_gradcheck_both_layouts_and_partial (mutant s06)
leaving the interleaved output in half orderthe next layer reads permuted featurestest_interleaved_layout_golden (mutant s07)
DirectionModuleHow it uses this
BackL0.2F.concat, F.stack, F.reshape give the backward
BackL0.1Tensor slicing
BackM00.2the 2D rotation and rotate_pairs
ForwardL7.5rotates qq and kk with a RopeSpec before the scores
ForwardL7.4its scaled inv_freq and attention_scaling feed rope_cos_sin
ForwardL7.6MLA rotates only a decoupled rope part of each key
ForwardL7.7sliding-window attention rotates with the same tables
ForwardL7.9config.json’s rope_theta and rope_scaling build the spec
ForwardL8.2decoding passes each new token’s absolute position
ForwardL9.6tl_rope_f32 in C must match this
Your pieceProduction equivalentWhat it addsWhere to look
rope_cos_sinHF LlamaRotaryEmbeddingcaches tables, recomputes them for dynamic scalingtransformers/models/llama/modeling_llama.py
apply_ropevLLM RotaryEmbedding (rotary_embedding kernel)in place on q and k in one fused CUDA kernel, is_neox_style selects the layoutvLLM csrc/pos_encoding_kernels.cu
partial rotaryGPT-NeoX rotary_pct, Phi partial_rotary_factorthe fraction from the configHF GPT-NeoX and Phi modeling files
position handlingmultimodal RoPE (M-RoPE, Qwen2-VL)separate time, height, width sections of the pairsQwen2-VL report