Ten problems on rotary position embeddings (RoPE): the 2-D rotation matrix and its algebra, why the score of a rotated query and key depends only on their relative position, the score as a sinusoid in the offset, the block-diagonal form and shift invariance, the complex-number form, the frequencies and wavelengths at d = 128, the backward pass through the rotation, the derivative with respect to the offset and the frequencies, RoPE inside attention backward, and position interpolation, with worked solutions and the mistakes that lose the position or flip its sign.
Before you start
Rotary position embeddings put position into attention without adding anything to the token: each query and key is rotated, two coordinates at a time, by an angle proportional to its position, and the dot product of two rotated vectors then depends only on how far apart they are. That one fact is the whole design, and it follows from three lines of trigonometry. These ten problems derive it for one pair of coordinates, extend it to the full block-diagonal rotation, rewrite it with complex numbers, work out the frequencies and wavelengths used in practice, and then do what a framework does: backpropagate through the rotation, inside attention. The five mistakes are the ones that leave the model running: both vectors rotated by the same angle, a transposed rotation with the sign of the offset flipped, values rotated too, a backward pass that rotates forwards, and a pairing of coordinates that differs between queries and keys.
The conventions are those of the attention page: vectors are columns, a gradient has the shape of its variable, and a token at position m has query q∈Rd and key k∈Rd with d even.
The 2-D rotation by angle α is ;R(α)=(cosαsinα−sinαcosα); it turns a vector anticlockwise by .α. The trig-identities page's addition formulas are used freely.
The coordinates of q are grouped into d/2 pairs: pair p is q(p)=(q2p−1,q2p)⊤∈R2 for .p=1,…,d/2. Each pair has its own frequency ,θp=10000−2(p−1)/d, so θ1=1 and the frequencies decrease geometrically. The ratio θp+1/θp is .10000−2/d.
RoPE at position m rotates pair p by the angle :mθp:,Rm=blockdiag(R(mθ1),…,R(mθd/2)), a d×d matrix, and the rotated query is .q~=Rmq. A key at position n becomes .k~=Rnk. The score is s=q~⊤k~ (the attention page's Smn before the division by ).d). Values are not rotated.
Implementations that pair coordinate j with j+d/2 instead ("rotate half") are the same construction after one fixed permutation of the coordinates, applied to every q and k alike.
Complex form: a pair (x,y)⊤ is the complex number x+iy with ;i2=−1;zˉ is the conjugate and Re the real part. Euler's formula is .eiα=cosα+isinα.
L is a scalar loss and g=∂L/∂s the upstream gradient of a score; inside attention, ,G,G~ and ∇SL are the attention page's matrices.
Show that ,R(α)⊤R(α)=I,,detR(α)=1,R(α)R(β)=R(α+β) and .R(α)⊤=R(−α).
·
One pair with frequency :θ: the query q∈R2 at position m and the key k∈R2 at position .n. Show that .(R(mθ)q)⊤(R(nθ)k)=q⊤R((n−m)θ)k.
··
For one pair, write s(φ)=q⊤R(φ)k in the form ,acosφ+bsinφ, giving a and b in terms of the entries of q and .k. Show that ,a2+b2=∥q∥2∥k∥2, and hence that s(φ)=∥q∥∥k∥cos(φ−ψ) for some angle .ψ. What is ?ψ?
··
Full dimension .d. Show that ,Rm⊤Rn=Rn−m, that the score (Rmq)⊤(Rnk) is a sum over pairs that depends on m and n only through n−m (so it is unchanged if both positions shift by the same ),s), and that .∥Rmq∥=∥q∥.
··
Complex form. Write pair p of q as zp=q2p−1+iq2p and of k as .wp=k2p−1+ik2p. Show that rotating a pair by α is multiplying its complex number by ,eiα, that ,q(p)⊤k(p)=Re(zpwˉp), and that the RoPE score is .Re∑pzpwˉpei(m−n)θp.
·
With d=128 and base :10000: give θ1 and ,θ64, the wavelength 2π/θp of the first and last pairs in positions, and the number of pairs whose wavelength is at most a context length of 4096 (the pairs that complete at least one full turn inside the context).
···
Backward through the rotation. For s=(Rmq)⊤(Rnk) with upstream gradient ,g=∂L/∂s, compute ∇qL and .∇kL. Then, for a general vector u=∇q~L of upstream gradient at the rotated query ,q~=Rmq, give .∇qL.
···
Show that .dφdR(φ)=R(φ+π/2). Treating the offset Δ=n−m as a continuous variable, compute ∂s/∂Δ for the full score ,s(Δ)=∑pq(p)⊤R(Δθp)k(p), and compute ∂s/∂θp for learned or rescaled frequencies.
···
Attention with RoPE: Q~ and K~ have rows (Rmqm)⊤ and (Rnkn)⊤ for tokens at positions ,m,n=1,…,N,,S=Q~K~⊤/d,A=softmax(S) row-wise, .O=AV. Using the attention page's results for ∇Q~L and ,∇K~L, write ∇qmL and .∇knL. Then show that shifting every position by the same s leaves A and O unchanged.
··
Position interpolation extends a model trained at context L0 to context sL0 by using positions m/s in place of .m. Show that this equals keeping the positions and dividing every frequency by ,s, that every wavelength is multiplied by ,s, and compute, for d=128 and ,s=4, how many pairs now complete a turn within 4096 positions.
,∇qmL=R−m(d1(∇SL)K~)m,:⊤,:∇knL=R−n(d1(∇SL)⊤Q~)n,:⊤: each row of the attention page's gradient rotated back by its own position; A and O are invariant to a common shift of all positions
Positions m/s with frequencies θp equal positions m with frequencies ;θp/s; wavelengths scale by ;s; at ,d=128,,s=4,36 pairs complete a turn within 4096 positions, down from 46
Worked solutions
Problem 1
Show that ,R(α)⊤R(α)=I,,detR(α)=1,R(α)R(β)=R(α+β) and .R(α)⊤=R(−α).
.R(α)⊤R(α)=(cosα−sinαsinαcosα)(cosαsinα−sinαcosα)=(cos2α+sin2α00sin2α+cos2α)=I.Multiply out; the off-diagonal entries are ,−cosαsinα+sinαcosα=0, and .cos2+sin2=1.
.detR(α)=cos2α+sin2α=1.ad−bc with .b=−sinα.
.R(α)R(β)=(cosαcosβ−sinαsinβsinαcosβ+cosαsinβ−cosαsinβ−sinαcosβ−sinαsinβ+cosαcosβ)=R(α+β).The entries are the addition formulas: cos(α+β)=cosαcosβ−sinαsinβ and .sin(α+β)=sinαcosβ+cosαsinβ.
.R(α)⊤=(cosα−sinαsinαcosα)=(cos(−α)sin(−α)−sin(−α)cos(−α))=R(−α).cos is even and sin is odd.
,R(α)⊤R(α)=I,,detR(α)=1,,R(α)R(β)=R(α+β),R(α)⊤=R(−α)A rotation preserves lengths and angles (orthogonal, determinant +1 so no reflection), rotations compose by adding angles, and the inverse of a rotation is its transpose, which is the rotation back. Step 3 also shows rotations in the plane commute: .R(α)R(β)=R(β)R(α).
Problem 2
One pair with frequency :θ: the query q∈R2 at position m and the key k∈R2 at position .n. Show that .(R(mθ)q)⊤(R(nθ)k)=q⊤R((n−m)θ)k.
(R(mθ)q)⊤(R(nθ)k)=q⊤R((n−m)θ)kThe absolute positions m and n have cancelled: the score is the unrotated q against k rotated by the relative offset n−m times the frequency. Nothing was assumed about q or ,k, so this holds for whatever the projections produce.
Problem 3
For one pair, write s(φ)=q⊤R(φ)k in the form ,acosφ+bsinφ, giving a and b in terms of the entries of q and .k. Show that ,a2+b2=∥q∥2∥k∥2, and hence that s(φ)=∥q∥∥k∥cos(φ−ψ) for some angle .ψ. What is ?ψ?
.R(φ)k=(k1cosφ−k2sinφ,k1sinφ+k2cosφ)⊤.Matrix times vector.
.s(φ)=q1k1cosφ−q1k2sinφ+q2k1sinφ+q2k2cosφ=(q1k1+q2k2)cosφ+(q2k1−q1k2)sinφ.Dot with q and collect the cos and sin terms.
a=q1k1+q2k2=q⊤k and .b=q2k1−q1k2.Read off step 2. a is the dot product; b is the 2-D cross product of k and q (the signed area of their parallelogram).
.a2+b2=q12k12+2q1k1q2k2+q22k22+q22k12−2q2k1q1k2+q12k22=(q12+q22)(k12+k22).Expand both squares; the cross terms cancel and the remaining four terms factor.
Let r=∥q∥∥k∥ and choose ψ with rcosψ=a and .rsinψ=b. Then .s(φ)=r(cosψcosφ+sinψsinφ)=rcos(φ−ψ).Step 4 says (a,b) lies on the circle of radius ,r, so such a ψ exists (when );r>0); then the difference formula for cosine (the trig-identities page).
ψ is the angle from k to .q.With k along the first axis, k=(∥k∥,0)⊤ and q=∥q∥(cosψ′,sinψ′)⊤ at angle :ψ′: then a=∥q∥∥k∥cosψ′ and ,b=q2k1=∥q∥∥k∥sinψ′, so .ψ=ψ′. The dot product is unchanged by rotating both vectors, so this holds in general.
,s(φ)=(q⊤k)cosφ+(q2k1−q1k2)sinφ=∥q∥∥k∥cos(φ−ψ),ψ the angle from k to qAs the offset n−m grows, the score of one pair oscillates between ±∥q∥∥k∥ with period ,2π/θ, and is largest when the rotation by φ=(n−m)θ brings k into line with .q. The full score (Problem 4) is a sum of d/2 such sinusoids with different periods, which is what lets it single out particular offsets.
Problem 4
Full dimension .d. Show that ,Rm⊤Rn=Rn−m, that the score (Rmq)⊤(Rnk) is a sum over pairs that depends on m and n only through n−m (so it is unchanged if both positions shift by the same ),s), and that .∥Rmq∥=∥q∥.
.Rm⊤Rn=blockdiag(R(mθp)⊤R(nθp))p=blockdiag(R((n−m)θp))p=Rn−m.Block-diagonal matrices multiply block by block, and the transpose of a block-diagonal matrix is block-diagonal with transposed blocks; each block is Problem 2 with .θ=θp.
.(Rmq)⊤(Rnk)=q⊤Rm⊤Rnk=q⊤Rn−mk.(Ab)⊤=b⊤A⊤ and step 1.
.q⊤Rn−mk=∑p=1d/2q(p)⊤R((n−m)θp)k(p).A block-diagonal matrix acts on each pair separately, and the dot product adds up the pairs.
Replacing (m,n) by (m+s,n+s) leaves n−m unchanged, so the score is unchanged.Step 3 contains m and n only in the combination .n−m.
.∥Rmq∥2=q⊤Rm⊤Rmq=q⊤R0q=q⊤q.Step 1 with ,n=m, and R0=I since .R(0)=I.
;Rm⊤Rn=Rn−m;,(Rmq)⊤(Rnk)=∑pq(p)⊤R((n−m)θp)k(p), a function of n−m alone; ∥Rmq∥=∥q∥The score depends on relative position, and so do the attention weights built from it (Problem 9): a sequence shifted by s positions attends the same way. Each pair contributes ∥q(p)∥∥k(p)∥cos((n−m)θp−ψp) by Problem 3, and the rotation never changes the size of a query or key, so the 1/d scaling argument of the attention page is untouched.
Problem 5
Complex form. Write pair p of q as zp=q2p−1+iq2p and of k as .wp=k2p−1+ik2p. Show that rotating a pair by α is multiplying its complex number by ,eiα, that ,q(p)⊤k(p)=Re(zpwˉp), and that the RoPE score is .Re∑pzpwˉpei(m−n)θp.
.eiα(x+iy)=(cosα+isinα)(x+iy)=(xcosα−ysinα)+i(xsinα+ycosα).Euler's formula, then multiply out using .i2=−1.
The real and imaginary parts are the two entries of .R(α)(x,y)⊤.Compare with R(α) applied to :(x,y)⊤: first entry ,xcosα−ysinα, second .xsinα+ycosα.
,zwˉ=(x1+iy1)(x2−iy2)=(x1x2+y1y2)+i(y1x2−x1y2), so .Re(zwˉ)=x1x2+y1y2.Multiply out; the real part is the dot product of the two pairs, and the imaginary part is Problem 3's .b.
z~p=eimθpzp and w~p=einθpwp are the rotated pairs.Step 2 with α=mθp and .nθp.
.z~pw~p=eimθpzpe−inθpwˉp=zpwˉpei(m−n)θp.,eiβw=e−iβwˉ, and exponents add.
Rotation by α is multiplication by ;eiα;;q(p)⊤k(p)=Re(zpwˉp); score =Re∑pzpwˉpei(m−n)θpStep 3 applied to the rotated pairs and summed over .p. This is how RoPE is usually implemented: view the last axis as complex, multiply by a precomputed table of ,eimθp, view as real again. The relative-position property is the exponent rule ,eimθe−inθ=ei(m−n)θ, which is Problem 1's R(α)R(β)=R(α+β) in different clothes.
Problem 6
With d=128 and base :10000: give θ1 and ,θ64, the wavelength 2π/θp of the first and last pairs in positions, and the number of pairs whose wavelength is at most a context length of 4096 (the pairs that complete at least one full turn inside the context).
θ1=100000=1 and .θ64=10000−126/128.p=1 gives exponent ;0;p=64 gives .−2⋅63/128=−126/128.
,10000126/128=e(126/128)ln10000=e0.984375⋅9.2103≈e9.0664≈8660, so .θ64≈1.15×10−4..ln10000=4ln10≈9.2103.
Wavelength of pair :1:2π/1≈6.28 positions; of pair :64:2π⋅8660≈54,400 positions.A pair rotates by θp per position, so it completes a turn every 2π/θp positions.
Pair p completes a turn within 4096 positions iff 2π/θp≤4096 iff .100002(p−1)/128≤4096/(2π).Substitute θp and invert; the inequality direction is preserved because .θp>0.
,1282(p−1)ln10000≤ln2π4096≈ln651.9≈6.480, so .p−1≤2⋅9.2103128⋅6.480≈45.03.Take logs; .4096/(2π)≈651.9.
,θ1=1,;θ64≈1.15×10−4; wavelengths ≈6.3 and ≈54,400 positions; 46 of the 64 pairs complete a turn within 4096 positions,p≤46.03, so .p=1,…,46. The remaining 18 pairs turn through less than a full circle over the whole context and act like a slowly varying, almost monotone position signal; the fast pairs resolve nearby offsets. Trained at a shorter context, those slow pairs have never seen the angles a longer context produces, which is the problem position interpolation addresses (Problem 10).
Problem 7
Backward through the rotation. For s=(Rmq)⊤(Rnk) with upstream gradient ,g=∂L/∂s, compute ∇qL and .∇kL. Then, for a general vector u=∇q~L of upstream gradient at the rotated query ,q~=Rmq, give .∇qL.
.s=q⊤Rn−mk.Problem 4, step 2.
.∇qs=Rn−mk.s=q⊤c with c=Rn−mk not depending on ,q, and ∇q(q⊤c)=c (the matrix-calculus page).
,s=k⊤Rn−m⊤q=k⊤Rm−nq, so .∇ks=Rm−nq.A scalar equals its transpose; Rn−m⊤=Rm−n by Problem 1, step 4, block by block.
∇qL=gRn−mk and .∇kL=gRm−nq.L depends on q and k only through ,s, so the chain rule multiplies by .g.
q~=Rmq gives ,∂q~/∂q=Rm, so .∇qL=Rm⊤∇q~L=Rm⊤u=R−mu.A linear map's Jacobian is its matrix, and for a scalar ,L,∇qL=(∂q~/∂q)⊤∇q~L (the Jacobians page); then .Rm⊤=R−m.
,∇qL=gRn−mk,;∇kL=gRm−nq; in general :∇qL=R−m∇q~L: rotate the upstream gradient back by the token's own angleStep 4 is step 5 applied to :∇q~L=gk~=gRnk:.R−mRnk=Rn−mk. The backward pass of RoPE is RoPE with the sign of the position flipped, which in the complex form is multiplication by ,e−imθp, the conjugate table. It costs the same as the forward pass and needs no stored activations beyond the positions.
Problem 8
Show that .dφdR(φ)=R(φ+π/2). Treating the offset Δ=n−m as a continuous variable, compute ∂s/∂Δ for the full score ,s(Δ)=∑pq(p)⊤R(Δθp)k(p), and compute ∂s/∂θp for learned or rescaled frequencies.
.dφdR(φ)=(−sinφcosφ−cosφ−sinφ).Differentiate each entry.
.R(φ+π/2)=(cos(φ+π/2)sin(φ+π/2)−sin(φ+π/2)cos(φ+π/2))=(−sinφcosφ−cosφ−sinφ).cos(φ+π/2)=−sinφ and sin(φ+π/2)=cosφ (the trig-identities page). The two matrices agree.
.∂Δ∂q(p)⊤R(Δθp)k(p)=θpq(p)⊤R(Δθp+π/2)k(p).Chain rule: the angle Δθp has derivative θp with respect to ,Δ, and step 2 gives the derivative of R with respect to its angle. q(p) and k(p) are constants here.
.∂s/∂Δ=∑pθpq(p)⊤R(Δθp+π/2)k(p).Sum step 3 over the pairs.
θp appears only in pair p's term, and the angle Δθp has derivative Δ with respect to .θp.Each frequency belongs to one block.
.∂s/∂θp=Δq(p)⊤R(Δθp+π/2)k(p).Chain rule as in step 3 with the roles of Δ and θp swapped.
;dφdR(φ)=R(φ+π/2);;∂s/∂Δ=∑pθpq(p)⊤R(Δθp+π/2)k(p);∂s/∂θp=Δq(p)⊤R(Δθp+π/2)k(p)For one pair, by Problem 3 with ,φ→φ+π/2, the derivative is ,−asinφ+bcosφ, the derivative of .acosφ+bsinφ. The frequency gradient carries a factor :Δ: distant pairs of tokens push the frequencies hardest, which is why learned or fine-tuned frequencies are sensitive to the longest offsets in the training data.
Problem 9
Attention with RoPE: Q~ and K~ have rows (Rmqm)⊤ and (Rnkn)⊤ for tokens at positions ,m,n=1,…,N,,S=Q~K~⊤/d,A=softmax(S) row-wise, .O=AV. Using the attention page's results for ∇Q~L and ,∇K~L, write ∇qmL and .∇knL. Then show that shifting every position by the same s leaves A and O unchanged.
∇Q~L=d1(∇SL)K~ and .∇K~L=d1(∇SL)⊤Q~.The attention page's Problems 5 and 6, with ,Q~,K~ in place of ,Q,:K: the attention computation is unchanged, it only receives rotated inputs.
Row m of Q~ depends on qm alone, through .q~m=Rmqm.Each token is rotated by its own position; no other token's query enters its row.
.∇qmL=Rm⊤(row m of ∇Q~L)⊤=R−m(∇Q~L)m,:⊤.Problem 7, step 5, applied to row :m: the upstream gradient at q~m is row m of ,∇Q~L, written as a column.
.∇knL=R−n(∇K~L)n,:⊤.The same argument for keys, with .k~n=Rnkn.
Smn=(Rmqm)⊤(Rnkn)/d depends on m and n only through .n−m.Problem 4, step 4, for every entry.
Shifting all positions by s leaves every Smn unchanged, hence A and .O=AV.;(n+s)−(m+s)=n−m;A is a function of S and V is not rotated.
,∇qmL=R−m(d1(∇SL)K~)m,:⊤,:∇knL=R−n(d1(∇SL)⊤Q~)n,:⊤: each row of the attention page's gradient rotated back by its own position; A and O are invariant to a common shift of all positionsSo RoPE is a layer wrapped around attention: forward, rotate Q and K row by row; backward, rotate the gradients back row by row and hand the rest to the attention page. The shift invariance means a RoPE model sees no absolute position at all; the only asymmetry it has is the causal mask, which tells a token how many keys lie before it.
Problem 10
Position interpolation extends a model trained at context L0 to context sL0 by using positions m/s in place of .m. Show that this equals keeping the positions and dividing every frequency by ,s, that every wavelength is multiplied by ,s, and compute, for d=128 and ,s=4, how many pairs now complete a turn within 4096 positions.
Pair p at position m/s is rotated by .(m/s)θp=m(θp/s).The angle is a product, and the factor 1/s can be attached to either side.
So the score with positions ,m/s,n/s and frequencies θp equals the score with positions ,m,n and frequencies .θp/s.Step 1 applies to every pair of every query and key; the rotated vectors are identical.
The wavelength becomes .2π/(θp/s)=s⋅2π/θp.Problem 6's definition with the new frequency.
Pair p completes a turn within 4096 iff s⋅2π/θp≤4096 iff .100002(p−1)/128≤4096/(2πs)=4096/(8π)≈163.0.Problem 6, step 4, with 2πs in place of .2π.
,p−1≤2ln10000128ln163.0≈18.42128⋅5.094≈35.4, so .p≤36.4.Take logs as in Problem 6, step 5.
Positions m/s with frequencies θp equal positions m with frequencies ;θp/s; wavelengths scale by ;s; at ,d=128,,s=4,36 pairs complete a turn within 4096 positions, down from 46The first pair's wavelength goes from 2π to 8π positions. Every angle the model sees at the long context is one it saw during training (its range is the same as at ),L0), which is why interpolation needs little fine-tuning; the cost is that adjacent positions now differ by a quarter of the angle they used to, so the fast pairs resolve neighbours less sharply. Schemes such as NTK-aware scaling and YaRN divide the slow pairs' frequencies by more than the fast pairs', trading between the two effects, and Problems 6 and 10 are the arithmetic they rest on.
Where this goes wrong
1. Rotating the query and the key by the same angle
RoPE is applied to queries and keys with one function, and the position passed in has to be the right one for each.
q~=Rmq for the query at position mRight so far.
“Apply the same rotation to the key.”The slip that causes the mistake: the key belongs to position ,n, and its rotation must use ,n, not the query's .m.
s=(Rmq)⊤(Rmk)=q⊤Rm⊤Rmk=q⊤kProblem 4, step 1 with :n=m:.Rm⊤Rm=I. Every score is the plain dot product, the positions have cancelled completely, and the model is a bag of tokens with no position information at all; the loss still goes down, so nothing fails loudly.
2. Transposing the wrong rotation
The relative-position step has a transpose in it, and which rotation carries it decides the sign of the offset.
(Rmq)⊤(Rnk)=q⊤Rm⊤RnkRight so far: Problem 2, step 1.
“Rm⊤Rn is the rotation by .m−n.”The shortcut that causes the mistake: the transposed factor is the one whose angle is negated, so .Rm⊤Rn=R−mRn=Rn−m.
s=q⊤Rm−nkThe offset's sign is flipped. For one pair the score is acosφ+bsinφ (Problem 3), and φ→−φ flips the sign of the b term, so the two agree only when ,b=q2k1−q1k2=0, that is, when q and k are parallel. In Problem 7 the same slip gives ∇qL=gRm−nk in place of ,gRn−mk, which fails a gradient check for every pair of positions except .m=n.
3. Rotating the values as well
The rotation is applied to two of the three projections, and the third looks like an omission.
q~m=Rmqm and k~n=Rnkn give Smn depending on n−m onlyRight so far: Problem 4.
“For consistency, rotate V by position too.”The analogy that causes the mistake: ,Q,K and V are produced alike, so they should be treated alike; but position enters attention through the scores, and V never touches a score.
Om=∑nAmnRnvnShifting all positions by s leaves A unchanged (Problem 9) but replaces Rn by ,Rn+s, so the output of every token changes with the absolute position of the sequence, and the relative-position property is lost at the output. With the rotation on Q and K only, Om=∑nAmnvn is invariant to the shift.
4. Backward pass that rotates forwards
The backward pass of a layer reuses the forward code, and for RoPE the forward code rotates by .+mθp.
q~=Rmq and the upstream gradient u=∇q~LRight so far.
“The gradient flows back through the same rotation.”The shortcut that causes the mistake: the backward of a linear map is its transpose, and Rm⊤=R−m is the rotation the other way.
∇qL=RmuProblem 7 gives .R−mu. For the score s=q~⊤k~ this makes ∇qL=gRmRnk=gRm+nk instead of :gRn−mk: a gradient with the right norm and the wrong direction, by the angle 2mθp in pair .p. In the complex implementation the fix is one conjugate: multiply by e−imθp on the way back.
5. Pairing the coordinates differently for queries and keys
The paper pairs coordinate 2p−1 with ;2p; many implementations pair j with .j+d/2. Both are RoPE, until they are mixed.
Pair (q2p−1,q2p) rotated by ,mθp, and the same pairing for ,k, gives a score depending on n−mRight so far: Problem 4, whichever fixed pairing is used, provided it is the same for q and .k.
“The pairing is a convention, so the query code can use one and the key code (or the pretrained weights) another.”The assumption that causes the mistake: Problem 2 cancels Rm⊤Rn block by block, which needs the blocks of Rm and Rn to sit on the same pairs of coordinates.
Rotate (q2p−1,q2p) by mθp but (kj,kj+d/2) by nθjRm⊤Rn is no longer block-diagonal on matching pairs, so the score is not a function of :n−m: it changes when the sequence is shifted, and the model trained with one convention is garbage under the other. This is the bug behind permuting the query and key projection weights when converting checkpoints between the two layouts; the test is Problem 4, step 4: compute the score at (m,n) and at (m+s,n+s) and check they agree.