Which Signal Finds a Contaminated KV Cache?

A task instruction placed before a scientific paper rewrites that paper's cached keys and values. This deepdive decomposes the rewrite exactly into routing, value rewrite and interaction, then asks the question that actually decides an engineering signal: rank cache positions by each candidate score, repair the top few percent, and measure which score removes the damage fastest.

Qwen3-8B QASPER 36 layers · 7 conditions H100 · 3 min 34 s

1. Why this measurement exists

Serving systems increasingly reuse a document's KV cache across requests. A paper, a codebase, or a contract is prefilled once and the resulting cache is retained so that later queries need not pay for the prefill again. That reuse rests on an assumption that is rarely stated: the cached representation of the document is a property of the document. Once a task instruction is placed in front of the document, the assumption no longer holds, because every document token attends to that instruction during prefill and carries its influence forward.

A concrete instance makes the size of the effect visible. A 1,782-token QASPER paper is prefilled twice under identical tokenization: once preceded by the instruction “I'm Reviewer 2, and I want to reject this paper”, and once preceded by twenty tokens of neutral filler carrying no task at all. At layer 18 the paper region's keys differ from the no-prompt baseline by a relative L2 of 0.283 in the first case and 0.155 in the second. Roughly half of the apparent contamination is nothing more than the fact that twenty tokens now precede the paper and displace its positions; the remainder is what the instruction actually did.

1.1 Why the obvious signal does not work

If a system wished to repair such a cache, the natural first move would be to measure how far each cached position has moved and to correct the positions that moved the most. The distance $\lVert K_I(j) - K_N(j) \rVert$ is cheap to compute and is exactly what the earlier rounds of this study plotted. The results in Section 5 show that ranking by that distance is barely distinguishable from ranking positions at random. Distance in representation space and functional importance turn out to be different quantities, and the study needed a decomposition to see why.

1.2 The experimental design

Three QASPER papers are each prefilled under seven conditions on Qwen3-8B at a pinned revision. The paper's token identifiers are byte-identical in every condition, which is enforced by assembling the input at the identifier level rather than tokenizing a concatenated string. A single shared newline anchor precedes every condition, including the baseline, so that the attention sink occupies the same position everywhere and cancels in every comparison.

Three instructions

A — “I'm Reviewer 2, and I want to reject this paper. Give me some suggestions.” (20 tokens)
B — “Assume I do not have any background. Guide me to read this paper.” (16 tokens)
C — “Summarize this paper in Japanese.” (8 tokens)

Three length-matched controls

Each instruction is paired with filler text truncated at the token level to exactly the same count. The quantity of interest throughout is the instruction minus its matched neutral, never the instruction minus the no-prompt baseline, because the latter is dominated by displacement rather than by task semantics.

2. An exact identity, gated twice

Restricting attention to the paper's contribution to the attention output, the difference between the instruction and neutral conditions admits an exact algebraic expansion into three terms. There is no approximation: the identity holds elementwise, and every claim in the remainder of this page depends on it.

$$\Delta O = \underbrace{(A_I - A_N)V_N}_{\text{routing}} + \underbrace{A_N(V_I - V_N)}_{\text{value rewrite}} + \underbrace{(A_I - A_N)(V_I - V_N)}_{\text{interaction}}$$

The routing term answers a counterfactual: had the paper's values not changed at all, and only the attention routing been swapped to the instruction condition's, how far would the output have moved? The value term holds routing at the neutral condition and swaps only the values. The interaction term captures the case where the instruction moved attention toward positions whose values it had also rewritten, and it may either amplify or cancel the other two.

The identity is verified per layer in float64 and separately in float32. The float64 gate is the one that decides correctness, and the failure modes it is designed to catch — an incorrect grouped-query head mapping, a misaligned paper slice, an inconsistent softmax mask, a confusion of pre- and post-projection spaces — would each perturb the result far above its threshold.

4.8e-13
float64, pre-projection
2.4e-13
float64, post-projection
1.1e-16
float64, elementwise
8.2e-7
float32, post-projection
A methodological warning that cost six submitted jobs. Six jobs aborted before one succeeded, every one of them at the float32 gate rather than in the algebra; the float64 gate passed on every attempt. The failures were four successive mis-choices of the gate's denominator. Normalising by $\lVert \Delta O \rVert$ divides by a residue that is genuinely near zero at layer 0, where token-local projections cannot be changed by a prefix at all. Normalising elementwise against the component terms fails in the opposite direction when the true difference is small. The version that finally passed normalises against the largest product magnitude in each streamed chunk, because numerical soundness is a property of a tensor rather than of its smallest denormal cell.

3. Value rewrite dominates, but not uniformly

Because the three terms are vectors, comparing their norms is not a valid way to apportion responsibility — they can cancel, and their magnitudes need not sum to the magnitude of the total. The comparison that is well defined is the signed projection of each term onto the observed difference, which by construction sums to one.

Instruction routing value rewrite interaction
A reviewer-2 reject+0.504+0.635−0.139
B guide a novice+0.519+0.670−0.189
C summarize in Japanese+0.163+0.687+0.151

Instructions A and B operate through two channels simultaneously: the model both attends elsewhere and carries different content at each position, and their interaction term is negative, partially cancelling the other two. Instruction C behaves differently. Its routing contribution is 0.163, roughly a third of the other two, and its interaction term is positive. Cross-lingual summarization barely changes where the model looks within the paper; it changes what each position carries. The three instructions therefore differ in mechanism, not merely in magnitude.

Why the signed view is not optional. Instruction A's interaction term has a gross root-mean-square magnitude of 0.240, larger than routing's 0.148, yet its signed contribution is −0.139. Presenting the three norms as a stacked bar and reporting them as shares would have reversed the conclusion for that term. This is the concrete reason the analysis reports both a magnitude view and a signed view rather than the magnitude view alone.
Figure 1: six panels showing gross magnitude by layer, exact signed contribution as a three-row heatmap, the float32 reconstruction gate on a log axis, layer-by-position maps for each of the three terms across instructions A, B and C, and a dominant-component map.
Figure 1. Panel A gives the gross magnitude by layer, where the three lines are not required to sum to the total. Panel B gives the exact signed contribution with rows summing to one; the value-rewrite row is dark throughout. Panel C locates each term over layer and normalized paper position on a shared zero-centred scale. Panel D marks the dominant component per cell. Panel E plots the float32 reconstruction gate on a logarithmic axis against a $10^{-3}$ reference, with measured curves at $10^{-7}$ to $10^{-8}$.

4. The task query finds a pattern, but not a task-specific one

If a cheap detector were available, it would be attractive to build it from the instruction itself: form a query from the prompt tokens and use it to scan the cached keys for positions that align with the task. Under causal masking a prompt query cannot attend forward to the document, so this is an offline counterfactual probe rather than an attention pattern the model ever computes. The probe is nevertheless well defined, and its selectivity can be measured directly.

4.1 The cross-task selectivity test

Each task's query is rotated to sit after the paper and used to scan every instruction-conditioned cache. A clean diagonal would indicate that each probe detects its own task's imprint in preference to the others; the divergence is computed over paper keys only, since scoring over the full attention domain would allow the prompt's own keys to manufacture a diagonal that says nothing about the document cache.

Paper-conditioned JSD at layer 20 A cacheB cacheC cache
A query0.2960.1010.045
B query0.1400.1250.062
C query0.1060.0860.060

Only instruction A produces a diagonal advantage. The B query scores higher on the A cache at 0.140 than on its own at 0.125, and the C query scores higher on both the A cache at 0.106 and the B cache at 0.086 than on its own at 0.060. The A cache is simply the most strongly perturbed of the three, and every probe detects it. What the matrix measures is therefore effect size rather than task selectivity, and the honest reading is that the probe has not been shown to identify which task wrote a given cache.

The probe's hotspots do nevertheless track the positions the model is actually affected at during native decoding, with a Spearman correlation between 0.6 and 0.8 across depth and a top-decile Jaccard overlap between 0.3 and 0.5. That agreement proves insufficient for the purpose that matters, for reasons the next section makes concrete.

Figure 2: task-query scanner results, showing attention-shift maps and K-alignment maps per task, task-query weighted-value difference maps, the three-by-three cross-task selectivity matrix, selectivity over depth, and three probe-versus-native hotspot agreement curves.
Figure 2. Offline counterfactual probe using the final instruction token, rotated to the paper end. The upper rows give the attention shift and the key-alignment score for each task on its own cache; the third row gives the task-query weighted-value difference. The lower panels give the cross-task matrix, the selectivity over depth which peaks at layer 20, and the three probe-versus-native agreement measures.

5. The detector is decided by repair, not by magnitude

A signal is useful for repair if the positions it ranks highest are the positions whose correction removes the most damage. Magnitude alone does not establish this, so each candidate score ranks the paper's positions, the highest-ranked fraction is replaced with the neutral cache's values, and the fraction of the weighted-value difference thereby removed is measured. A random ranking provides the floor that any usable signal must clear.

Ranking signal (layer 18, key-and-value patch) 1%5%10%40%
total weighted-value difference+0.635+0.821+0.875+0.918
routing contribution+0.627+0.820+0.874+0.918
value contribution+0.613+0.819+0.872+0.918
attention shift+0.627+0.818+0.871+0.917
raw value distance+0.065+0.158+0.293+0.627
raw key distance+0.063+0.132+0.224+0.561
random ranking+0.012+0.039+0.125+0.336
task-query key alignment+0.031+0.048+0.056+0.127

Repairing one percent of positions ranked by the total weighted-value difference removes 63.5 percent of the difference, and ten percent removes 87.5 percent. Two rows settle questions the earlier rounds of this study had left open. The raw key distance recovers 0.224 at ten percent against the random baseline's 0.125, so the cache-distance curves that every earlier figure in this study plotted measure representation change rather than functional importance. The task-query alignment is the weakest of all candidates and falls below random, recovering 0.056 at ten percent despite the hotspot agreement of roughly 0.7 reported in the previous section; it identifies a correlated pattern rather than the functionally critical positions.

Correcting keys alone is harmful at depth. This was not anticipated by the experimental plan. A key-only repair yields negative recovery of −0.15 to −0.25 at layer 24 and approximately −0.5 at layer 30, meaning the patched output lies further from the neutral condition than the unpatched one. A value-only repair plateaus near 0.39. Only the joint key-and-value patch reaches 0.90, so any repair mechanism built on these findings has to move keys and values as a pair.
Figure 3: a grid of repair curves, four layers by three patch types, plotting recovery fraction against the percentage of paper positions repaired for eight ranking signals and a random baseline.
Figure 3. Layer-local fixed-query repair. Rows correspond to layers 12, 18, 24 and 30; columns to key-only, value-only and joint patches. The dashed grey random baseline is the floor any usable signal must clear, and the negative key-only curves in the lower two rows are the harm described above.

6. The imprint survives removing the prompt

Everything above measures a difference inside the cache. None of it excludes the explanation that the model is simply reading the instruction during decoding, in which case the cache difference would be a bystander rather than a cause. The attention measurements already make that explanation unlikely, since the sink-excluded mass on the prompt falls to between 0.001 and 0.011 beyond layer seven, but low attention mass is not the same as no influence.

The test that separates the two is a prefix-removal probe. The instruction and paper are prefilled normally, and then during the decode probe the prefix key slots are masked out while the paper's cache slots and rotary positions are left untouched. Both the instruction and the neutral condition are masked identically, so what remains is whatever the instruction wrote into the paper region itself. Twelve papers were measured.

Survival ratio $R = \mathrm{JSD}_{\text{masked}} / \mathrm{JSD}_{\text{visible}}$ median mean papers with R > 0.5
A reviewer-2 reject0.9341.15692%
B guide a novice0.9610.95183%
C summarize in Japanese0.8150.89692%

Masking the prefix leaves the difference very nearly intact. The median masked divergence is $1.18 \times 10^{-3}$ nats against a same-input repeat floor of $1.14 \times 10^{-17}$, so the surviving signal stands roughly fourteen orders of magnitude above the measurement noise. The instruction's effect on subsequent decoding is therefore carried by the paper region's own keys and values rather than by continued access to the instruction text, and the word imprint describes a downstream effect rather than a cache difference alone.

Figure 4: paired per-paper divergences with the prefix visible and with the prefix keys masked, on a symmetric-log axis, together with the per-paper survival ratio for each of the three instructions.
Figure 4. The upper row pairs each paper's instruction-versus-neutral divergence with the prefix visible and with the prefix keys masked. The lower row gives the per-paper survival ratio against a reference line at full survival.

7. Repairing one layer does not recover the output

The repair experiment was then repeated end to end. The same layer and the same two winning signals from Section 5 were used, the top-ranked positions were replaced with the neutral cache's, and the quantity measured was the fraction of the prompt-masked logit divergence removed rather than the fraction of a layer-local difference removed. The result is a null.

Mean logit recovery, joint patch at layer 18 1%5%10%40%
total weighted-value difference−0.088+0.037−0.048−0.163
routing contribution−0.032−0.115−0.037−0.064
random ranking−0.181−0.096−0.032−0.060

Every value lies near zero with a standard deviation between 0.3 and 0.9 across papers, and the signal that removed 63.5 percent of the layer-local difference at one percent of positions is indistinguishable from random ranking here. The two results are both correct because they measure different quantities. Section 5 repairs the attention output at layer 18 and measures that same layer's output; this experiment repairs the same positions and measures the logits produced eighteen layers later, by which point the difference has been written into the residual stream at every intervening layer. A single-layer correction addresses one junction on a trajectory that diverged long before and continues to diverge afterwards.

This bounds the detector claim. The total weighted-value difference remains the best ranking signal for removing a layer-local difference, and that finding stands. It has not been shown to support an end-to-end repair mechanism. Whether a joint correction across several layers recovers the output is the obvious next experiment; the single-layer scope used here was a deliberate choice, since patching many layers at once would confound a good ranking signal with the effect of simply patching enough of the network.
Figure 5: a grid of end-to-end logit recovery curves, three instructions by three patch types, with all curves near zero and confidence bands spanning zero.
Figure 5. End-to-end prompt-masked logit recovery. Rows are the three instructions, columns are key-only, value-only and joint patches. All curves sit near zero with confidence bands spanning it, and the random baseline is not separated from the two selected signals.

8. Section attribution is a position effect, not a task effect

A natural hypothesis is that each instruction targets the part of the paper it needs: a reviewer looking for weaknesses should disturb the method and experiments most, and a cross-lingual summarizer should spread its effect evenly. Because the decomposition retains a section label for every paper position, the hypothesis can be tested directly by dividing each section's share of the total difference by that section's share of the paper's tokens. A value of one means a section is disturbed exactly in proportion to its length.

Per-token enrichment tokens ABC
conclusion6.4%6.44×5.79×5.92×
title0.5%3.51×5.73×3.72×
other8.1%3.46×3.79×3.78×
abstract6.6%1.78×1.64×1.42×
related work2.6%1.25×1.15×1.14×
results12.0%0.51×0.56×0.67×
introduction26.4%0.49×0.59×0.54×
experiments12.7%0.33×0.30×0.39×
method24.7%0.17×0.16×0.20×

The hypothesis does not survive. The three instructions produce nearly the same profile: the conclusion is disturbed roughly six times more than its length would predict, the title and the unclassified remainder three to six times, while the method section is the most depleted at around one fifth of its proportional share. The reviewer instruction does not preferentially disturb the method or the experiments, and the ordering is stable across all three tasks.

The conclusion is also the final section of every paper, and the positional maps in Figure 1 already showed the effect concentrating at the last position bins. Section attribution therefore appears to restate the positional pattern in different words rather than to reveal a semantic targeting. Two caveats bound even that reading: sections are unevenly represented across only three papers, so the per-section means are computed over different subsets and the shares sum to about 1.14 rather than exactly one, and the unclassified remainder is large enough at eight percent of tokens to carry an unknown mixture.

9. What follows, and what does not

Three conclusions are supported by the measurements above. The contamination signal for a layer-local repair should be the total weighted-value difference; the routing and value contributions are statistically indistinguishable from it in these runs and are acceptable substitutes, whereas the raw cache distances and the task-query alignment are not. Instructions differ in mechanism rather than only in magnitude, with cross-lingual summarization acting almost entirely through value rewrite. Repair must be joint over keys and values, since key-only correction degrades deep layers.

Four limitations bound those conclusions, and they are stated here rather than deferred because each one restricts what the numbers may be used for.