A task instruction placed before a scientific paper rewrites that paper's cached keys and values. This deepdive decomposes the rewrite exactly into routing, value rewrite and interaction, then asks the question that actually decides an engineering signal: rank cache positions by each candidate score, repair the top few percent, and measure which score removes the damage fastest.
Serving systems increasingly reuse a document's KV cache across requests. A paper, a codebase, or a contract is prefilled once and the resulting cache is retained so that later queries need not pay for the prefill again. That reuse rests on an assumption that is rarely stated: the cached representation of the document is a property of the document. Once a task instruction is placed in front of the document, the assumption no longer holds, because every document token attends to that instruction during prefill and carries its influence forward.
A concrete instance makes the size of the effect visible. A 1,782-token QASPER paper is prefilled twice under identical tokenization: once preceded by the instruction “I'm Reviewer 2, and I want to reject this paper”, and once preceded by twenty tokens of neutral filler carrying no task at all. At layer 18 the paper region's keys differ from the no-prompt baseline by a relative L2 of 0.283 in the first case and 0.155 in the second. Roughly half of the apparent contamination is nothing more than the fact that twenty tokens now precede the paper and displace its positions; the remainder is what the instruction actually did.
If a system wished to repair such a cache, the natural first move would be to measure how far each cached position has moved and to correct the positions that moved the most. The distance $\lVert K_I(j) - K_N(j) \rVert$ is cheap to compute and is exactly what the earlier rounds of this study plotted. The results in Section 5 show that ranking by that distance is barely distinguishable from ranking positions at random. Distance in representation space and functional importance turn out to be different quantities, and the study needed a decomposition to see why.
Three QASPER papers are each prefilled under seven conditions on Qwen3-8B at a pinned revision. The paper's token identifiers are byte-identical in every condition, which is enforced by assembling the input at the identifier level rather than tokenizing a concatenated string. A single shared newline anchor precedes every condition, including the baseline, so that the attention sink occupies the same position everywhere and cancels in every comparison.
A — “I'm Reviewer 2, and I want to reject this paper. Give me some suggestions.” (20 tokens)
B — “Assume I do not have any background. Guide me to read this paper.” (16 tokens)
C — “Summarize this paper in Japanese.” (8 tokens)
Each instruction is paired with filler text truncated at the token level to exactly the same count. The quantity of interest throughout is the instruction minus its matched neutral, never the instruction minus the no-prompt baseline, because the latter is dominated by displacement rather than by task semantics.
Restricting attention to the paper's contribution to the attention output, the difference between the instruction and neutral conditions admits an exact algebraic expansion into three terms. There is no approximation: the identity holds elementwise, and every claim in the remainder of this page depends on it.
The routing term answers a counterfactual: had the paper's values not changed at all, and only the attention routing been swapped to the instruction condition's, how far would the output have moved? The value term holds routing at the neutral condition and swaps only the values. The interaction term captures the case where the instruction moved attention toward positions whose values it had also rewritten, and it may either amplify or cancel the other two.
The identity is verified per layer in float64 and separately in float32. The float64 gate is the one that decides correctness, and the failure modes it is designed to catch — an incorrect grouped-query head mapping, a misaligned paper slice, an inconsistent softmax mask, a confusion of pre- and post-projection spaces — would each perturb the result far above its threshold.
Because the three terms are vectors, comparing their norms is not a valid way to apportion responsibility — they can cancel, and their magnitudes need not sum to the magnitude of the total. The comparison that is well defined is the signed projection of each term onto the observed difference, which by construction sums to one.
| Instruction | routing | value rewrite | interaction |
|---|---|---|---|
| A reviewer-2 reject | +0.504 | +0.635 | −0.139 |
| B guide a novice | +0.519 | +0.670 | −0.189 |
| C summarize in Japanese | +0.163 | +0.687 | +0.151 |
Instructions A and B operate through two channels simultaneously: the model both attends elsewhere and carries different content at each position, and their interaction term is negative, partially cancelling the other two. Instruction C behaves differently. Its routing contribution is 0.163, roughly a third of the other two, and its interaction term is positive. Cross-lingual summarization barely changes where the model looks within the paper; it changes what each position carries. The three instructions therefore differ in mechanism, not merely in magnitude.
If a cheap detector were available, it would be attractive to build it from the instruction itself: form a query from the prompt tokens and use it to scan the cached keys for positions that align with the task. Under causal masking a prompt query cannot attend forward to the document, so this is an offline counterfactual probe rather than an attention pattern the model ever computes. The probe is nevertheless well defined, and its selectivity can be measured directly.
Each task's query is rotated to sit after the paper and used to scan every instruction-conditioned cache. A clean diagonal would indicate that each probe detects its own task's imprint in preference to the others; the divergence is computed over paper keys only, since scoring over the full attention domain would allow the prompt's own keys to manufacture a diagonal that says nothing about the document cache.
| Paper-conditioned JSD at layer 20 | A cache | B cache | C cache |
|---|---|---|---|
| A query | 0.296 | 0.101 | 0.045 |
| B query | 0.140 | 0.125 | 0.062 |
| C query | 0.106 | 0.086 | 0.060 |
Only instruction A produces a diagonal advantage. The B query scores higher on the A cache at 0.140 than on its own at 0.125, and the C query scores higher on both the A cache at 0.106 and the B cache at 0.086 than on its own at 0.060. The A cache is simply the most strongly perturbed of the three, and every probe detects it. What the matrix measures is therefore effect size rather than task selectivity, and the honest reading is that the probe has not been shown to identify which task wrote a given cache.
The probe's hotspots do nevertheless track the positions the model is actually affected at during native decoding, with a Spearman correlation between 0.6 and 0.8 across depth and a top-decile Jaccard overlap between 0.3 and 0.5. That agreement proves insufficient for the purpose that matters, for reasons the next section makes concrete.
A signal is useful for repair if the positions it ranks highest are the positions whose correction removes the most damage. Magnitude alone does not establish this, so each candidate score ranks the paper's positions, the highest-ranked fraction is replaced with the neutral cache's values, and the fraction of the weighted-value difference thereby removed is measured. A random ranking provides the floor that any usable signal must clear.
| Ranking signal (layer 18, key-and-value patch) | 1% | 5% | 10% | 40% |
|---|---|---|---|---|
| total weighted-value difference | +0.635 | +0.821 | +0.875 | +0.918 |
| routing contribution | +0.627 | +0.820 | +0.874 | +0.918 |
| value contribution | +0.613 | +0.819 | +0.872 | +0.918 |
| attention shift | +0.627 | +0.818 | +0.871 | +0.917 |
| raw value distance | +0.065 | +0.158 | +0.293 | +0.627 |
| raw key distance | +0.063 | +0.132 | +0.224 | +0.561 |
| random ranking | +0.012 | +0.039 | +0.125 | +0.336 |
| task-query key alignment | +0.031 | +0.048 | +0.056 | +0.127 |
Repairing one percent of positions ranked by the total weighted-value difference removes 63.5 percent of the difference, and ten percent removes 87.5 percent. Two rows settle questions the earlier rounds of this study had left open. The raw key distance recovers 0.224 at ten percent against the random baseline's 0.125, so the cache-distance curves that every earlier figure in this study plotted measure representation change rather than functional importance. The task-query alignment is the weakest of all candidates and falls below random, recovering 0.056 at ten percent despite the hotspot agreement of roughly 0.7 reported in the previous section; it identifies a correlated pattern rather than the functionally critical positions.
Everything above measures a difference inside the cache. None of it excludes the explanation that the model is simply reading the instruction during decoding, in which case the cache difference would be a bystander rather than a cause. The attention measurements already make that explanation unlikely, since the sink-excluded mass on the prompt falls to between 0.001 and 0.011 beyond layer seven, but low attention mass is not the same as no influence.
The test that separates the two is a prefix-removal probe. The instruction and paper are prefilled normally, and then during the decode probe the prefix key slots are masked out while the paper's cache slots and rotary positions are left untouched. Both the instruction and the neutral condition are masked identically, so what remains is whatever the instruction wrote into the paper region itself. Twelve papers were measured.
| Survival ratio $R = \mathrm{JSD}_{\text{masked}} / \mathrm{JSD}_{\text{visible}}$ | median | mean | papers with R > 0.5 |
|---|---|---|---|
| A reviewer-2 reject | 0.934 | 1.156 | 92% |
| B guide a novice | 0.961 | 0.951 | 83% |
| C summarize in Japanese | 0.815 | 0.896 | 92% |
Masking the prefix leaves the difference very nearly intact. The median masked divergence is $1.18 \times 10^{-3}$ nats against a same-input repeat floor of $1.14 \times 10^{-17}$, so the surviving signal stands roughly fourteen orders of magnitude above the measurement noise. The instruction's effect on subsequent decoding is therefore carried by the paper region's own keys and values rather than by continued access to the instruction text, and the word imprint describes a downstream effect rather than a cache difference alone.
The repair experiment was then repeated end to end. The same layer and the same two winning signals from Section 5 were used, the top-ranked positions were replaced with the neutral cache's, and the quantity measured was the fraction of the prompt-masked logit divergence removed rather than the fraction of a layer-local difference removed. The result is a null.
| Mean logit recovery, joint patch at layer 18 | 1% | 5% | 10% | 40% |
|---|---|---|---|---|
| total weighted-value difference | −0.088 | +0.037 | −0.048 | −0.163 |
| routing contribution | −0.032 | −0.115 | −0.037 | −0.064 |
| random ranking | −0.181 | −0.096 | −0.032 | −0.060 |
Every value lies near zero with a standard deviation between 0.3 and 0.9 across papers, and the signal that removed 63.5 percent of the layer-local difference at one percent of positions is indistinguishable from random ranking here. The two results are both correct because they measure different quantities. Section 5 repairs the attention output at layer 18 and measures that same layer's output; this experiment repairs the same positions and measures the logits produced eighteen layers later, by which point the difference has been written into the residual stream at every intervening layer. A single-layer correction addresses one junction on a trajectory that diverged long before and continues to diverge afterwards.
A natural hypothesis is that each instruction targets the part of the paper it needs: a reviewer looking for weaknesses should disturb the method and experiments most, and a cross-lingual summarizer should spread its effect evenly. Because the decomposition retains a section label for every paper position, the hypothesis can be tested directly by dividing each section's share of the total difference by that section's share of the paper's tokens. A value of one means a section is disturbed exactly in proportion to its length.
| Per-token enrichment | tokens | A | B | C |
|---|---|---|---|---|
| conclusion | 6.4% | 6.44× | 5.79× | 5.92× |
| title | 0.5% | 3.51× | 5.73× | 3.72× |
| other | 8.1% | 3.46× | 3.79× | 3.78× |
| abstract | 6.6% | 1.78× | 1.64× | 1.42× |
| related work | 2.6% | 1.25× | 1.15× | 1.14× |
| results | 12.0% | 0.51× | 0.56× | 0.67× |
| introduction | 26.4% | 0.49× | 0.59× | 0.54× |
| experiments | 12.7% | 0.33× | 0.30× | 0.39× |
| method | 24.7% | 0.17× | 0.16× | 0.20× |
The hypothesis does not survive. The three instructions produce nearly the same profile: the conclusion is disturbed roughly six times more than its length would predict, the title and the unclassified remainder three to six times, while the method section is the most depleted at around one fifth of its proportional share. The reviewer instruction does not preferentially disturb the method or the experiments, and the ordering is stable across all three tasks.
The conclusion is also the final section of every paper, and the positional maps in Figure 1 already showed the effect concentrating at the last position bins. Section attribution therefore appears to restate the positional pattern in different words rather than to reveal a semantic targeting. Two caveats bound even that reading: sections are unevenly represented across only three papers, so the per-section means are computed over different subsets and the shares sum to about 1.14 rather than exactly one, and the unclassified remainder is large enough at eight percent of tokens to carry an unknown mixture.
Three conclusions are supported by the measurements above. The contamination signal for a layer-local repair should be the total weighted-value difference; the routing and value contributions are statistically indistinguishable from it in these runs and are acceptable substitutes, whereas the raw cache distances and the task-query alignment are not. Instructions differ in mechanism rather than only in magnitude, with cross-lingual summarization acting almost entirely through value rewrite. Repair must be joint over keys and values, since key-only correction degrades deep layers.
Four limitations bound those conclusions, and they are stated here rather than deferred because each one restricts what the numbers may be used for.