Experiment Deepdive #19

How Deep Must a KV Cache Be Repaired?

Round 5 replaces the repair target with the one that cache reuse actually requires — the paper-only baseline — and measures how much of the network must be patched before a document's cache behaves as though the instruction had never preceded it. Spans of eight layers or fewer recover essentially nothing; recovery arrives only at eighteen contiguous layers, half the stack. A complete restoration of the document region appeared to leave a residual; Round 6 traced that appearance to the evaluation itself — under prompt-masked evaluation, the definition cache reuse actually requires, the same complete patch restores exactly 1.000.

Prompt-conditioned KV cache study · round 5 · Qwen3-8B at revision b968826d

Qwen3-8B QASPER 36 layers · 3 instructions H200 · 8 min 54 s

1. The repair target was measuring the wrong quantity

The predecessor of this page established that a task instruction placed before a scientific paper rewrites that paper's cached keys and values, decomposed the rewrite exactly into routing, value rewrite and interaction, and showed that the imprint survives masking the instruction's keys during decoding. Every repair curve in that study was measured against a length-matched neutral prefix, and that choice quietly restricts what the curves are entitled to mean. Deepdive #18 covers that ground and is assumed here.

A length-matched neutral control is the correct instrument for isolating task semantics. The instruction and its control occupy the same positions and the same token count, so the difference between the two caches contains only the content of the instruction and nothing of the fact that some prefix was present at all. Repair measured against that control therefore answers a well-posed scientific question: how much of the task-semantic difference has been removed.

The engineering goal is a different question, and it is the question this series exists to answer. A serving stack that caches a document wants a cache that behaves as though only the shared anchor and the document itself had been prefilled, because that is the state any later request would have reused. Call that condition the paper-only baseline, which the figures on this page label P0. A cache repaired toward the neutral condition still carries the entire generic-prefix component and is still not the document's own. Round 5 therefore re-runs the repair experiments with the paper-only baseline as the target, retaining the neutral target as a comparison arm rather than discarding it.

What changes when the target changes. Under the neutral target the repair problem is geometrically trivial in one respect: the control is length-matched, so the document occupies identical absolute positions in the two caches and a patch is a plain copy. Under the paper-only baseline the document sits at a different absolute offset, and that mismatch has to be corrected before any patch is meaningful at all. The correction is the subject of the next section, and it is the reason this round required new code rather than a changed flag.

2. The patch source must be re-rotated before it is copied

Under the paper-only baseline the document begins immediately after the shared anchor. Under an instruction it begins after the anchor and the instruction together, so every document token sits further along the sequence by exactly the instruction's token length. Rotary position embedding writes absolute position into the key by rotating it, which means the key the baseline stores for a document token is rotated by an angle appropriate to a position that token does not occupy in the instruction-conditioned layout.

Copying those keys unmodified would install a correct content vector at a wrong angle, producing a state that is neither the instruction's nor the baseline's. The patch source is therefore the baseline key advanced by the instruction's length in rotary position, while values carry no positional rotation and are copied directly.

$$\tilde{K}_{p} \;=\; R(\ell_{\text{pre}})\, K^{P_0}_{p}, \qquad \tilde{V}_{p} \;=\; V^{P_0}_{p}$$

Here $K^{P_0}_p$ and $V^{P_0}_p$ are the key and value that the paper-only baseline stores at document position $p$, $R(\ell_{\text{pre}})$ is the rotary rotation that advances a key by $\ell_{\text{pre}}$ positions, and $\ell_{\text{pre}}$ is the token length of the instruction. Reversing the sign of that shift produces a patch that is wrong in a way no monotonic check would catch, because a wrongly rotated key still moves the layer output away from the instruction's state.

The guard against that failure is therefore stated as an equality rather than as an improvement. Patching every document position at a layer must make that layer's output equal the baseline's output, not merely bring it closer, and the test fails on either sign convention that does not satisfy the equality. A second cross-check protects the partial-patch path, in which only a subset of positions is replaced: a fresh contiguous prefill of the anchor and the document is run through the same decode, and this canonical rebuild reaches a recovery of exactly 1.000 on all four measured endpoints by construction. It is the reference against which every partial patch in this page is read.

Agreement is measured, not assumed. The canonical rebuild and a full patch of the document region are not asserted to coincide. Section 5 reports the gap between them, and that gap turns out to be one of the results rather than an artefact to be tuned away.

3. Recovery arrives only at eighteen contiguous layers

The sweep ranks document positions once, at layer 18, by the total weighted-value difference, which was the ranking that won the detector comparison in the predecessor study. It then takes the top five percent of those positions and patches them toward the re-rotated baseline cache across a contiguous span of layers ending at the deepest block, sweeping the span length over one, two, four, eight, eighteen and thirty-six layers. The endpoint is the signed end-to-end recovery of the next-token distribution.

A random ranking over the same position budget and the same spans is run alongside every configuration. This separates two questions that are easy to conflate: whether the ranking selects the right positions, and whether the span reaches deep enough for the choice of positions to matter at all.

The table below gives the signed logit recovery for the selected ranking at each span length, averaged over the three papers.

Contiguous patched layers · top 5% of positions 12481836
A reviewer-2 reject+0.166−0.124+0.115+0.054+0.748+0.887
B guide a novice+0.120−0.140+0.130+0.118+0.729+0.688
C summarize in Japanese−0.195−0.217−0.307+0.145+0.252+0.395

Spans of eight layers or fewer achieve essentially nothing. The entries in those four columns are small, they change sign between adjacent span lengths and between instructions, and they are not separated from the random ranking, whose entries over the same cells range from −0.247 to +0.156. This extends the single-layer null of the predecessor study: the failure was not specific to one layer but holds for any shallow suffix of the network.

Recovery appears somewhere between eight and eighteen layers, and it is large when it appears. At eighteen layers, half of the thirty-six-block network, the selected ranking recovers 0.748 and 0.729 for the two English instructions while the random ranking over the same positions recovers 0.011 and 0.121, so both the depth and the choice of positions carry weight at that scale. At the full thirty-six layers the separation widens further: the random baseline reaches 0.081, 0.219 and 0.084 against the selected ranking's 0.887, 0.688 and 0.395.

The cross-lingual instruction is the hardest to repair by a wide margin. Instruction C is negative at every span up to four layers, reaches only 0.252 at eighteen layers and 0.395 with every layer patched. The comparison with a full position budget locates the constraint: at full depth, patching every document position recovers 0.914 for A against the 0.887 that five percent of positions already achieves, so for A the sparse budget nearly saturates what depth makes available, whereas for C the same comparison is 0.810 against 0.395, so a cross-lingual instruction writes its imprint across positions that a five percent budget does not cover.

Suffix-span restoration sweep
Figure 1. Suffix-span restoration toward the paper-only baseline at the top five percent of positions, one panel per instruction. The purple curve is the end-to-end logit recovery and the dashed curves are the matched random rankings; the three remaining solid curves are intermediate endpoints measured in the same run, namely the local attention output, the post-block hidden state and the final hidden state. All four endpoints sit near zero out to eight layers and the logit curve separates sharply at eighteen. The intermediate endpoints rise more modestly, reaching 0.62, 0.54 and 0.60 for A at full depth, 0.34, 0.38 and 0.39 for B, and 0.11, 0.17 and 0.13 for C.
The imprint is written deeply and diffusely. There is no shallow suffix of the network that holds enough of the damage to be worth repairing, which removes the cheapest engineering option from the table. Any repair that works has to touch roughly half the stack, and the cost of such a repair is close to the cost of the prefill it was meant to avoid.

4. The harder target is not measurably harder

Retaining the neutral target as a comparison arm makes it possible to ask whether the change of target changed the difficulty. If the generic component of a prefix, the part that is present regardless of what the prefix says, contributed substantially to the damage, then repairing toward the paper-only baseline would be measurably harder than repairing toward a length-matched control at every depth.

The comparison is only meaningful where recovery exists, so the table pairs the two targets at the two span lengths that section 3 identified as informative.

Logit recovery · top 5% of positions baseline, 18 layers baseline, 36 layers neutral, 18 layers neutral, 36 layers
A reviewer-2 reject+0.748+0.887+0.725+0.930
B guide a novice+0.729+0.688+0.477+0.708
C summarize in Japanese+0.252+0.395+0.226+0.409

At these depths the two targets are of comparable difficulty. Instructions A and C are nearly indistinguishable across the pair, differing by 0.023 and 0.026 at eighteen layers and by 0.043 and 0.014 at thirty-six. Instruction B at eighteen layers is the only cell where the arms differ by more than a tenth, and it differs in the direction that makes the paper-only baseline the easier target rather than the harder one. The generic component of a prefix therefore contributes little beyond what the task-semantic component already contributes, at least once the patched span is deep enough for recovery to exist at all.

At shallow spans the two arms disagree in ways that should not be interpreted. The neutral arm reports +0.268 for A at eight layers where the baseline arm reports +0.054, and it reports −0.546 for B at four layers; the random baselines in the same region swing between −0.641 and +0.235. These are the same shallow-span numbers that section 3 declined to read, and the disagreement between targets is further evidence that nothing stable is being measured there.

5. The full-patch residual belonged to the measurement, not the cache

The sweep in section 3 patches a sparse subset of positions, so its ceiling could always be explained by the budget. Replacing every key and value cell of the document region, at every one of the thirty-six layers, asks a different question: what does a document-cache patch fail to reach even when nothing about the document's cache is left unpatched.

The answer is that it falls short, and by an amount that is stable enough across endpoints to be structural rather than incidental. The table reports the full-document patch on all four endpoints, with the canonical rebuild of section 2 as the reference.

100% document patch · 36 layers · baseline target local attention post-block hidden final hidden logit
A reviewer-2 reject0.8740.9110.8750.914
B guide a novice0.7780.7990.8240.876
C summarize in Japanese0.8410.8230.8860.810
canonical rebuild1.0001.0001.0001.000
Round-6 verification. On the Round-6 canonical set (Qwen3-8B, twelve calibration papers; four papers × forty-eight structured-repair cases), the 100% document patch evaluated with the instruction masked at every layer — the evaluation that matches the reuse objective, and one Round 6 verified to be functionally identical to physically deleting the instruction slots and re-rotating — recovers logit fidelity of exactly 1.000 for every prompt, every geometry S1/S2/S3 and every selector. The table above is retained as the Round-5 measurement under an instruction-conditioned evaluation.

Round 6 resolved where this residual lives. The candidate explanations were two: the instruction's cache slots are masked from attention rather than absent, and the probe query that drives the decode was formed under the instruction. The first contributes nothing — masking a key at every layer is functionally identical to deleting the slot, and Round 6 verified the masked evaluation against the physical rebuild. What produced the numbers above is the second: Round 5 decoded with an instruction-conditioned query. Under Round 6's prompt-masked evaluation, in which the probe never attends the instruction at any layer, the same one-hundred-percent document patch reaches 1.000 exactly, on every prompt, geometry and selector of the canonical set.

The neutral arm had already hinted at this: with a length-matched control the geometry matches, masking the prefix slots suffices, and the same full patch reached 1.000. Round 6 closes the question this run left open. The entire residual was the evaluation's, and it vanishes once the evaluation stops attending the instruction; the Round-5 absolute recovery numbers on this page are correspondingly deflated lower bounds, while their qualitative shapes — the eighteen-layer threshold, the head-diffuse imprint, the cross-lingual ordering — survive the correction.

Full-paper partial patch versus canonical rebuild
Figure 2. Full-document partial patch against the canonical rebuild, one panel per instruction and one marker per endpoint. The left column is the intervention that replaces every document key and value at every layer; the right column is a fresh contiguous prefill of the anchor and the document. The four endpoints cluster between 0.78 and 0.92 on the left and coincide at 1.000 on the right, and the vertical gap between the two columns is the part of the difference that lies outside the document region.
Perfect cache restoration is not full restoration. Any repair strategy that operates on a cached document region inherits this ceiling, whatever signal it uses to choose positions and whatever depth it patches. Between eight and nineteen percent of the end-to-end difference lives in the request rather than in the document, and only rebuilding the request removes it.

6. No key-value head carries the imprint

Depth is not the only dimension along which a repair could be made cheap. If the imprint were concentrated in a small number of attention heads, a strategy could patch every layer but only the heads that matter, and the cost would fall by the head ratio rather than the layer ratio. The run tests this directly by patching a single key-value head at a single layer, holding the position budget at five percent and the layer at eighteen.

Qwen3-8B uses grouped-query attention with eight key-value heads, so eight single-head patches per instruction exhaust the head dimension at that layer. The table gives the per-head end-to-end logit recovery, averaged over the three papers.

Patched KV head · layer 18 · top 5% of positions 01234567
A reviewer-2 reject+0.204+0.173+0.211+0.121+0.072−0.067+0.129+0.129
B guide a novice+0.152+0.181−0.177+0.014+0.011+0.238−0.020+0.036
C summarize in Japanese+0.259−0.281−0.034−0.232−0.223−0.097+0.041+0.081

No head carries the effect. The largest single cell is +0.259, for the cross-lingual instruction at head zero, and the same instruction produces −0.281 at the adjacent head; before averaging, individual documents range from −0.873 to +0.576. The ranking of heads is not stable across instructions either, since head five is the best head for B and among the worst for C, so there is no head subset that a repair strategy could select once and reuse.

This measurement sits against the family of serving methods that reuse key-value state head-selectively, of which the head-aware reuse in RedKnot is a representative example. The claim here is narrow: the prompt-conditioned imprint measured in this setting offers no small set of heads to select. It is not a refutation of head-selective reuse in general, whose head selection is motivated by properties of long-context serving that this experiment does not measure. RedKnot is covered separately in the paper library.

One-head restoration at layer 18
Figure 3. One-head restoration toward the paper-only baseline at layer 18, one panel per instruction and one column of markers per key-value head. The purple diamonds are the end-to-end logit recovery and dominate the vertical spread; the three remaining endpoints stay within roughly a twentieth of zero at every head. No column reaches even a third of the recovery that an eighteen-layer span attains, and the sign of the logit marker changes from head to head within every panel.

7. The largest deviation is not where the output difference is produced

The layer-span result raises a descriptive question that the repair experiments cannot answer on their own, namely where along the stack the deviation actually lives. A separate pass records, for every block and for three points within each block, the relative size of the hidden-state difference between the instruction condition and the paper-only baseline, the angular part of that difference, and the projection of the layer-wise difference onto the direction in which the final states differ.

The three measurements disagree about where the action is, and that disagreement is the content of this section. Magnitude and alignment peak in different parts of the network, which is exactly the situation in which a repair guided by magnitude alone would be misdirected.

Hidden-state propagation across blocks
Figure 4. Hidden-state propagation against the paper-only baseline, one column per instruction. The upper row is the relative norm of the hidden-state difference on a logarithmic axis, the middle row is the angular distance on the same kind of axis, and the lower row is the signed projection of each layer's difference onto the final difference. Within every block the three traces are the block input, the post-attention state and the post-MLP state, which stay close to one another throughout. The upper two rows peak near block nineteen and decline thereafter, whereas the lower row is flat at zero until roughly block twenty.

The relative deviation grows steeply through the first ten blocks, peaks at block nineteen at 0.149 of the target hidden norm for instruction A and at 0.086 and 0.088 for B and C, then declines slowly to 0.086, 0.048 and 0.058 at the last block. The projection onto the final difference behaves differently. It is indistinguishable from zero through block fifteen, reaches only 0.02 at block twenty and 0.09 at block twenty-five, and rises to roughly 0.21 to 0.25 at block thirty before closing at one by construction at the last block.

The mid-stack therefore carries the largest deviation in magnitude while that deviation is nearly orthogonal to the direction in which the final states differ. A repair confined to the last few blocks cannot remove what has already accumulated below it, and the span at which recovery first appears, eighteen layers reaching back to block eighteen, is also the first span that covers the magnitude peak. That coincidence is suggestive rather than demonstrated: the projection is a signed association measured on unperturbed forward passes, and the figure states as much in its own title.

8. Locally functional, not downstream causal

The depth result makes one question urgent: whether any of the cheap signals that can be computed at a single layer predicts downstream restoration at all. The final experiment measures four endpoints for one and the same intervention, patching the top five percent of positions at layer eighteen alone, so that a single patch can be read at increasing distance from the site at which it was applied.

The four endpoints are the post-projection attention output at layer eighteen, the block's output hidden state one step later, the final hidden state after the last normalization, and the next-token distribution. The table averages the first three columns over the three instructions because they are stable, and reports the logit column per instruction because it is not; Figure 5 carries the full per-instruction detail for all four.

Ranking signal · layer 18 · top 5% of positions local attn post-block final hidden logit A logit B logit C
aligned key distance+0.140+0.026+0.011+0.231+0.089−0.336
value distance+0.165+0.031−0.005+0.158+0.052−0.265
attention source contribution+0.262+0.075+0.010+0.188−0.232−0.277
routing+0.245+0.071+0.007+0.111−0.111+0.241
value rewrite+0.275+0.073+0.011+0.322+0.150−0.406
interaction+0.264+0.074−0.007+0.302+0.059−0.388
total weighted-value+0.269+0.073+0.020+0.285+0.162+0.070
random ranking+0.035+0.008+0.000+0.035−0.036+0.108

Every decomposition-derived signal is locally functional. At the attention output of the patched layer they recover between 0.20 and 0.32 per instruction, against a random ranking that recovers between 0.02 and 0.05, which reproduces the detector result of the predecessor study now measured against the paper-only baseline rather than against a neutral control.

The signal then decays with every step away from the intervention. One block later the same patches recover between 0.05 and 0.10; at the final hidden state they recover approximately nothing, with every entry within 0.04 of zero and the ordering among signals no longer meaningful. At the logit the column is not merely small but unstable and instruction-dependent: under the cross-lingual instruction the value-rewrite ranking scores −0.406 and the interaction ranking −0.388 while the random ranking scores +0.108, an ordering that inverts the local result completely.

Four-level signal comparison heatmap
Figure 5. Signed recovery for every ranking signal at four endpoints, one panel per instruction, on a shared zero-centred colour scale with the random ranking placed last in a fixed row order. The first column is uniformly warm for the decomposition-derived signals and pale for random; the second column is uniformly pale; the third is white throughout. The fourth column is the only one that carries deep colour of both signs, and its pattern differs between the three panels.
A cheap signal is not thereby a useless one. The local result is real and the depth result explains why it does not transfer: an intervention confined to one layer is absorbed by the thirty-five blocks that follow it, whatever positions it selected. Whether these rankings predict downstream restoration when they are applied across a deep span is not answered here, because the span sweep of section 3 used only the total weighted-value ranking. That is the measurement the next round has to make.

9. What follows, and what does not

Two conclusions survive this round. The imprint that a task instruction leaves on a document's cache is written deeply and diffusely, since no suffix shorter than half the network and no individual key-value head carries enough of it to be worth repairing. And restoring the document region perfectly is still not restoring the request, because the instruction's own cache slots and the probe query lie outside the region that any document-cache repair can touch.

Four limitations bound those conclusions, and they are stated here rather than deferred because each one restricts what the numbers may be used for.