Round 6 removes the last oracle assumption from the repair problem. A serving stack that wants to restore a contaminated document cache must decide which cells to repair without ever reading the clean cache it is trying to reach. Three information tiers, three repair geometries and a budget sweep from a quarter of a percent to everything measure what that decision actually costs. A single probe forward pass buys most of the information the clean cache would have supplied; a search that queries the model directly buys more, at nine hundred forward evaluations per document; and a learned per-cell predictor buys nothing at all, for a reason the data makes explicit.
Prompt-conditioned KV cache study · round 6 · Qwen3-8B at revision b968826d
The two predecessors of this page established what a task instruction does to a document's cached keys and values, and how much of the network a repair has to touch before that damage comes back out. Both of them repaired with an oracle in hand: the patch source was a clean paper-only cache that the experiment had prefilled for the purpose, and the ranking that chose which positions to patch was computed by comparing the contaminated cache against that same clean cache. A serving stack has neither. If it had the clean cache it would serve from it and there would be nothing to repair. Deepdive #18 and Deepdive #19 cover that ground and are assumed here.
Round 6 separates the two halves of the oracle and retires only one of them. The patch source remains oracle-supplied, because a system that recomputes a document region can obtain the correct values by recomputing it; what a deployed system genuinely cannot obtain in advance is the knowledge of which cells are worth recomputing. This round therefore holds the patch source fixed and varies only the selector, asking a question that is answerable from measurement rather than from architecture: how much does a selector need to know, and what does that knowledge cost?
Every selector in this round is assigned to exactly one tier, and the tier is a contract about what the selector is permitted to read, enforced in code rather than by convention. The paper-only cache is written to disk as tensors; a tier-zero selector is given no handle to those tensors and no capability to open the file.
The atomic unit of repair is a cell, meaning one document token at one layer, with its key and value replaced jointly across all key-value heads. A budget is a number of cells, expressed as a percentage of the document region, and a geometry is a constraint on which sets of cells a selector is allowed to form.
| Geometry | Constraint | Why it is on the list |
|---|---|---|
| S1 | Any set of cells | Unconstrained; the most expressive selector and the loosest upper bound on what a mask can do. |
| S2 | A layer prefix per token | A token is repaired from layer zero up to some depth. Matches a partial-recompute scheduler that stops early. |
| S3 | Whole-token columns | A token is repaired at every layer or at none. Matches a system that recomputes selected tokens end to end. |
The three geometries are not merely three restrictions of the same problem. They price different actions. Under S1 and S2 a single action is one cell; under S3 a single action is thirty-six cells, because a token column spans the whole stack. That asymmetry turns out to matter more than anything else the geometry axis measures, and section 4 is about why.
Round 5 reported that patching every key and value of the document region, at every layer, still fell short of the clean cache by between eight and nineteen percent, and read that shortfall as a structural property: something about the request lived outside the document region and no document-cache repair could reach it. Round 6 traced the shortfall to the evaluation instead. Round 5 decoded with a probe query that had itself been formed under the instruction; masking the instruction's keys removes the instruction from attention but does not undo a query that was already shaped by it.
Under the prompt-masked evaluation the shortfall is not merely small. It is absent. Across the whole structured-repair set, the hundred-percent patch recovers logit fidelity of 1.000 on every single row, and the largest departure from unity anywhere in that set is two parts in a hundred trillion, which is float32 arithmetic and not a residual.
This is the result that makes the rest of the page interpretable. When a partial repair reaches 0.86 rather than 1.00, there is now exactly one place for the missing 0.14 to have gone: the selector chose the wrong cells. It has not been absorbed by an unreachable part of the request, and it is not an artefact of how the patch was installed. A recovery figure on this page is a statement about a decision, not about a mechanism.
Twenty-five candidate signals were screened in phase two by the only criterion that matters, namely the repair quality each one actually achieves when its ranking is used to spend a budget. Three winners went forward, one per information tier below the oracle, and the difference between them is not what they compute but what they are allowed to look at.
The prefill-only winner is an accumulated-inflow score. During the contaminated prefill it measures, at each layer, how large a contribution the instruction's keys and values make to a document token's attention output, expressed as a fraction of that token's hidden-state norm, and then accumulates that quantity causally down the stack with a decay of 0.8. It reads nothing but the prefill the system already performed, and costs no additional forward pass.
Here $p$ indexes a document token, $\ell$ a layer, $\mathcal{I}$ the instruction's key positions, $W_O$ the output projection and $h^{(\ell)}_p$ the block's hidden state. The score answers the question a contamination detector should ask: how much of this cell was written by the instruction?
The one-probe winner asks a second question and multiplies the two answers together. A single dry-run forward pass of the actual request query over the contaminated cache yields, for every cell, the weighted-value readout that query performs on it. The signal is the product of the two ranks.
Read plainly, $S^{\mathrm{IR}}$ repairs a cell when the instruction wrote to it and the query reads from it. Neither factor alone is competitive: the readout term by itself, without the inflow factor, is the native-weighted-value signal, which in screening reached 0.710 at the ten-percent budget against 0.791 for the product. The P0-dependent winner asks the question the deployable tiers cannot. It computes the probe's attention distribution over document keys twice, once against the contaminated cache and once against the clean one, and ranks cells by how far the distribution moved.
The table gives mean logit recovery under the unconstrained geometry, averaged over three instructions, four held-out papers and eight evaluation queries. Read down a column to compare tiers at a fixed budget; read across a row to see how each tier spends more.
| Selector · S1 · logit recovery | 0.25% | 0.5% | 1% | 2% | 5% | 10% | 20% |
|---|---|---|---|---|---|---|---|
| P0-dependent · attention shift | 0.600 | 0.669 | 0.773 | 0.802 | 0.863 | 0.915 | 0.932 |
| P0-free · one probe ($S^{\mathrm{IR}}$) | 0.384 | 0.482 | 0.602 | 0.677 | 0.763 | 0.859 | 0.901 |
| P0-free · prefill only | 0.167 | 0.250 | 0.326 | 0.331 | 0.502 | 0.597 | 0.682 |
| Baseline · evenly spaced tokens | 0.121 | 0.174 | 0.221 | 0.288 | 0.385 | 0.429 | 0.514 |
| Baseline · complete shallow layers | −0.007 | 0.028 | −0.009 | 0.004 | −0.010 | −0.005 | 0.008 |
| Baseline · random cells | 0.008 | 0.019 | 0.017 | 0.059 | 0.169 | 0.254 | 0.438 |
| Information gap (dependent − probe) | +0.215 | +0.186 | +0.172 | +0.125 | +0.100 | +0.056 | +0.031 |
Three readings follow. First, the information gap is real but it closes. Knowing the target cache exactly is worth 0.215 of recovery at a quarter-percent budget and 0.031 at twenty percent; by ten percent the deployable selector reaches 0.859 where perfect knowledge of the target reaches 0.915. The advantage of the oracle-grade signal is concentrated where the budget is too small to absorb a mistake, and it evaporates once the budget is generous enough that approximately-right choices overlap with exactly-right ones.
Second, the probe forward pass is where the deployable performance comes from, not the accumulated-inflow term. The prefill-only selector, which is free, trails the one-probe selector by 0.217 to 0.346 at every budget in the sweep, a gap larger at most budgets than the one between the one-probe selector and full knowledge of the target. A contamination map alone does not identify what to repair; the query has to be consulted about what it is going to read.
Third, both structural baselines are informative in different directions. Evenly spaced token columns beat random cells at small budgets and are beaten by them at twenty percent, so a naive stride is not a substitute for a signal. Patching complete shallow layers is worse than random at every budget in the sweep: at the ten-percent budget it restores the bottom four layers exactly, for every one of the document's tokens, and returns an end-to-end recovery of −0.005. That baseline is the sharpest available restatement of Deepdive #19's depth result, and section 7 returns to it.
The per-instruction rows repeat the difficulty ordering of the two predecessor studies without inverting it. Instruction A, the reviewer-2 rejection frame, is the easiest at every budget, reaching 0.786 at a quarter of a percent under the P0-dependent ranking against 0.543 for the novice-guidance instruction and 0.470 for the cross-lingual one. The ordering is preserved but the spread narrows sharply with budget: by ten percent the three instructions sit at 0.942, 0.882 and 0.921, and the cross-lingual instruction is no longer the hardest. The gap between instructions is a small-budget phenomenon, exactly like the gap between tiers.
The three geometries restrict a selector to increasingly system-shaped masks: arbitrary cells, then per-token layer prefixes, then whole-token columns. The engineering hope is that the restrictions are cheap, because only the third of them corresponds to something a recompute scheduler can actually execute without per-layer bookkeeping. For signal-driven repair that hope is largely met.
| Cost of the geometry restriction (S1 − S3) | 0.25% | 0.5% | 1% | 2% | 5% | 10% | 20% |
|---|---|---|---|---|---|---|---|
| P0-dependent · attention shift | +0.087 | +0.081 | +0.099 | +0.037 | +0.019 | +0.037 | +0.009 |
| P0-free · one probe | +0.079 | +0.140 | +0.226 | +0.179 | +0.135 | +0.125 | +0.068 |
| P0-free · prefill only | +0.022 | +0.047 | +0.020 | +0.016 | +0.119 | +0.072 | +0.067 |
For the P0-dependent ranking the restriction to whole-token columns costs at most 0.099, at the one-percent budget, and no more than 0.037 from two percent onward. The deployable one-probe selector pays more, up to 0.226 at one percent, which is worth noting: a weaker signal benefits more from the freedom to place cells individually, presumably because it needs the extra degrees of freedom to compensate for ranking errors. At the budgets a system would plausibly operate at, the whole-token column costs the deployable selector 0.125 at ten percent and 0.068 at twenty, which is a real but affordable price for a mask a scheduler can execute.
A signal ranks all cells at once, so a geometry merely reshapes the top of a list it has already produced. A search does not rank; it takes actions, and it is charged per action. The phase-three oracle was given sixteen actions and the three geometries turned that budget into wildly different masks.
| Greedy oracle · 16-action budget · paper 1908.05908, instruction A | cells bought | % of region | evaluations | recovery |
|---|---|---|---|---|
| S1 arbitrary cells | 2 | 0.0014% | 190 | 0.075 |
| S2 layer prefixes | 2 | 0.0014% | 193 | 0.076 |
| S3 token columns | 576 | 0.389% | 905 | 0.818 |
Under S1 and S2 the search exhausted its sixteen actions after two of them — every further single cell it could evaluate had utility too small to select — and bought two cells out of 148,068, recovering 0.075. Under S3 the same sixteen actions bought sixteen token columns, which is 576 cells, and recovered 0.818. The gap between 0.075 and 0.818 is not a statement about search quality. It is a statement about what a single action is worth, and at a tiny budget the whole-token column is the only atomic action of the three that buys anything at all.
A signal predicts which cells will help. The intervention oracle does not predict; it patches a candidate, measures the resulting distance to the clean cache, keeps the best candidate and repeats. It cannot be deployed, because each of its steps requires the very evaluation the repair was supposed to make unnecessary, but it bounds what any selector could achieve on the same geometry at the same cell count.
Because the geometry result of section 4 makes S1 and S2 oracle rows uninformative about search quality, the meaningful comparison is the S3 one, and it must be read at matched cell counts. The oracle bought 576 cells, 0.389 percent of the region. The table places it between the two signal budgets that bracket it, on exactly the same paper, instruction and evaluation query.
| S3 · paper 1908.05908 · instruction A · one query | cells | recovery | P0 bytes read | selector latency |
|---|---|---|---|---|
| P0-free one probe, 0.5% budget | 740 | 0.280 | 0 | 0.74 s |
| P0-dependent, 0.25% budget | 370 | 0.475 | 583 MB | 5.07 s |
| P0-dependent, 0.5% budget | 740 | 0.593 | 583 MB | 5.07 s |
| Intervention oracle, 16 actions | 576 | 0.818 | 1.69 GB | 86.5 s |
| P0-dependent, 1% budget | 1,480 | 0.834 | 583 MB | 5.07 s |
At a cell count between the 0.25 and 0.5 percent budgets, the oracle recovers 0.818 where the best signal at 0.5 percent recovers 0.593. That is an advantage of roughly 0.23 with about a fifth fewer cells, and the same margin appears against the four-paper mean of the P0-dependent selector at 0.5 percent, which is 0.587. The advantage is not unbounded, however: the same signal at one percent, using 2.6 times the oracle's cells, reaches 0.834 and passes it. The oracle is buying cell efficiency, not a capability the signals lack.
What it pays for that efficiency is the entire point. Nine hundred and five teacher-forced forward evaluations of the model were spent selecting one mask for one document, taking 86.5 seconds of selection latency and reading 1.69 GB of clean-cache tensors. The deployable selector produced its mask in 0.74 seconds having read none. If the objective is to avoid recomputing a document, spending nine hundred forward passes to decide how to avoid it is not a trade a serving stack can make; the oracle exists to measure the headroom, not to be run.
The natural response to a hand-designed signal is to learn one instead. Phase one built the supervision for exactly that. For twelve calibration papers, three instructions and seven layers spanning the stack, it patched one cell at a time and measured the exact change in distance to the clean cache, producing 22,176 labelled single-cell interventions. Two gradient-boosted predictors were trained on the 11,088 rows in the training half of a paper-disjoint split, one restricted to P0-free features and one allowed the P0-dependent ones.
Both failed the screening outright, at a level indistinguishable from choosing cells at random, and neither was promoted to phase three.
| Phase-2 screening · S1 · logit recovery | 1% | 5% | 10% |
|---|---|---|---|
| Best hand-designed P0-dependent (attention shift) | 0.603 | 0.790 | 0.778 |
| Best hand-designed P0-free ($S^{\mathrm{IR}}$) | 0.489 | 0.706 | 0.791 |
| Learned predictor, P0-dependent features | 0.044 | 0.087 | 0.226 |
| Learned predictor, P0-free features | 0.082 | 0.070 | 0.208 |
| Random cells | 0.019 | 0.090 | 0.200 |
The explanation is in the labels themselves, and it is not a defect of the model class. A single cell almost never matters. Across the 22,176 labelled interventions the mean absolute utility is 3.5×10−4 against a mean starting distance of 1.0×10−2, so the typical single cell moves the cache by three percent of the damage, and 44.9 percent of all single-cell interventions have negative utility: patching them moves the cache further from the clean one. That fraction is stable across the stack, ranging only from 43.6 percent at layer 18 to 46.0 percent at layer 24.
Against labels like these, no signal in the study has any rank correlation with the truth. Every one of the twenty-five candidates has a Spearman coefficient against exact single-cell utility that is visually indistinguishable from zero, and every one of them selects a set in which between 42 and 47 percent of the chosen actions turn out to be harmful in isolation. The learned predictors were trained on a target that carries almost no learnable structure, so they reproduce the noise faithfully and rank no better than chance.
Deepdive #19 closed on an open question. Its ranking signals were locally functional and downstream inert: a patch confined to one layer restored that layer's attention output and left the next-token distribution untouched. It could not tell whether the failure belonged to the signals or to the shallowness of the intervention. Round 6 patches across the whole stack, so the question is now answerable, and the answer has two parts.
The first part is that the signals do transfer once the patch is deep. At the one-percent budget the rank correlation between local recovery at the patched sites and end-to-end logit recovery is 0.896 across all 507 configurations. The predecessor's null was a property of the shallow intervention, not of the rankings.
The second part is that the correspondence degrades as budgets grow, falling to 0.471 at five percent and 0.441 at ten. Local fidelity saturates while end-to-end recovery keeps climbing, so the two stop tracking one another. Between five and seven percent of configurations at those budgets are locally positive and end-to-end negative, which means a local read is not a safe acceptance test for a repair.
The extreme case is the shallow-layer baseline. At the ten-percent budget it restores the bottom four layers of the network completely, every key and value of all 4,113 document tokens, which produces a local recovery of 0.955 — the highest of any selector in the study, and higher than the 0.535 achieved by the P0-dependent ranking that reaches 0.892 end to end. Its end-to-end recovery is −0.005.
A perfect repair of the shallow tenth of the stack and a partial repair spread across all of it therefore sit at opposite ends of both measures, with the ordering reversed. This is Deepdive #19's depth result restated as a contrast between two selectors at an identical budget, and it disposes of the cheapest deployment idea in the space: there is no shallow prefix of the network whose complete restoration is worth anything downstream.
A separate pass reads the same interventions at every block of the network rather than at the endpoints alone, which shows where along the stack a repair actually takes hold. Recovery is near zero for the first several blocks under every selector, rises steeply between blocks eight and twelve, and then accumulates slowly to the readout. At the one-percent budget the P0-dependent selector's hidden-state recovery is 0.002 at block zero, 0.073 at block eight, 0.408 at block eighteen and 0.558 at the last block; the one-probe selector traces the same shape at roughly half the height, 0.000, 0.059, 0.238 and 0.312.
Every oracle number on this page comes from greedy forward selection, which is itself an approximation. If greedy search were substantially worse than the best achievable mask at the same budget, the oracle would be understating the headroom and the signal-oracle gap of section 5 would be a lower bound of unknown looseness. Phase four measures that approximation directly, running greedy, beam and exhaustive search over the same candidate set at budgets of one to four actions, on one paper under instruction A.
The metric here is distance to the clean cache rather than recovery, so lower is better and the starting distance is what a search has to reduce.
| Distance to clean cache · lower is better | 1 action | 2 actions | 3 actions | 4 actions | evals at 4 |
|---|---|---|---|---|---|
| S3 exhaustive | 0.0534 | 0.0371 | 0.0236 | 0.0170 | 12,951 |
| S3 beam | 0.0534 | 0.0371 | 0.0236 | 0.0170 | 483 |
| S3 greedy | 0.0534 | 0.0371 | 0.0254 | 0.0176 | 91 |
| S1 exhaustive | 0.0831 | 0.0808 | 0.0800 | 0.0792 | 12,951 |
| S1 greedy | 0.0831 | 0.0826 | 0.0826 | 0.0826 | 70 |
| S2 exhaustive | 0.0791 | 0.0776 | 0.0767 | 0.0739 | 20,475 |
| S2 greedy | 0.0791 | 0.0780 | 0.0780 | 0.0780 | 73 |
Greedy search is well calibrated where the actions are worth taking. Under S3 at four actions it reaches 0.0176 where exhaustive search over all 12,951 four-subsets reaches 0.0170, so greedy is 3.6 percent worse using 142 times fewer evaluations, and beam search matches exhaustive exactly at 483 evaluations. The oracle numbers elsewhere on this page are therefore close to the true optimum for their geometry and budget, and the signal-oracle gap of section 5 is not an artefact of a weak search.
The table also reproduces the geometry result at a resolution the phase-three run could not reach. Under S3 the distance falls from 0.0534 to 0.0170 across four actions, a reduction of 68 percent, because each action is a whole token column. Under S1 the same four actions move the distance from 0.0831 to 0.0792 even with exhaustive search, a reduction of five percent, and greedy stalls completely after the second action. Exhaustive search over every possible four-cell mask cannot make four cells matter.
Four conclusions survive this round. A complete document-region patch, evaluated as cache reuse requires, restores the clean cache exactly, so the repair problem is entirely a selection problem. A deployable selector that never reads the clean cache recovers 0.859 of the difference at a ten-percent cell budget, and the value of knowing the clean cache exactly falls from 0.215 at a quarter-percent budget to 0.031 at twenty percent. The single probe forward pass, not the contamination map, is what makes that selector work. And repair quality is a property of masks rather than of cells, which is why every per-cell learned predictor in this study ranks no better than chance.
Six limitations bound those conclusions. Each one is stated because it restricts what a specific number on this page may be used for.