Experiment Deepdive #20

What Must a KV-Cache Repair Know to Choose?

Round 6 removes the last oracle assumption from the repair problem. A serving stack that wants to restore a contaminated document cache must decide which cells to repair without ever reading the clean cache it is trying to reach. Three information tiers, three repair geometries and a budget sweep from a quarter of a percent to everything measure what that decision actually costs. A single probe forward pass buys most of the information the clean cache would have supplied; a search that queries the model directly buys more, at nine hundred forward evaluations per document; and a learned per-cell predictor buys nothing at all, for a reason the data makes explicit.

Prompt-conditioned KV cache study · round 6 · Qwen3-8B at revision b968826d

Qwen3-8B QASPER 3 tiers × 3 geometries × 9 budgets 4 phases · 2 h 17 min

1. Repair is a selection problem before it is a patching problem

The two predecessors of this page established what a task instruction does to a document's cached keys and values, and how much of the network a repair has to touch before that damage comes back out. Both of them repaired with an oracle in hand: the patch source was a clean paper-only cache that the experiment had prefilled for the purpose, and the ranking that chose which positions to patch was computed by comparing the contaminated cache against that same clean cache. A serving stack has neither. If it had the clean cache it would serve from it and there would be nothing to repair. Deepdive #18 and Deepdive #19 cover that ground and are assumed here.

Round 6 separates the two halves of the oracle and retires only one of them. The patch source remains oracle-supplied, because a system that recomputes a document region can obtain the correct values by recomputing it; what a deployed system genuinely cannot obtain in advance is the knowledge of which cells are worth recomputing. This round therefore holds the patch source fixed and varies only the selector, asking a question that is answerable from measurement rather than from architecture: how much does a selector need to know, and what does that knowledge cost?

Three information tiers

Every selector in this round is assigned to exactly one tier, and the tier is a contract about what the selector is permitted to read, enforced in code rather than by convention. The paper-only cache is written to disk as tensors; a tier-zero selector is given no handle to those tensors and no capability to open the file.

P0-free
Tier 0 · deployable
Reads only the contaminated cache and, optionally, one extra probe forward pass. Zero access to the clean cache, enforced by capability.
0 B P0
P0-dep
Tier 1 · diagnostic upper bound
May compare the contaminated cache against the clean cache directly. Not deployable; it measures what perfect knowledge of the target would be worth.
583 MB P0
Oracle
Tier 2 · intervention oracle
May run the intervention and read the outcome, then choose again. Greedy forward selection against the true objective; the ceiling of what any selector could achieve at a given budget.
905 evals

Three repair geometries

The atomic unit of repair is a cell, meaning one document token at one layer, with its key and value replaced jointly across all key-value heads. A budget is a number of cells, expressed as a percentage of the document region, and a geometry is a constraint on which sets of cells a selector is allowed to form.

Geometry Constraint Why it is on the list
S1Any set of cellsUnconstrained; the most expressive selector and the loosest upper bound on what a mask can do.
S2A layer prefix per tokenA token is repaired from layer zero up to some depth. Matches a partial-recompute scheduler that stops early.
S3Whole-token columnsA token is repaired at every layer or at none. Matches a system that recomputes selected tokens end to end.

The three geometries are not merely three restrictions of the same problem. They price different actions. Under S1 and S2 a single action is one cell; under S3 a single action is thirty-six cells, because a token column spans the whole stack. That asymmetry turns out to matter more than anything else the geometry axis measures, and section 4 is about why.

The campaign

1
Calibration
12 papers, 66 cases. Exact single-cell utilities at seven layers; 22,176 labelled interventions.
2
Screening
25 candidate signals plus random, ranked by actual repair quality at 1, 5 and 10 percent under S1.
3
Structured repair
4 held-out papers, 48 cases, 3 geometries, budgets 0 to 100 percent. The main result.
4
Search calibration
Greedy against beam against exhaustive at one to four actions, to price the oracle's own approximation.
One evaluation, used throughout. Every number on this page is measured with the instruction's keys masked at every layer, so the probe query never attends the instruction. That is the evaluation the reuse objective requires — the goal is a cache that behaves as though the instruction had never been prefilled — and Round 6 verified it is functionally identical to physically deleting the instruction's slots and re-rotating the document. Every number is also measured against the paper-only target, never the length-matched neutral control, which is retained in the artefacts but is not the objective.

2. The ceiling is exactly one, so every deficit below it belongs to the selector

Round 5 reported that patching every key and value of the document region, at every layer, still fell short of the clean cache by between eight and nineteen percent, and read that shortfall as a structural property: something about the request lived outside the document region and no document-cache repair could reach it. Round 6 traced the shortfall to the evaluation instead. Round 5 decoded with a probe query that had itself been formed under the instruction; masking the instruction's keys removes the instruction from attention but does not undo a query that was already shaped by it.

Under the prompt-masked evaluation the shortfall is not merely small. It is absent. Across the whole structured-repair set, the hundred-percent patch recovers logit fidelity of 1.000 on every single row, and the largest departure from unity anywhere in that set is two parts in a hundred trillion, which is float32 arithmetic and not a residual.

1.000
logit recovery at a 100% patch, on all 507 rows
2.4×10−14
largest absolute deviation from 1.000 anywhere
4,552
measured repair rows on the paper-only target
11
selectors, all reaching 1.000 at full budget

This is the result that makes the rest of the page interpretable. When a partial repair reaches 0.86 rather than 1.00, there is now exactly one place for the missing 0.14 to have gone: the selector chose the wrong cells. It has not been absorbed by an unreachable part of the request, and it is not an artefact of how the patch was installed. A recovery figure on this page is a statement about a decision, not about a mechanism.

What this corrects, and what it does not. Deepdive #19's section 5 has been amended with this finding. Its absolute recovery numbers are deflated lower bounds measured under an instruction-conditioned probe, and its qualitative conclusions — the eighteen-layer depth threshold, the absence of a load-bearing attention head, the difficulty ordering across the three instructions — are unaffected, because all three are comparisons within a single evaluation. What does not survive is the reading of the residual as a hard ceiling on document-cache repair. There is no such ceiling.

3. One probe forward buys most of what the clean cache would have told the selector

Twenty-five candidate signals were screened in phase two by the only criterion that matters, namely the repair quality each one actually achieves when its ranking is used to spend a budget. Three winners went forward, one per information tier below the oracle, and the difference between them is not what they compute but what they are allowed to look at.

The prefill-only winner is an accumulated-inflow score. During the contaminated prefill it measures, at each layer, how large a contribution the instruction's keys and values make to a document token's attention output, expressed as a fraction of that token's hidden-state norm, and then accumulates that quantity causally down the stack with a decay of 0.8. It reads nothing but the prefill the system already performed, and costs no additional forward pass.

$$\mathrm{AI}^{\lambda}_{p,\ell} \;=\; \mathrm{inflow}_{p,\ell-1} \;+\; \lambda\,\mathrm{AI}^{\lambda}_{p,\ell-1}, \qquad \mathrm{inflow}_{p,\ell} \;=\; \frac{\big\lVert W_O \sum_{j\in\mathcal{I}} a^{(\ell)}_{p\to j}\, v^{(\ell)}_{j}\big\rVert}{\lVert h^{(\ell)}_{p}\rVert}$$

Here $p$ indexes a document token, $\ell$ a layer, $\mathcal{I}$ the instruction's key positions, $W_O$ the output projection and $h^{(\ell)}_p$ the block's hidden state. The score answers the question a contamination detector should ask: how much of this cell was written by the instruction?

The one-probe winner asks a second question and multiplies the two answers together. A single dry-run forward pass of the actual request query over the contaminated cache yields, for every cell, the weighted-value readout that query performs on it. The signal is the product of the two ranks.

$$S^{\mathrm{IR}}_{p,\ell} \;=\; \operatorname{rank}\!\big(\mathrm{AI}^{0.8}_{p,\ell}\big)\;\times\;\operatorname{rank}\!\big(\rho_{p,\ell}\big)$$

Read plainly, $S^{\mathrm{IR}}$ repairs a cell when the instruction wrote to it and the query reads from it. Neither factor alone is competitive: the readout term by itself, without the inflow factor, is the native-weighted-value signal, which in screening reached 0.710 at the ten-percent budget against 0.791 for the product. The P0-dependent winner asks the question the deployable tiers cannot. It computes the probe's attention distribution over document keys twice, once against the contaminated cache and once against the clean one, and ranks cells by how far the distribution moved.

What the information is worth

The table gives mean logit recovery under the unconstrained geometry, averaged over three instructions, four held-out papers and eight evaluation queries. Read down a column to compare tiers at a fixed budget; read across a row to see how each tier spends more.

Selector · S1 · logit recovery 0.25%0.5%1%2%5%10%20%
P0-dependent · attention shift0.6000.6690.7730.8020.8630.9150.932
P0-free · one probe ($S^{\mathrm{IR}}$)0.3840.4820.6020.6770.7630.8590.901
P0-free · prefill only0.1670.2500.3260.3310.5020.5970.682
Baseline · evenly spaced tokens0.1210.1740.2210.2880.3850.4290.514
Baseline · complete shallow layers−0.0070.028−0.0090.004−0.010−0.0050.008
Baseline · random cells0.0080.0190.0170.0590.1690.2540.438
Information gap (dependent − probe)+0.215+0.186+0.172+0.125+0.100+0.056+0.031

Three readings follow. First, the information gap is real but it closes. Knowing the target cache exactly is worth 0.215 of recovery at a quarter-percent budget and 0.031 at twenty percent; by ten percent the deployable selector reaches 0.859 where perfect knowledge of the target reaches 0.915. The advantage of the oracle-grade signal is concentrated where the budget is too small to absorb a mistake, and it evaporates once the budget is generous enough that approximately-right choices overlap with exactly-right ones.

Second, the probe forward pass is where the deployable performance comes from, not the accumulated-inflow term. The prefill-only selector, which is free, trails the one-probe selector by 0.217 to 0.346 at every budget in the sweep, a gap larger at most budgets than the one between the one-probe selector and full knowledge of the target. A contamination map alone does not identify what to repair; the query has to be consulted about what it is going to read.

Third, both structural baselines are informative in different directions. Evenly spaced token columns beat random cells at small budgets and are beaten by them at twenty percent, so a naive stride is not a substitute for a signal. Patching complete shallow layers is worse than random at every budget in the sweep: at the ten-percent budget it restores the bottom four layers exactly, for every one of the document's tokens, and returns an end-to-end recovery of −0.005. That baseline is the sharpest available restatement of Deepdive #19's depth result, and section 7 returns to it.

Recovery curves for three selector tiers across three geometries and three instructions
Figure 1. Logit recovery toward the paper-only cache under prompt-masked evaluation, with one row per instruction and one column per geometry. The horizontal axis is the patched cell budget as a percentage of the document region, on a categorical scale. Orange squares are the P0-dependent attention-shift ranking, blue diamonds the one-probe $S^{\mathrm{IR}}$, green triangles the prefill-only accumulated inflow, grey crosses the random baseline. All four curves meet at exactly 1.000 at the full budget in all nine panels. The black circle is the greedy intervention oracle, plotted at the cell fraction its action budget actually consumed rather than at a nominal budget slot; it is present only for instruction A, the only instruction on which the oracle was run. Note that the original module figure restricted every curve to the keys the oracle covered, which left the B and C rows empty; this version lifts that restriction for the signal curves and keeps it only for the oracle marker.

The per-instruction rows repeat the difficulty ordering of the two predecessor studies without inverting it. Instruction A, the reviewer-2 rejection frame, is the easiest at every budget, reaching 0.786 at a quarter of a percent under the P0-dependent ranking against 0.543 for the novice-guidance instruction and 0.470 for the cross-lingual one. The ordering is preserved but the spread narrows sharply with budget: by ten percent the three instructions sit at 0.942, 0.882 and 0.921, and the cross-lingual instruction is no longer the hardest. The gap between instructions is a small-budget phenomenon, exactly like the gap between tiers.

A deployable selector exists. This is the first result in the series that a serving system could act on. One extra forward pass of the request query over the cache it already holds, combined with an inflow map computed during the prefill it already performed, selects ten percent of the document's key-value cells and recovers 0.859 of the difference to a clean cache. The selector reads zero bytes of the clean cache, which it does not have, and its selection latency is 743 ms against 5,074 ms for the P0-dependent ranking that must read 583 MB of clean-cache tensors.

4. Geometry barely constrains a signal and completely determines a search

The three geometries restrict a selector to increasingly system-shaped masks: arbitrary cells, then per-token layer prefixes, then whole-token columns. The engineering hope is that the restrictions are cheap, because only the third of them corresponds to something a recompute scheduler can actually execute without per-layer bookkeeping. For signal-driven repair that hope is largely met.

Cost of the geometry restriction (S1 − S3) 0.25%0.5%1%2%5%10%20%
P0-dependent · attention shift+0.087+0.081+0.099+0.037+0.019+0.037+0.009
P0-free · one probe+0.079+0.140+0.226+0.179+0.135+0.125+0.068
P0-free · prefill only+0.022+0.047+0.020+0.016+0.119+0.072+0.067

For the P0-dependent ranking the restriction to whole-token columns costs at most 0.099, at the one-percent budget, and no more than 0.037 from two percent onward. The deployable one-probe selector pays more, up to 0.226 at one percent, which is worth noting: a weaker signal benefits more from the freedom to place cells individually, presumably because it needs the extra degrees of freedom to compensate for ranking errors. At the budgets a system would plausibly operate at, the whole-token column costs the deployable selector 0.125 at ten percent and 0.068 at twenty, which is a real but affordable price for a mask a scheduler can execute.

Selected repair masks under the three geometries
Figure 2. The masks the one-probe selector actually forms, for a single case at the five-percent budget: paper 1908.05908 under instruction A, 4,113 document tokens by 36 layers. The faint background is the selector score on a shared scale; dark marks are the cells chosen for repair. The three panels spend almost identical budgets — 7,403, 7,403 and 7,380 cells — on visibly different shapes. Under S1 the mask is scattered; under S3 it collapses into full-height stripes, since choosing a token commits all thirty-six of its layers. The score field itself is strongly banded by layer, with a bright band between roughly layers 15 and 26 and a nearly dark band below layer 4, which is why the shallow-layer baseline of section 3 selects cells the signal considers worthless.

The same restriction is decisive for a search

A signal ranks all cells at once, so a geometry merely reshapes the top of a list it has already produced. A search does not rank; it takes actions, and it is charged per action. The phase-three oracle was given sixteen actions and the three geometries turned that budget into wildly different masks.

Greedy oracle · 16-action budget · paper 1908.05908, instruction A cells bought % of region evaluations recovery
S1 arbitrary cells20.0014%1900.075
S2 layer prefixes20.0014%1930.076
S3 token columns5760.389%9050.818

Under S1 and S2 the search exhausted its sixteen actions after two of them — every further single cell it could evaluate had utility too small to select — and bought two cells out of 148,068, recovering 0.075. Under S3 the same sixteen actions bought sixteen token columns, which is 576 cells, and recovered 0.818. The gap between 0.075 and 0.818 is not a statement about search quality. It is a statement about what a single action is worth, and at a tiny budget the whole-token column is the only atomic action of the three that buys anything at all.

These oracle rows are not budget-matched to the signal columns. The oracle was run against an action budget, the signals against a cell budget, and the two do not coincide. The oracle appears in Figure 1 at the cell fraction it actually consumed, 0.389 percent, which falls between the 0.25 and 0.5 percent slots of the signal sweep; it never occupies the 1, 5 or 10 percent slots even though the run recorded it there. It was also run on one paper under one instruction and one query only. Any comparison between oracle and signal in this page states its cell counts explicitly for that reason.

5. The search advantage is real and costs nine hundred forward evaluations

A signal predicts which cells will help. The intervention oracle does not predict; it patches a candidate, measures the resulting distance to the clean cache, keeps the best candidate and repeats. It cannot be deployed, because each of its steps requires the very evaluation the repair was supposed to make unnecessary, but it bounds what any selector could achieve on the same geometry at the same cell count.

Because the geometry result of section 4 makes S1 and S2 oracle rows uninformative about search quality, the meaningful comparison is the S3 one, and it must be read at matched cell counts. The oracle bought 576 cells, 0.389 percent of the region. The table places it between the two signal budgets that bracket it, on exactly the same paper, instruction and evaluation query.

S3 · paper 1908.05908 · instruction A · one query cells recovery P0 bytes read selector latency
P0-free one probe, 0.5% budget7400.28000.74 s
P0-dependent, 0.25% budget3700.475583 MB5.07 s
P0-dependent, 0.5% budget7400.593583 MB5.07 s
Intervention oracle, 16 actions5760.8181.69 GB86.5 s
P0-dependent, 1% budget1,4800.834583 MB5.07 s

At a cell count between the 0.25 and 0.5 percent budgets, the oracle recovers 0.818 where the best signal at 0.5 percent recovers 0.593. That is an advantage of roughly 0.23 with about a fifth fewer cells, and the same margin appears against the four-paper mean of the P0-dependent selector at 0.5 percent, which is 0.587. The advantage is not unbounded, however: the same signal at one percent, using 2.6 times the oracle's cells, reaches 0.834 and passes it. The oracle is buying cell efficiency, not a capability the signals lack.

What it pays for that efficiency is the entire point. Nine hundred and five teacher-forced forward evaluations of the model were spent selecting one mask for one document, taking 86.5 seconds of selection latency and reading 1.69 GB of clean-cache tensors. The deployable selector produced its mask in 0.74 seconds having read none. If the objective is to avoid recomputing a document, spending nine hundred forward passes to decide how to avoid it is not a trade a serving stack can make; the oracle exists to measure the headroom, not to be run.

Decomposition of the gap between deployable and oracle selectors
Figure 3. The gap between the deployable selector and the oracle, split into the part attributable to information and the part attributable to search, at three nominal budgets and three geometries. Green circles are the best P0-free selector, orange squares the best P0-dependent one, black diamonds the oracle under that geometry and purple triangles the oracle's S1 reference. Two cautions apply and both are printed on the figure itself. All four levels are restricted to the single paper, instruction and query on which every level was run, so each column is n=1; and the annotated signal/search gaps compare the oracle's 0.389-percent mask against the signals at their nominal 1, 5 and 10 percent slots, which is why several of them are negative. The information gaps, which are budget-matched between the two signal tiers, are the reliable annotations here: +0.36 to +0.52 at one percent, +0.16 at five percent, +0.01 to +0.15 at ten.
Cost-quality Pareto frontier across selectors
Figure 4. Recovery against four cost axes over the whole structured-repair set. Green triangles are P0-free selectors, orange squares P0-dependent, black circles the oracle; the dashed line is the empirical Pareto frontier. The top-left panel separates the three tiers into three vertical bands of selection latency at roughly 0.74, 5.1 and 13–87 seconds, and the P0-free band spans the full recovery range from 0 to 1.000, so at equal latency the tier reaches every quality level the others do. The top-right panel is the sharpest: P0-free selectors sit on the zero line of clean-cache bytes read and still reach 1.000, while the other two tiers require between 0.58 and 1.7 GB before producing any mask at all. The lower-left panel is the only one on which the oracle is Pareto-optimal, since its axis is patched bytes rather than the cost of deciding.

6. Single-cell utility is unpredictable, so the learned predictors fail

The natural response to a hand-designed signal is to learn one instead. Phase one built the supervision for exactly that. For twelve calibration papers, three instructions and seven layers spanning the stack, it patched one cell at a time and measured the exact change in distance to the clean cache, producing 22,176 labelled single-cell interventions. Two gradient-boosted predictors were trained on the 11,088 rows in the training half of a paper-disjoint split, one restricted to P0-free features and one allowed the P0-dependent ones.

Both failed the screening outright, at a level indistinguishable from choosing cells at random, and neither was promoted to phase three.

Phase-2 screening · S1 · logit recovery 1%5%10%
Best hand-designed P0-dependent (attention shift)0.6030.7900.778
Best hand-designed P0-free ($S^{\mathrm{IR}}$)0.4890.7060.791
Learned predictor, P0-dependent features0.0440.0870.226
Learned predictor, P0-free features0.0820.0700.208
Random cells0.0190.0900.200

The explanation is in the labels themselves, and it is not a defect of the model class. A single cell almost never matters. Across the 22,176 labelled interventions the mean absolute utility is 3.5×10−4 against a mean starting distance of 1.0×10−2, so the typical single cell moves the cache by three percent of the damage, and 44.9 percent of all single-cell interventions have negative utility: patching them moves the cache further from the clean one. That fraction is stable across the stack, ranging only from 43.6 percent at layer 18 to 46.0 percent at layer 24.

Against labels like these, no signal in the study has any rank correlation with the truth. Every one of the twenty-five candidates has a Spearman coefficient against exact single-cell utility that is visually indistinguishable from zero, and every one of them selects a set in which between 42 and 47 percent of the chosen actions turn out to be harmful in isolation. The learned predictors were trained on a target that carries almost no learnable structure, so they reproduce the noise faithfully and rank no better than chance.

Signal quality against exact single-cell oracle utility
Figure 5. Every candidate signal measured against exact single-cell utility, on the twelve-paper calibration set. Upper left: three ranking-quality measures per signal, with the P0-free group left of the dashed line and the P0-dependent group right of it. The blue Spearman bars are flat against zero for all twenty-five signals, while the orange and green bars, which measure top-of-list quality, sit between 0.13 and 0.28 — the signals identify a somewhat enriched head of the list but carry no global ordering. Upper right: the fraction of each signal's selected actions that have negative utility, between 0.42 and 0.47 everywhere, against a base rate of 0.449. Lower row: exact utility against signal percentile for the best signal in each tier, showing the vertical scatter that the flat Spearman coefficients summarise.
Repair quality is a property of masks, not of cells. The two facts on this page are only apparently in tension. Cell-level utility is unpredictable, and yet the same signals, used to select tens of thousands of cells at once, recover 0.86 of the difference to a clean cache. Both are true because the cells cooperate: a mask works through the joint effect of its members on a deep non-linear network, and the marginal contribution of any one member, measured alone, is small, noisy and frequently negative. This is why a per-cell regression is the wrong learning target, and why a selector should be evaluated by the repair its mask achieves rather than by its agreement with single-cell labels.

7. Local repair and downstream repair are dissociable

Deepdive #19 closed on an open question. Its ranking signals were locally functional and downstream inert: a patch confined to one layer restored that layer's attention output and left the next-token distribution untouched. It could not tell whether the failure belonged to the signals or to the shallowness of the intervention. Round 6 patches across the whole stack, so the question is now answerable, and the answer has two parts.

The first part is that the signals do transfer once the patch is deep. At the one-percent budget the rank correlation between local recovery at the patched sites and end-to-end logit recovery is 0.896 across all 507 configurations. The predecessor's null was a property of the shallow intervention, not of the rankings.

The second part is that the correspondence degrades as budgets grow, falling to 0.471 at five percent and 0.441 at ten. Local fidelity saturates while end-to-end recovery keeps climbing, so the two stop tracking one another. Between five and seven percent of configurations at those budgets are locally positive and end-to-end negative, which means a local read is not a safe acceptance test for a repair.

The extreme case is the shallow-layer baseline. At the ten-percent budget it restores the bottom four layers of the network completely, every key and value of all 4,113 document tokens, which produces a local recovery of 0.955 — the highest of any selector in the study, and higher than the 0.535 achieved by the P0-dependent ranking that reaches 0.892 end to end. Its end-to-end recovery is −0.005.

0.955
local recovery, complete bottom four layers
−0.005
end-to-end recovery, same intervention
0.535
local recovery, best P0-dependent ranking
0.892
end-to-end recovery, same ranking

A perfect repair of the shallow tenth of the stack and a partial repair spread across all of it therefore sit at opposite ends of both measures, with the ordering reversed. This is Deepdive #19's depth result restated as a contrast between two selectors at an identical budget, and it disposes of the cheapest deployment idea in the space: there is no shallow prefix of the network whose complete restoration is worth anything downstream.

Local recovery against final logit recovery
Figure 6. Local recovery at the patched sites against end-to-end logit recovery, one panel per budget, coloured by geometry and shaped by tier. The dashed diagonal is equality. Almost every point lies above it, so a repair achieves more at the output than at the site of the intervention, which is the opposite of the predecessor study's single-layer finding. The cloud is tight and steep at one percent and fans out at five and ten, matching the fall in rank correlation from 0.896 to 0.441. The isolated column of points at local recovery near 1.0 with end-to-end recovery near zero, visible in the five and ten percent panels, is the complete-shallow-layers baseline.

A separate pass reads the same interventions at every block of the network rather than at the endpoints alone, which shows where along the stack a repair actually takes hold. Recovery is near zero for the first several blocks under every selector, rises steeply between blocks eight and twelve, and then accumulates slowly to the readout. At the one-percent budget the P0-dependent selector's hidden-state recovery is 0.002 at block zero, 0.073 at block eight, 0.408 at block eighteen and 0.558 at the last block; the one-probe selector traces the same shape at roughly half the height, 0.000, 0.059, 0.238 and 0.312.

Hidden-state recovery accumulating across downstream blocks
Figure 7. Hidden-state recovery as a function of downstream block, one column per geometry and one row per read point within the block, with the final readout on the bottom row. Orange squares are the P0-dependent selector, green triangles the one-probe P0-free selector and black circles the greedy oracle, averaged over the one, five and ten percent budgets. The two signal tiers separate within the first ten blocks and hold their separation for the remaining twenty-six, so the information gap of section 3 is established early in the stack rather than accumulated late. The oracle traces are the geometry result again: flat at zero across the whole network under S1 and S2, where its sixteen actions bought two cells, and climbing past the P0-free curve under S3.

9. What follows, and what does not

Four conclusions survive this round. A complete document-region patch, evaluated as cache reuse requires, restores the clean cache exactly, so the repair problem is entirely a selection problem. A deployable selector that never reads the clean cache recovers 0.859 of the difference at a ten-percent cell budget, and the value of knowing the clean cache exactly falls from 0.215 at a quarter-percent budget to 0.031 at twenty percent. The single probe forward pass, not the contamination map, is what makes that selector work. And repair quality is a property of masks rather than of cells, which is why every per-cell learned predictor in this study ranks no better than chance.

Six limitations bound those conclusions. Each one is stated because it restricts what a specific number on this page may be used for.

What a repair still costs. Selecting well is now possible without an oracle, but selecting is only half the operation. The cells a selector chooses still have to be filled with correct values, and this round supplied them from a prefilled clean cache exactly as its predecessors did. Whether a system can produce those values more cheaply than recomputing the document — and at ten percent of cells spread over every layer, a partial recompute is not obviously cheap — is the question the next round has to answer. This page establishes that the selection problem is tractable, not that the repair is economical.