Hardware-level measurement campaigns on real datacenter GPUs. Where the engine deep-dives read source code, these read performance counters: each entry runs a controlled experiment, reports the raw numbers, and explains what they mean for serving systems.
Nsight Compute byte ledgers for every kernel class vLLM dispatches for Llama-3.1-8B — GEMM, FlashAttention, RMSNorm, RoPE, KV-cache writes, sampling — across the full memory hierarchy, from HBM through L2 and shared memory into the Tensor Cores. Includes per-step traffic rollups and a measured ranking of where optimization potential actually lives.
A plain-language walkthrough of a measurement study, written for a reader with no GPU background. When a model runs, leftover data sits briefly in the on-chip L2 cache between computation steps — folklore says reusing it saves several percent of every step. A first careless measurement agreed; careful measurement showed it is worth under two-thirds of one percent, and that three standard profiling instruments are biased toward overstating it by 7–246×. The story of how a quick experiment lied, and the pre-registered referee that caught it.
A closed-form decode-latency model splits one step into a GPU-compute term and a host-overhead term. CUDA graphs are the production feature built to erase host overhead — so enabling capture should subtract exactly the model's host term. Measured on three datacenter GPUs, graph capture removes 26.7–46.0% of the step, monotone in HBM bandwidth as predicted, but its batch-shape is the opposite of the modeled scheduler residual and the predictor over-predicts the regime by ~22%. A pre-registered kill-switch fired; the page reports what graphs actually remove (a batch-constant non-overlapped host cost) rather than re-fitting to a positive story.
A decode-latency calculator, calibrated once on Llama-3.1-8B, claims its GPU efficiencies are properties of the silicon — so swapping only the architecture numbers should predict a new model for free. We freeze every efficiency, swap shapes, and test on a 1B, a 7B, and a mixture-of-experts model. The GPU-compute shape transfers cleanly to a sibling 7B (5.5% residual, below the anchor's 8.4%), but absolute latency does NOT transfer: its host floor is model-size dependent, and the MoE carries an unmodeled +18 ms host floor with a 34.6%-off long-context cell. Reported as an honest map of where fit-free holds and where it breaks.
To predict one decode step under tensor parallelism, the natural model shards the single-GPU compute as one over k and adds the all-reduce cost. The all-reduce term, calibrated alone against a standalone NCCL sweep, is faithful (R² > 0.999). The composition is not: accurate at TP-2 (2.86% MAPE) but failing its 15% kill-switch at TP-4 (42.3% MAPE), because the measured Llama-3.1-70B step shrinks only 1.12× from TP-2 to TP-4 against the ideal 2×. The compute saving never materializes; a same-node control rules out node variance and localizes the missing term to a non-shrinking eager per-step host cost.
A prior study (Deepdive #2) measured a survival cliff for dirty data left in a GPU's L2 cache. This work fits that cliff with a capacity-ratio law — half the data is gone once the intervening traffic reaches 0.30× the cache — calibrated on three GPUs (R² = 0.99) and pre-registers a single prediction for a held-out L40S: half-survival at 28.9 MB for every blob size. The measurement falsifies it: the half-survival pressure moves strongly with blob size, from 9.4 MB for a 96 MB blob (the clean monotone falsifier) to about 80 MB for a 25.2 MB blob. The L40S points are only consistent-with an absolute-capacity form (L − b) + 9.3 on a thin 3-point sweep — which does not fit the trained anchors either, so neither law transfers cleanly across all four SKUs. An honest, pre-registered falsification.
GPU-centric latency models measure the wrong term: a directly-measured host floor governs eager decode and transforms predictably across five system axes.
Advisor-facing status snapshot of the GPU-frequency measurement note: clock droop 20.0–42.7% (L40S −15% throughput), $\eta_c=f_{clk}\cdot\eta_{noclk}$ reconstruction MAPE 29%→10%, 5090 causal chain, NEW cross-card reproducibility (2–3 distinct cards/SKU), cold Pro-gate 6.5/10.
A complete re-examination of the per-step inference predictor ('m2'): the full t_step equation, the five-factor GEMM efficiency that makes the SM clock explicit ($\eta_c = f_{clk}\cdot\eta_{cluster}\cdot\eta_{active}\cdot\eta_{duty}$, 18–44% power-cap clock droop, L40S −15% throughput), and an honest account of the Vidur comparison: a forest that memorises its own profiling table wins on absolute MAPE on already-profiled hardware (0.88% vs our 21%), yet on the capacity-planning decision Vidur exists for — rank unprofiled hardware — non-intrusive m2 gets top-1 100% / Spearman +1.00 with zero profiling while the memoriser cannot even play. Three held-out-validated GEMM root-cause fixes (η_wave occupancy, per-op η̄, attention re-fit) are why m2 ranks correctly; plus the lesson that per-operator accuracy ≠ per-step accuracy.
A data-grounded self-audit of the two-half project. Every claim traced to a raw file: the model is ~7 terms (not 3), microbenchmark-calibrated (not online learning); the frequency droop is real but declines throughput only on tight-cap parts. Part 2 is honest about losing to Vidur's RF, tying DistServe, and winning QLM only on estimator accuracy — with cross-GPU transfer and QLM SLO unmeasured.
A self-contained, advisor-facing report of the two-part project: Part 1 the predictor (full per-step model, five-factor GEMM ceiling, clock factor, Hopper cluster cap, L2 memory law, not online learning, where accuracy holds/fails) and Part 2 the utility (Vidur decision-win, DistServe, QLM honest gap), with a status scorecard.
The round-1 experiment program (29 analysed, 36 figures), each section mapped to its blueprint question: GEMM root-cause fixes, cross-SKU η_wave + Hopper persistence, an L2 parameter law, the light-kernel model, the key honest finding that fitting Vidur's GT hurts real wall-clock serving, online-learning adapt-only, decision quality 92.3%, power-aware scheduling +10.5%, Sarathi chunk +30.7%, QLM SLO mixed — every number traced, negatives included.
An advisor-facing closing report: the t_step equation, the five-factor ceiling and the H100 989-TFLOP breakdown, per-SKU clock curves on six GPUs, the cross-GPU transfer win (1.56× on unseen H100), the metric-divergence lesson (Vidur GT ≠ wall-clock), utility results (SKU 92.3%, power-aware +10.5%, Sarathi +30.7%), and a submission-directions map of three candidate papers with maturity, gaps, and venues. In-flight items marked, not invented.
A semester closing report across the two main PhD lines — agent-kvcache and MAS_comm — organised around a mind map of how every sub-topic connects, the problem each attacks, and where each stands. Validated: the four-policy scheduling campaign (Continuum 3.3–4.7× JCT, ~98% prefix hit), CPU-offloading (2.95× under VRAM pressure), the live trace viewer, a paused continuum sweep. Open and honestly marked: the forward-latency model (max-form rejected on 301K rows, branch unmerged) and MAS_comm KV communication (92%→67% accuracy for ~1.05× speed; CacheBlend dataset-dependent). JitServe and SimAI reported not-started. Cross-links to the predictor deepdives (#1–6, #10–12).
A multi-turn agent job re-reads its growing prefix every turn but idles through a multi-second tool call between turns while its KV prefix sits in scarce GPU memory. vLLM's native offloading can spill it to a DRAM tier, but its content-blind LRU evicts exactly the shared prefix blocks the next turn re-reads. On a PACE H200 with Llama-3.1-8B under VRAM pressure, an 8 GB pool under LRU returns 0.0% hits — worse than no offloading. Continuum-D routes agent job structure through kv_transfer_params into an out-of-tree four-class eviction policy (finished, overdue, far-wake-up, recency) plus admission control and predictive warm promotion, zero core changes: 0%→11.7% hits, −17% late-turn TTFT (−26% with warm promotion). Design validated; publishing gated.
The mainline system paper Continuum-D (#15) grew into, now measured on real agent traces at scale. Under 42–82 GB working-set pressure with 6–16 s reuse windows and an 11.3× thrash tax, faithful reimplementations of LRU, MORI-idleness, Marconi-utility, and Continuum-TTL all hit ≤1.6% and land at-or-below no DRAM tier at all — on production TraceLab traces LRU/TTL are CI-clean worse than no tier. A capacity-aware admission gate (necessary substrate, 33% alone) plus one client-owned bit last_turn (22% alone) recover 97% superadditively; gap prediction is unnecessary. Result: 19.2% lower p95 vs LRU (H100), 18.5% hit where every baseline is below 2%, CI-clean under three bootstraps. Honest negatives: synchronous quantized offload is net-negative; hybrid SSM transfer works (+8.2% JCT) but is granularity-bounded. Zero-fork on stock vLLM v0.23. ASPLOS target, iterating.
When agents hand off their KV-cache instead of re-sending text, the receiver skips a full prefill but inherits a spliced cache that never ran the cross-passage attention — so a small budget of tokens must be recomputed, and the whole method reduces to which tokens. Every prior selector (CacheBlend, RelayCaching, KVCOMM, Ada-KV) scores a sender-side signal: KV deviation or attention, blind to what the receiver is deciding. CADENCE scores decision-saliency instead — one backward pass of the receiver's own decision loss measures how much its answer would move if a token's cache were repaired. On HotpotQA·Llama-3.1-8B at b=0.2 it reaches 60.0% against a 64.2% full-reprefill ceiling, ahead of CacheBlend's 43.3% (which dips to random), RelayCaching's 52.5%, and the exact-deviation DenseDev at 54.2%; it exceeds CacheBlend's best-ever accuracy while recomputing about 2.4× fewer tokens. Reported honestly and framed as a selection-quality claim only: the timing verdict is final on both models — there is no wall-clock speedup, the method runs about 13% slower because the selection backward pass (~84 ms) exceeds the prefill it saves (~45 ms), and the score is a first-order saliency, not a validated causal effect. The efficiency path — a cheap approximate saliency estimator — plus MuSiQue + 5-model breadth, a GAIA multi-agent pipeline, and multi-hop are still running.
Serving stacks reuse a document's KV cache across requests, which assumes the cached representation belongs to the document. Place a task instruction in front of the document and it no longer does. Holding a QASPER paper byte-identical and comparing three instructions against length-matched neutral prefixes, this study decomposes the attention output difference exactly into routing, value rewrite and interaction, and finds that value rewrite dominates for all three while cross-lingual summarization acts almost entirely through it. The engineering question is settled by repair rather than by magnitude: ranking positions by the total weighted-value difference removes 63.5% of the damage at 1% of positions and 87.5% at 10%, whereas the raw key distance that every earlier figure plotted recovers 0.224 against a random baseline's 0.125, and a task-query alignment score falls below random despite 0.7 hotspot correlation. Correcting keys without values is actively harmful at depth, reaching -0.5 recovery at layer 30. Two follow-ups sharpen the picture and one of them is negative: masking the prefix keys during decoding leaves the difference almost intact (survival ratio 0.93 / 0.96 / 0.82), so the imprint really is carried by the paper's own cache rather than by continued reading of the instruction -- but repairing a single layer recovers approximately zero of the end-to-end logit difference and is indistinguishable from random, so the detector result is scoped to layer-local repair. Section attribution shows no semantic targeting either: all three instructions enrich the conclusion about sixfold and deplete the method section, which restates position rather than task.
Deepdive #18 measured every repair against a length-matched neutral prefix, which isolates task semantics but is the wrong target for the engineering goal: a serving stack needs a document's cache to behave as though only the anchor and the document had been prefilled. Round 5 switches the target to that paper-only baseline. The switch is not a flag change, because the document sits at a different absolute offset under the baseline than under an instruction and rotary embedding writes absolute position into every key, so the patch source must be the baseline key re-rotated forward by the instruction's length while values copy directly; the correction is guarded by an equality unit test and by a canonical rebuild that reaches 1.000 by construction. The headline is a depth result. Patching the top 5% of positions across a contiguous span of layers ending at the last block recovers essentially nothing at spans of 1, 2, 4 or 8 layers, where the entries change sign and are not separated from a random ranking, and reaches +0.748 / +0.729 / +0.252 only at 18 layers and +0.887 / +0.688 / +0.395 at all 36. The imprint is therefore written across roughly half the network, and no cheap shallow shortcut exists. Cross-lingual summarization is the hardest instruction throughout, and at full depth its sparse budget captures 0.395 of the 0.810 that patching every position achieves. Two further nulls tighten the picture: no single key-value head carries the effect, with a best cell of +0.259 and −0.281 at the adjacent head and no ranking stable across instructions, and every cheap single-layer signal is locally functional (+0.20 to +0.32 at the patched layer's attention output against random's +0.02 to +0.05) but not downstream causal, decaying to approximately zero at the final hidden state and inverting at the logit under the cross-lingual instruction. The ceiling is reported rather than tuned away: replacing every document key and value at every layer recovers only 0.914 / 0.876 / 0.810, because the instruction's own cache slots and the probe query lie outside the document region, so perfect cache restoration is not full restoration. The repair is oracle-style on three papers with one probe each, which measures where the damage lives and what removing it would cost rather than how to remove it in deployment.
Deepdives #18 and #19 repaired with an oracle in hand: the ranking that chose which cells to patch was computed by comparing the contaminated cache against the very clean cache the repair was trying to reach. A serving stack has no such cache. Round 6 retires that half of the oracle and measures what a selector actually needs to know, across three information tiers, three repair geometries and budgets from 0.25% of key-value cells to all of them. The ceiling comes first and it is clean: evaluated as cache reuse requires, with the instruction's keys masked at every layer, a 100% document patch recovers exactly 1.000 on all 507 measured rows — the largest deviation anywhere is 2.4×10−14 — which retires Deepdive #19's apparent full-patch residual as an artefact of decoding with an instruction-conditioned probe, and means every deficit below 1.000 belongs to the selector. The headline is that one probe forward pass is nearly enough. A deployable selector reading zero bytes of the clean cache — the product of an accumulated-inflow contamination map computed during the prefill the system already ran and a weighted-value readout from one dry-run forward of the request query — recovers 0.859 at a 10% cell budget against 0.915 for a selector granted exact knowledge of the target; the information gap falls from +0.215 at 0.25% to +0.031 at 20%, and the probe, not the contamination map, is what buys the performance, since the prefill-only signal trails by 0.22 to 0.35 at every budget. Geometry is nearly free for a signal, costing the best ranking at most 0.099 to restrict to whole-token columns, but decisive for a search: given 16 actions the greedy intervention oracle bought 2 cells and recovered 0.075 under arbitrary cells, and 576 cells (0.389% of the region) for 0.818 under token columns. That oracle advantage is about +0.23 at matched cell counts, priced at 905 forward evaluations, 86.5 s and 1.69 GB of clean-cache reads against the deployable selector's 0.74 s and zero. Two negative results carry weight. Oracle-supervised per-cell predictors trained on 22,176 exact single-cell utilities rank no better than random (0.04–0.23 against random's 0.02–0.20), because 44.9% of single-cell interventions have negative utility and every one of the 25 candidate signals has a Spearman coefficient of approximately zero against exact utility — repair quality is a property of masks, not cells. And local repair is dissociable from downstream repair: completely restoring the bottom four layers for all 4,113 document tokens gives the highest local recovery in the study, 0.955, and an end-to-end recovery of −0.005. Greedy search is calibrated at 3.6% worse than exhaustive using 142× fewer evaluations. Bounded by a four-paper test split, an oracle run on one paper under one instruction, and 16-token evaluation probes.