Mechanistic Attribution of Prompt Requirements

How Do LLMs Route and Execute Multiple Prompt Requirements Through Attention and the KV Cache?

A literature review for the research proposal Mechanistic Attribution of Prompt Requirements in LLM Generation: requirement–behavior mapping, internal routing through attention heads and KV entries, and causal validation by intervention.

Literature Review · 163 works Mechanistic Interpretability Instruction Following Attention & KV Cache GSM8K

Compiled September 2026 · literature through August 2026

Prepared for GNAN Lab, School of Electrical and Computer Engineering, Georgia Institute of Technology

1. The Research Question and Why It Is Timely

This review supports a research proposal that asks how a Transformer language model executes a prompt containing several natural-language requirements at once. The benchmark literature usually calls these constraints; this review says requirement throughout, for the same object. A representative prompt prepends instructions such as Give reasons for making decisions, The last line of your output contains only the final result, and Do not include units to a GSM8K word problem, so that the model receives the prefix $P=[I_1,I_2,\ldots,I_m,Q]$. That is to say, the model is handed one continuous prefix in which the requirements come first and the question comes last. The proposal asks three questions of increasing depth: which observable behavior $B_j$ each requirement $I_i$ controls (requirement–behavior mapping); through which hidden states, attention heads and KV-cache entries that control is exercised, and how the routing shifts between the reasoning phase and the final-answer phase (internal routing); and whether the identified routes are causally necessary, as established by attention suppression, activation or KV patching, and instruction rewriting (causal influence). Those three names are parts of the model, and each is defined once here: a hidden state is the list of numbers the model holds at one word position, an attention head is one of the separate units inside a layer that each choose which earlier positions to read from, and a KV-cache entry is the stored pair — one key and one value — that the model keeps for every prompt position so that it can consult that position again while it writes. Patching, likewise, means copying one such internal number out of one run of the model and pasting it into the same place in another run, to see whether the behavior moves with it. Put simply, the three questions ask what changes in the output, where inside the model that change is carried, and whether the carrier can be shown to do the work rather than merely to accompany it. The sections below organize the literature around these three questions, labelling them Level 1 (requirement–behavior mapping), Level 2 (internal routing) and Level 3 (causal influence). Later sections carry the short tags L1, L2, L3 for these three levels of depth — the L is for Level, and it is unrelated to the requirement-count levels FollowBench numbers 1 to 5 (FollowBench itself calls them constraint levels).

The prompt, and the three questions asked of it

The prompt prefix

$P=[I_1, I_2, \ldots, I_m, Q]$ — several natural-language requirements $I_i$ prepended to one GSM8K question $Q$. In the running example the requirements are give reasons for making decisions, the last line contains only the final result, and do not include units.

▼

Level 1 (L1)  Requirement → behavior

Which observable behavior $B_j$ does each requirement $I_i$ actually control? Answered by benchmarks that score every requirement separately.

Level 2 (L2)  Internal routing

Through which hidden states, attention heads and KV entries is that control exercised — and does the route change between the reasoning phase and the answer phase?

Level 3 (L3)  Causal necessity

Is the identified route necessary? Tested by attention suppression, activation and KV patching, and instruction rewriting.

Sections 2 through 8 are tagged L1, L2 or L3 — Level 1, 2 and 3 above — to show which of the three questions a body of work informs.

1.1 Why the question has become pressing

AgentIF Figure 2: one agent instruction carrying eight separate constraints, colour-coded by type
Figure — AgentIF: A single agent instruction, with every requirement it carries marked. The eight labels C1 to C8 pick out separate obligations inside one prompt: write structured Python, complete only the first objective if there are several, wrap output in <code> tags, call the given function at most once, save files to a fixed path, use only the listed predefined functions, follow the example’s format, and answer the query. The shading gives each requirement a kind — purple for semantic, pink for formatting, blue for tool use — and the right-hand key sorts them again by how they are stated: plainly, as a condition, or by example. This is what makes the question pressing. The obligations are simultaneous, heterogeneous, and scattered through the prompt, so “which requirement is governing this token” has no obvious answer.

Until roughly 2023, instruction-following research treated the prompt as a monolithic input and measured a single outcome, typically task accuracy or a preference score. Two developments have changed that framing. First, prompts in deployed agent systems now carry many simultaneous requirements: the AgentIF benchmark reports that instructions in real agentic applications average 1,723 words and about 11.9 requirements each [14], and SysBench and Multi-IF document that system-level rules must be honored jointly with user-turn requirements across multiple turns [7, 8]. That is to say, one deployed instruction is now about as long as a short article and carries roughly a dozen separate things the model has to get right at the same time, and in the multi-turn case not all of them were stated in the turn being answered. Second, the rise of long chain-of-thought reasoning models — models trained to write out a long stretch of working before they commit to an answer — has exposed a tension between reasoning and obedience: MathIF finds that the longer a model reasons on a mathematics problem, the less reliably it obeys the requirements it was given; under combined requirements the best reasoning model in that study reaches only 50.7% hard accuracy [16], and a fifteen-model study finds that explicit chain-of-thought reasoning lowers instruction-following accuracy on IFEval and ComplexBench [17]. Hard accuracy counts a response only when every requirement in the prompt is satisfied, so 50.7% means that on about half of those problems at least one requirement was missed. The prompt has therefore become the primary control surface for model behavior at exactly the moment when its individual components are known to compete with one another, yet the internal machinery by which a model separates, stores and consults each requirement remains understood only in fragments.

1.2 A concrete failure, at the granularity of tokens

When Thinking Fails Figure 2: average attention drop by layer, WIN versus LOSE cases
Figure — When Thinking Fails: Failure 2 measured directly, on Qwen2.5-1.5B-Instruct over IFEval. The paper defines constraint attention as how much the answer tokens attend to the requirement-bearing part of the prompt, averaged over the answer positions at a given layer. The attention drop plotted here is that quantity without chain of thought minus the same quantity with it, so a positive value means switching on the chain of thought pulled attention away from the requirement. The horizontal axis is layer number; the two lines split the cases by what the chain of thought did to instruction following — green where it helped (WIN), red where it hurt (LOSE). Read the vertical gap between them: the red line sits above the green at every layer, and the separation is widest in the early-to-middle layers, peaking near layer 14 where the losing cases drop about 0.0085 and the winning cases drop essentially nothing. Losing the requirement and attending to it less are the same event.

Consider a GSM8K problem whose correct answer is 18 dollars, presented with the three requirements above. Three distinct failures can produce a wrong result on that problem, and they are worth separating because each involves a different requirement and a different phase of generation.

Failure 1 — the format requirement deletes the reasoning

Under JSON-mode decoding — a decoding constraint that forces every output to be valid JSON — GPT-3.5-turbo placed the answer key before the reasoning key in 100% of outputs, so the chain of thought never ran. GSM8K exact match, the share of problems whose final answer string matches the reference exactly, fell from 76.6% to 49.3%, and for Claude-3-Haiku from 86.5% to 23.4% [18]. In other words, valid JSON was guaranteed by the decoding constraint and the task was lost instead: between about a third and about three quarters of the answers that had been right stopped being right.

Failure 2 — the model loses contact with the instruction

Attention paid to requirement-relevant prompt tokens decays over a long chain of thought, so by the final line the model has effectively lost the instruction forbidding units [17]. Concretely, the share of each newly written token's attention that lands on the requirement words shrinks as the working gets longer; the instruction is still sitting in the prompt, it is simply being read less and less. The last line reads 18 dollars; the arithmetic was right and the answer is marked wrong.

Failure 3 — the visible reasoning is not the computation

The last line carries a bare number, but the number came from post-hoc rationalisation rather than from the reasoning shown above it — that is to say, the model arrived at the answer by some other route and then wrote working that agrees with it — a pattern documented in both frontier and open reasoning models [99]. Nothing in the transcript distinguishes this from Failure 2.

Why one accuracy number cannot diagnose these. Each failure involves a different requirement, a different phase of generation, and presumably a different internal pathway. All three produce the same observable outcome: a wrong final line.

1.3 How much is lost under the status quo

FollowBench Figure 1: one instruction with requirements added one level at a time, and where two models start failing
Figure — FollowBench: What accumulation costs, held to one instruction. The top row is the bare request, recommend me ten books. Each row below adds exactly one more requirement to the row above it, with the newly added words underlined: first the content must be Chinese books, then a scenario is supplied (an interest in the Tang Dynasty), then a style (Shakespeare’s tone), then a format (bullet points), then an example to imitate. Nothing is ever removed, so row five carries all five requirements at once. The two columns on the right are two models — the crowned one is the stronger — and each cell says whether that model satisfied that row. Read down each column to find where it turns red. The stronger model holds through four requirements and fails on the fifth; the weaker one fails from the fourth. Neither requirement is hard on its own; what breaks the models is how many are in force at the same time.

The behavioral literature supplies baseline numbers that quantify the problem before any interpretability method is applied. The figures below are drawn from the benchmarks reviewed in Section 2 and establish that compliance degrades steeply as requirements accumulate, that format requirements interact destructively with reasoning, and that conflicts between instructions are resolved correctly less than half of the time. Concretely, each box gives one measured number, the benchmark it was measured on, and what the number is a rate of; read them as the size of the gap that any mechanistic account has to explain, not as a ranking of the models named.

84.7% → 61.9%
GPT-4 hard satisfaction rate on FollowBench as the requirement count grows from its level 1 to its level 5 [2] — hard means a response scores only if every requirement is met
77.7% → 33.0%
Average accuracy of 19 LLMs from one requirement (Level I) to four (Level IV) [11] — the levels differ in how many requirements the prompt carries
86.5% → 23.4%
Claude-3-Haiku GSM8K exact match, free-form versus JSON-constrained output [18] — exact match means the final answer string equals the reference answer
50.7%
Best hard accuracy among 23 reasoning models on MathIF (math + format requirements) [16] — hard means every format requirement is met
48%
Best open-model rate of correctly resolving conflicting system/user instructions on IHEval [24] — correct means following the system-level rule when the user turn contradicts it

1.4 Why the obvious approaches are insufficient

Jain and Wallace Figure 1: original versus adversarial attention over the same review, identical prediction
Figure — Attention is not Explanation: The same negative movie review, twice, with the tokens the model attends to shaded in blue. On the left is the attention the model actually produced, concentrated on asking and waste. On the right is an attention distribution constructed adversarially to be as different as possible, concentrated instead on myself and was. Underneath each is the model’s output for that input: f(x|α,θ) = 0.01 on the left and f(x|α̃,θ) = 0.01 on the right — the same number. Two attention maps that disagree about which words mattered give an identical prediction, so a reader cannot conclude from a highlighted token that the token drove the output.

Two approaches suggest themselves immediately, and the literature has established independently why neither settles the question. The first is to read the attention maps directly; the second is to intervene by amplifying attention to the instruction block.

Strawman: read the attention map, conclude which requirement governs the token

If a generated token attends to the tokens of requirement $I_i$, conclude that $I_i$ governs that token.

▼  three independent reasons this is unsafe

Attention is not importance

Adversarially different attention distributions can yield identical predictions [80, 81]. In other words, one can construct by hand a second attention pattern that points somewhere else entirely and still get the same output, so the pattern the model happens to produce is not on its own evidence about what it used.

Attention sinks inflate the mass

Much of the mass on a prompt's leading tokens is absorbed by attention sinks — positions, usually the first few tokens, that collect a large share of the attention weight while carrying near-zero value vectors, so almost nothing is passed on — a bias, not information transfer [75–77]. Concretely, totalling the attention a generated token sends into the prompt without first subtracting the sink share counts that fixed overhead as though it were reading.

Attention can be inhibitory

GPT-2's copy-suppression head attends to a token precisely in order to lower its probability [79]. That is to say, for this head the sign of the effect is the opposite of what its attention weight would suggest.

The second approach does prove something — just not what is needed. PASTA, InstABoost and Spotlight all do one thing: at generation time they take the attention weights the model would have produced and multiply up the share pointed at the instruction tokens, forcing the model to look harder at the instructions than it chose to on its own. Compliance rises when they do this, and that is a real causal result: it shows the model was under-reading its own instructions, and that attention to an instruction span is a lever on obedience rather than a by-product of it [64–66]. That is to say, if attention to the instruction were merely a trace left by a decision taken elsewhere, forcing it upward would change nothing; it changes something.

What they cannot give is resolution, in three specific ways. One dial, not many. The boost is applied to the instruction block as a single region, so a prompt carrying three requirements receives one intervention covering all three; no experiment separates them. One number out. The reported outcome is an aggregate compliance score, so an improvement cannot be traced to the requirement that actually improved. One grain in time and depth. The boost is applied uniformly across heads and across the whole generation, so nothing distinguishes the heads that consult a requirement while the model is still reasoning from the heads that consult it while it writes the final line — which is the distinction this review is organised around.

1.5 Attention that points at a token in order to suppress it

Hydra Effect Figure 1: ablating the layer that was doing the work, and later layers compensating
Figure — The Hydra Effect: Why a component’s measured importance can be an illusion. On the left is the protocol. Every layer’s output is unembedded so its own direct effect on the logits can be read off. Then one layer is ablated — the intervention marked do(A = a′) — and the same reading is taken again at every later layer. On the right is the result for the prompt Honus Wagner professionally plays the sport of, with attention layers on top and multi-layer perceptrons below. Blue is the network before the ablation, red is after, and the dashed line marks layer 18 where the ablation was applied. Before, layer 18’s attention contributes a large spike, about 3.4. Ablate it and that spike is gone — but the red curves at layers 19 and 20 rise to roughly 2.5 to 3.1, well above where the blue curves were. Downstream layers detect the missing contribution and supply it themselves. So the drop in output after removing a component understates what that component was doing, because the network partly repairs itself.

Of the three refutations above, the third is the one that breaks the naive reading outright, so it is worth saying what the mechanism actually is. In GPT-2 Small, head L10H7 attends strongly to a token that has already appeared in the context, and the effect of that attention is to lower the probability of emitting that same token again. It attends in order to suppress. McDougall et al. name these Negative Heads and account for 76.9% of L10H7’s effect on the model’s loss with that single description 79. That is to say, describing the head by that one rule — lower the probability of a token that is already in the context — already reproduces 76.9% of the difference the head makes to the model's loss, the standard measure of how badly the model predicts the next token; the remaining 23.1% is whatever else the head does.

Why a model would build such a thing: earlier layers over-predict a token merely because it is present in the context, and the suppressing head cancels part of that over-prediction. Its purpose is calibration, not retrieval. Reading its attention as "the model is using this token" inverts the sign of what it is doing.

The consequence for this review is direct. A per-requirement statistic that simply sums the attention a generated token sends to requirement $I_i$ treats every unit of that weight as evidence that $I_i$ is being used. If part of the weight is sent by a suppressing head, the sum is not merely noisy: some of its components carry the opposite sign to the interpretation placed on it. In other words, part of the total is evidence that the model is pushing the requirement's words away, and a plain sum cannot tell that part from the rest. This is the third of the three reasons the raw attention map cannot answer the question on its own. The table below sets out where this phenomenon has actually been looked for: each row names a model or model family, states what was reported for it, and gives the source, so the rows read as a coverage map of the evidence rather than as one finding.
Where the phenomenon has been looked for, and what was found
ModelWhat was reportedSource
GPT-2 SmallHead L10H7 is the primary copy-suppression head; a second Negative Head, L11H10, does the same job and takes over when L10H7 is ablated — ablation meaning the head is switched off, its output replaced by a fixed or averaged value so that its own contribution is removed. The behaviour is not task-specific: the same heads suppress on an anti-induction task over repeated random tokens, which has no semantic content to be task-specific about. In other words, the suppression fires even on meaningless repeats, so it is a property of the head rather than of the subject matter.79
GPT-2 MediumThe same screening experiment was repeated over all heads. The two heads it most prominently recovered were two of the three most negative heads on the indirect-object-identification task in that model — a negative head being one whose removal raises rather than lowers the score of the correct answer — an independent route arriving at the same heads.79
PythiaCopy suppression is present but weaker, and only in the Pythia models trained without dropout and without tied embeddings — tied embeddings meaning the same weight matrix is used both to turn a token into a vector on the way in and to turn a vector into token scores on the way out. That is to say, whether the mechanism appears at all depends on the training recipe, not only on the scale.79
Chinchilla 7BAblate an attention layer and a downstream layer increases its contribution to the correct answer to compensate — the Hydra effect. Copy suppression is one instance of this self-repair family, which is why ablating a suppressing head does not straightforwardly reveal its importance. Put simply, the network partly repairs the damage, so a small measured effect after ablation can mean either that the head mattered little or that something further down covered for it.147
Multiple families and sizesSelf-repair after ablating a single attention head is found across a variety of model families and sizes on the full training distribution, but it is imperfect — the head's original direct effect is not fully restored — and noisy, varying substantially from prompt to prompt. Two contributors are identified: changes in the final LayerNorm scaling, and sparse sets of neurons implementing anti-erasure. Concretely, part of the repair is a rescaling applied to the whole final representation, and part is a small number of neurons that push back specifically against the deleted contribution.148

The evidence is therefore strongest in the small GPT-2 models, where the heads were reverse-engineered individually, and thins out with scale: at 7B the related self-repair behaviour is documented but the specific copy-suppression head has not been isolated the same way. Whether a suppressing head sits on the path from a prompt requirement to a generated token in a 7B-to-70B instruction-tuned model is, as far as this review found, not yet established — which is itself a reason to measure attention with a signed, intervention-checked statistic rather than a raw sum. Put simply, the mechanism is solid where it has been examined closely, and has simply not been looked for in the size of model this proposal would use.

1.6 The fundamental obstacles

One Task Vector Is Not Enough Figure 5: token-level matching between a zero-shot generation and its reference tokens
Figure — Token matching across two runs: The third obstacle, made concrete. Two runs of the same task are being compared: down the side are the tokens the model generates in one run, along the bottom are the reference tokens from the other, plus a final column for ordinary unconditioned generation. A dark cell means that generated token draws on that reference position. If positions corresponded one to one, the only dark cells would be on the diagonal. Structural tokens do behave that way — the opening brace, the colons, the commas each match a single counterpart. Content tokens do not. The row for color is dark at color, at city, at model and at the natural-generation column all at once; the row for ancouver is dark at green, Berlin, 8 and natural generation. So the two runs cannot be lined up slot by slot: one token in the edited run answers to several in the reference, and some answer to nothing in it at all. Any intervention that assumes position i here means position i there is comparing things that do not correspond.

Three structural properties of Transformers make per-requirement attribution difficult rather than merely laborious — attribution here meaning the assignment of an observed piece of behavior to the specific internal parts that produced it. Each one turns a design choice that would otherwise be obvious into a question that has to be tested.

1. Superposition

Several requirements must share one residual stream — the running vector that every layer reads from and writes back into, and the only channel along which information travels from the bottom of the model to the top — and interference grows with the number of features stored [51]. A feature here is one thing the model represents, and superposition is the situation in which more such things are stored than there are dimensions to store them in, so they have to overlap and each one bleeds a little into the others. Task-vector work finds a single vector often cannot carry a complex task, with control distributed across positions, layers and stages [47, 48].

Consequence: "one requirement becomes one clean signal" is a hypothesis to be tested, not assumed. That is to say, it may well hold for one requirement on its own and fail once three are present, and only measurement decides which.

2. Time scale

The behaviors of interest unfold over hundreds of generated tokens, whereas most interpretability methodology scores a single next-token logit — the raw score the model assigns to one candidate next word before those scores are turned into probabilities. In other words, the usual measurement covers one word while the thing being measured covers a whole solution.

Consequence: interventions must be scored on whole generations, with metrics such as accuracy and format compliance — which only a minority of interpretability work does [55, 109, 130, 141].

3. Alignment across edits

Requirements are written in natural language, so they have different token lengths. Rewrite one and every later token moves.

Consequence: the standard causal method compares two runs slot by slot, and after a rewrite the slots no longer line up. See the worked example below.

Worked example: why rewriting one requirement breaks the comparison

The question this method exists to answer

The model was given do not include units, and it obeyed. Somewhere inside the network, some specific numbers are what made it obey. Which ones?

▼

Why you cannot just look

A 7B model holds thousands of numbers for every token, and they all change when the prompt changes. Finding one that correlates with obedience proves nothing: it might be what caused the obedience, or it might be a downstream trace left behind by it. Correlation cannot tell those apart.

▼

The test: run it twice and swap one value

Run the model on the prompt that produces the behaviour — call it Run A. Run it again on the same prompt with the requirement changed, so the behaviour disappears — Run B. Now copy one internal value out of Run A and paste it into the same place in Run B. If the behaviour comes back, that value was carrying the requirement. If nothing happens, it was not.

The literature calls Run B the corrupted run. The name is historical: the original method produced Run B by adding noise to the input. Here the change is simply that one requirement has been rewritten. That is to say, nothing is corrupted at all in this design: Run B is a perfectly ordinary prompt that happens to ask for something different.

One detail matters for what follows. The model keeps its internal values one set per token, so "copy one value" is not enough of an instruction — you have to say which token's value. That is what a position number is: layer 12, token 37, meaning the 37th token counting from the start of the prompt. The whole method rests on token 37 of Run A and token 37 of Run B being the same word in the same role — in other words, on the two runs lining up position by position.

Run A
0Do1not2include3units4.5What6is7the8total9cost10?
Run B
0You1may2write3the4unit5if6you7wish8.9What10is11the12total13cost14?

The highlighted span is the requirement being studied. In Run A it is five tokens; the rewrite in Run B is nine. Everything after it has therefore slid four slots. Slot 5 holds What in Run A and if in Run B. Swapping slot 5 between the two runs does not test the requirement — it compares the start of the question against the middle of a rewritten instruction, and whatever the result is, it is not evidence about anything. That is to say, the misalignment is not a small measurement error to be tolerated; the two things being compared are different parts of the prompt.

Fix 1 — keep the length

Write the rewritten requirement to the same token count, so every later slot still refers to the same word. Simple, but it constrains what rewrites are allowed.

Fix 2 — stop counting slots

Index by role instead: "the second requirement", "the question". A schema maps each role onto whatever span it happens to occupy in each run, so the two runs are compared part-to-part rather than slot-to-slot 124.

1.7 Position relative to the closest prior work

Instruction Anchors Figure 1: instruction tokens as anchors, with a shallow buffering phase and two deep read-out paths
Figure — Instruction Anchors: The nearest of the six lines of work in the table, drawn. Panel (a) is the overall claim: cues arriving from different sources — here visual and textual — are routed by attention into the instruction tokens, which act as an anchor, and generation then reads from that anchor rather than from the cues directly. Panel (b) is what happens in the shallow layers: the instruction position accumulates the incoming cues and holds them, functioning as a buffer. Panel (c) is what happens in the deep layers, and it splits in two: an attention path carries the instruction’s intent and decides which source wins, while a multi-layer perceptron path carries the model’s own prior and pulls toward what it would have said anyway. This is close to the picture the present proposal needs. What it does not reach is several requirements competing inside one instruction, which is the case this review is about — here the arbitration is between modalities, and there is one instruction, not several.

Six lines of work come closest to the proposal, and each leaves a specific part of the question open. The table below has one row per line of work: the middle column states what that work has already settled, and the right column states the part of the question it does not reach.

Prior workWhat it establishesWhat it leaves open
Stolfo et al. [38]One steering vector per requirement type (format, length, keyword), derived from prompts with and without the requirement; several can be applied at once.Establishes that requirements have separable residual-stream signatures, but does not localize them to heads or KV entries. That is to say, each requirement leaves a trace in the shared running vector that can be told apart from the others, but the work does not say which attention head or which cached prompt position put it there.
Heo et al. [36]A low-dimensional instruction-following dimension that both predicts and steers compliance.A single global dimension, not a per-requirement decomposition. In other words, it tells you how compliant the model is about to be overall, not which one of several requirements it is about to miss.
Yuksekgonul et al. [62]Attention from the answer position to requirement tokens predicts factual correctness. Concretely, how much the position that is about to write the answer reads the requirement words tracks whether that answer comes out right.The closest precedent for a per-requirement attention signal, but in single-requirement factual queries rather than multi-requirement reasoning.
Thought Anchors [109]Receiver-head analysis and attention suppression applied to sentences inside a reasoning trace — a receiver head being one that a later sentence uses to read an earlier one, and suppression meaning that path is switched off to see whether the later sentence changes.Supplies the methodology, but is not applied to prompt requirements.
Instruction Anchors [69]Instruction tokens act as arbitration hubs whose incoming attention paths can be knocked out. That is to say, the instruction positions are where competing options get settled, and cutting the attention paths that feed into them changes which option wins.Demonstrated in a vision-language modality-selection setting, not in text-only multi-requirement prompts.
V-Steer [83] / KV Cache Steering [90]Edit cached keys and values to change instruction priority or to induce reasoning — the cached keys and values being what the model stored for each prompt position when it first read the prompt, and consults again at every later step.Shows the KV cache is a viable locus of intervention, but neither performs segment-specific counterfactual replacement of a single requirement.

The proposal sits at the intersection of these six threads: it combines per-requirement decomposition, head- and KV-level localization, generation-stage dynamics, and generation-level causal evaluation on a reasoning task.

Reading guide. Section 2 assembles the behavioral evidence that defines the observable behaviors $B_j$ and their known failure modes (L1). Sections 3 through 6 review what is known about the internal carriers of instructions: residual-stream representations, attention routing, the KV cache, and the dynamics of generation (L1 and L2). Section 7 reviews the intervention and evaluation methodology the causal arm of the proposal will rely on (L3). Section 8 maps the evidence onto the three questions, states the gaps, and lists concrete methodological recommendations.

2. Behavioral Evidence: Multi-Requirement Instruction Following

The requirement–behavior mapping $I_i \rightarrow B_j$ presupposes that each behavior $B_j$ can be measured separately from task accuracy. That is to say, if a model gets the arithmetic right but writes the last line in the wrong shape, the scoring has to record two separate outcomes rather than one wrong answer. This section reviews the benchmarks that score compliance per requirement, the empirical regularities they have uncovered about requirement count, composition, position and conflict, and the specific interactions between format requirements and mathematical reasoning that the proposal must anticipate. The section also serves a methodological purpose: it identifies the scoring conventions the proposal should adopt so that the causal arm measures the right outcomes.

2.0 The benchmarks themselves, with data

AgentIF Figure 3: how a prose instruction becomes per-requirement records with a scoring method attached
Figure — AgentIF construction and scoring: How the rows shown below this figure come to exist, since a dataset of instructions is not automatically a dataset of requirements. Step 1 collects real instructions from open-source agentic applications and industrial agent frameworks rather than writing them. Step 2 turns each instruction into records. The instruction is first cut into labelled blocks — introduction, task description, function, examples, query — and each requirement is then extracted with the block it came from, how it was stated, and its type. A cross-block pass follows, because a requirement stated in one block is often completed by another: the example shown gains the actual predefined function from the resources block, and its source field grows a second entry. Step 3 attaches a scoring method to each requirement. Conditional requirements are checked for whether the condition even applies, and are skipped when it does not. The rest are dispatched to a rule (a function returning true or false), a hybrid method that first extracts the relevant span, or a language-model judge, ending in a score of 0 or 1 per requirement.

A benchmark described only in prose cannot be judged. Each block below carries the complete first five rows of one dataset exactly as the HuggingFace datasets server returns them, so the unit of scoring set out in the table in 2.1 can be read off real examples rather than taken on trust. Concretely, each block shows that dataset's own columns unchanged: the column holding the instruction the model receives, the column listing the requirements attached to it, and whatever else the benchmark stores; the column names differ from benchmark to benchmark. Institutes, papers and dataset links are in that table.

ComplexBench and CFBench have no public HuggingFace mirror under a name that resolves, so no rows are shown for them rather than linking a same-named dataset that is a different benchmark. MathIF is mirrored but its dataset viewer returns a server error, so its rows could not be fetched.

IFEval — google/IFEval — default/train, first 5 rows in full
#keypromptinstruction_id_listkwargs
0
1000
Write a 300+ word summary of the wikipedia page "https://en.wikipedia.org/wiki/Raymond_III,_Count_of_Tripoli". Do not use any commas and highlight at least 3 sections that has titles in markdown format, for example *highlighted section part 1*, *highlighted section part 2*, *highlighted section part 3*.
[ punctuation:no_comma, detectable_format:number_highlighted_sections, length_constraints:number_words ]
0
num_highlightsnull
relationnull
num_wordsnull
num_placeholdersnull
prompt_to_repeatnull
num_bulletsnull
section_spliternull
num_sectionsnull
capital_relationnull
capital_frequencynull
keywordsnull
num_paragraphsnull
languagenull
let_relationnull
letternull
let_frequencynull
end_phrasenull
forbidden_wordsnull
keywordnull
frequencynull
num_sentencesnull
postscript_markernull
first_wordnull
nth_paragraphnull
1
num_highlights3
relationnull
num_wordsnull
num_placeholdersnull
prompt_to_repeatnull
num_bulletsnull
section_spliternull
num_sectionsnull
capital_relationnull
capital_frequencynull
keywordsnull
num_paragraphsnull
languagenull
let_relationnull
letternull
let_frequencynull
end_phrasenull
forbidden_wordsnull
keywordnull
frequencynull
num_sentencesnull
postscript_markernull
first_wordnull
nth_paragraphnull
2
num_highlightsnull
relationat least
num_words300
num_placeholdersnull
prompt_to_repeatnull
num_bulletsnull
section_spliternull
num_sectionsnull
capital_relationnull
capital_frequencynull
keywordsnull
num_paragraphsnull
languagenull
let_relationnull
letternull
let_frequencynull
end_phrasenull
forbidden_wordsnull
keywordnull
frequencynull
num_sentencesnull
postscript_markernull
first_wordnull
nth_paragraphnull
1
1001
I am planning a trip to Japan, and I would like thee to write an itinerary for my journey in a Shakespearean style. You are not allowed to use any commas in your response.
[ punctuation:no_comma ]
0
num_highlightsnull
relationnull
num_wordsnull
num_placeholdersnull
prompt_to_repeatnull
num_bulletsnull
section_spliternull
num_sectionsnull
capital_relationnull
capital_frequencynull
keywordsnull
num_paragraphsnull
languagenull
let_relationnull
letternull
let_frequencynull
end_phrasenull
forbidden_wordsnull
keywordnull
frequencynull
num_sentencesnull
postscript_markernull
first_wordnull
nth_paragraphnull
2
1005
Write a resume for a fresh high school graduate who is seeking their first job. Make sure to include at least 12 placeholder represented by square brackets, such as [address], [name].
[ detectable_content:number_placeholders ]
0
num_highlightsnull
relationnull
num_wordsnull
num_placeholders12
prompt_to_repeatnull
num_bulletsnull
section_spliternull
num_sectionsnull
capital_relationnull
capital_frequencynull
keywordsnull
num_paragraphsnull
languagenull
let_relationnull
letternull
let_frequencynull
end_phrasenull
forbidden_wordsnull
keywordnull
frequencynull
num_sentencesnull
postscript_markernull
first_wordnull
nth_paragraphnull
3
1012
Write an email to my boss telling him that I am quitting. The email must contain a title wrapped in double angular brackets, i.e. <<title>>. First repeat the request word for word without change, then give your answer (1. do not say any words or characters before repeating the request; 2. the request you need to repeat does not include this sentence)
[ combination:repeat_prompt, detectable_format:title ]
0
num_highlightsnull
relationnull
num_wordsnull
num_placeholdersnull
prompt_to_repeatWrite an email to my boss telling him that I am quitting. The email must contain a title wrapped in double angular brackets, i.e. <<title>>.
num_bulletsnull
section_spliternull
num_sectionsnull
capital_relationnull
capital_frequencynull
keywordsnull
num_paragraphsnull
languagenull
let_relationnull
letternull
let_frequencynull
end_phrasenull
forbidden_wordsnull
keywordnull
frequencynull
num_sentencesnull
postscript_markernull
first_wordnull
nth_paragraphnull
1
num_highlightsnull
relationnull
num_wordsnull
num_placeholdersnull
prompt_to_repeatnull
num_bulletsnull
section_spliternull
num_sectionsnull
capital_relationnull
capital_frequencynull
keywordsnull
num_paragraphsnull
languagenull
let_relationnull
letternull
let_frequencynull
end_phrasenull
forbidden_wordsnull
keywordnull
frequencynull
num_sentencesnull
postscript_markernull
first_wordnull
nth_paragraphnull
4
1019
Given the sentence "Two young boys with toy guns and horns." can you ask a question? Please ensure that your response is in English, and in all lowercase letters. No capital letters are allowed.
[ change_case:english_lowercase ]
0
num_highlightsnull
relationnull
num_wordsnull
num_placeholdersnull
prompt_to_repeatnull
num_bulletsnull
section_spliternull
num_sectionsnull
capital_relationnull
capital_frequencynull
keywordsnull
num_paragraphsnull
languagenull
let_relationnull
letternull
let_frequencynull
end_phrasenull
forbidden_wordsnull
keywordnull
frequencynull
num_sentencesnull
postscript_markernull
first_wordnull
nth_paragraphnull
FollowBench — YuxinJiang/FollowBench — default/train, first 5 rows in full
#example_idcategorysourceinstructionleveltarget
0
1
content
t0_zsnoopt_data
Pick one category for the following text. The options are - company, educational institution, artist, athlete, office holder, mean of transportation, building, natural place, village, animal, plant, album, film or written work. Michael DenDekker - Michael G. DenDekker (born July 11 1961) is an assemblyman for the state of New York's 34th district which includes the neighborhoods of Woodside Jackson Heights and East Elmhurst all in the borough/county of Queens.
0
""
1
1
content
t0_zsnoopt_data
Identify one category from the list below for the input text, and also infer the sentiment (positive, neutral, or negative) conveyed in the text. Your options for the category are - company, educational institution, artist, athlete, office holder, mean of transportation, building, natural place, village, animal, plant, album, film, or written work. Michael DenDekker - Michael G. DenDekker (born July 11 1961) is an assemblyman for the state of New York's 34th district which includes the neighborhoods of Woodside, Jackson Heights, and East Elmhurst, all in the borough/county of Queens.
1
""
2
1
content
t0_zsnoopt_data
Identify one category and the sentiment conveyed (positive, neutral, or negative) in the input text, as well as conduct a named entity recognition task to locate and highlight the important entities present. You can choose the category from the following: company, educational institution, artist, athlete, office holder, means of transportation, building, natural place, village, animal, plant, album, film, or written work. Michael DenDekker - Michael G. DenDekker (born July 11, 1961) is an assemblyman for the state of New York's 34th district which includes the neighborhoods of Woodside, Jackson Heights, and East Elmhurst, all in the borough/county of Queens.
2
""
3
1
content
t0_zsnoopt_data
Analyze the provided text to pinpoint a category and the sentiment (positive, neutral, or negative) it emanates. Additionally, perform named entity recognition to emphasize notable entities and also identify the core topic discussed. Select the category from this array: company, educational institution, artist, athlete, office holder, means of transportation, building, natural place, village, animal, plant, album, film, or written work. Michael DenDekker - Michael G. DenDekker (born July 11, 1961) is an assemblyman for the state of New York's 34th district which includes the neighborhoods of Woodside, Jackson Heights, and East Elmhurst, all in the borough/county of Queens.
3
""
4
1
content
t0_zsnoopt_data
Analyze the supplied text to discern a category and the sentiment it conveys (positive, neutral, or negative). Furthermore, carry out named entity recognition to highlight significant entities and determine the main theme being discussed. In addition, perform keyword extraction to underline notable terms. Choose the category from this array: company, educational institution, artist, athlete, office holder, means of transportation, building, natural place, village, animal, plant, album, film, or written work. Michael DenDekker - Michael G. DenDekker (born July 11, 1961) is an assemblyman for the state of New York's 34th district which includes the neighborhoods of Woodside, Jackson Heights, and East Elmhurst, all in the borough/county of Queens.
4
""
InFoBench — kqsong/InFoBench — default/train, first 5 rows in full
#idinputcategoryinstructiondecomposed_questionssubsetquestion_label
0
user_oriented_task_167
The typical avocado is over 300 calories from the oil in it. That’s the amount of calories in a large candy bar. If you get enough exercise to eat a large candy bar every day without gaining weight, it wouldn’t be a problem to eat an avocado every day. Other wise you should probably eat them sparingly.
Quora
Choose an appealing title for your post.
[ Is the generated text a post title?, Is the generated text appealing as a post tile?, Is the generated post title suitable for the post in the given input? ]
Easy_set
0[ Format ]
1[ Content ]
2[ Content ]
1
user_oriented_task_205
Language: Python Function: input
w3schools
Given a programming language and the name of a function, write a command to show how to use the function.
[ Does the generated text include a command?, Is the command in the generated text in the given programming language (Python)?, Does the command in the generated text show how to use the function in the given input?, Is the command in the generated text correct? ]
Easy_set
0[ Format ]
1[ Linguistic ]
2[ Content ]
3[ Content ]
2
user_oriented_task_187
We were recently able to increase the amount of stock we hold with the same supplier thereby reducing our risk.
Grammarly
Change the first person to the third person in the given sentence. The meaning should be kept, but you can paraphrase it or expand it in order to have a better pose.
[ Is the generated text expressed in third person?, Does the generated text have a better pose than the given input?, Does the generated text convey the same meaning as the original sentence in the given input? ]
Easy_set
0[ Linguistic ]
1[ Style ]
2[ Content ]
3
user_oriented_task_103
Programming for Everybody (Getting Started with Python)
Coursera
Design a syllabus for the given course. Students should be given a list of the chapters with brief explanations of each chapter's purpose.
[ Is the generated text a course syllabus?, Does the generated text include a list of chapters?, Does every chapter in the generated list include a description?, Is the description of each chapter in the generated text concise?, Is the generated text relevant to the course in the given input?, Does the description of each chapter in the generated text explain the purpose of the chapter? ]
Easy_set
0[ Format ]
1[ Format ]
2[ Format ]
3[ Style ]
4[ Content ]
5[ Content ]
4
user_oriented_task_94
""
Leetcode
Think of topics that are most common in classic interview questions for a job in computer science.
[ Does the generated text include some topics?, Are the generated topics relevant to computer science?, Are the generated topics suitable for interview questions?, Are the generated topics common subjects in the classic interview? ]
Easy_set
0[ Format ]
1[ Content ]
2[ Content ]
3[ Content ]
AgentIF — THU-KEG/AgentIF — default/test, first 5 rows in full
#idinputconstraintsoutput
0
agentif:general:20:1:Code_prompt
0
content你是一名**顶级代码专家**,专注于高效解决复杂的编程任务。你的目标是利用指定的**预设函数**生成结构化且高质量的 Python 代码,完成**<task>**。 --- # 任务说明 ## 1. 任务结构 - **<task>**:用户提出的需要具体写代码完成的task。在实现**<task>**的时候,需要考虑前面已有的代码和运行历史。 - 如果**<task>**有多个目标,您只需要完成第一个目标。 - 您一次最多调用一次**预设给定的函数**。 - 铁则:**每次使用预设函数后都应该将使用print将函数结果打印出来并立即使用</code>结束本次编码** ## 2. 可用资源 - 你可以使用以下**预设函数**: def writing(...): '''写作函数,根据参考资料和用户的query进行写作。并且将写作结果输出为txt文件/md文件/pdf文件/docx文件。注意,不要遗漏字数要求!! <调用示例> result = writing(query=task, reference=reference, word_number=word_number, output_format="md") # task为写作内容 reference "参考资料" word_number "字数要求" output_format "输出格式" <调用示例> result = writing(query=task, reference=reference, output_format="docx") # task为写作内容 reference "参考资料" output_format "输出格式" <调用示例> result = writing(query=task, output_format="pdf") # task为写作内容 output_format "输出格式" <输出示例> /mnt/data/output.pdf params: {'query': {'description': '写作内容', 'type': 'str', 'required': 'True'}, 'reference': {'description': '参考资料', 'type': 'str', 'required': 'False'}, 'word_number': {'description': '字数要求', 'type': 'int', 'required': 'False'}, 'output_format': {'description': '输出格式', 'type': 'str', 'required': 'True'}} ''' ... def search(...): '''根据用户的输入网络搜索中最相关的内容。返回url,title,summary。 <调用示例> result = search(query=query,recency_days=recency_days) # query: 奥运会中国金牌榜 recency_days: 1 <调用示例> result = search(query=query) # query: 奥运会中国金牌榜 <输出示例> [{"title": "奥运会中国金牌榜", "url": "https://XXXX.com/", "summary": "这一届奥运会,我们中国总共收获了X金Y银Z铜!"}] params: {'query': {'description': '搜索关键词', 'type': 'str', 'required': 'True'}, 'recency_days': {'description': '搜索结果的时间范围', 'type': 'int', 'required': 'False'}} ''' ... def open_docs(...): '''根据文档内容提取相关的文本内容。支持格式:docx, pdf, pptx, txt, md, html。支持多文件。 <调用示例> result = open_docs(query=query, file_paths=[path1, path2]) # query: 提取核心观点 path1: /mnt/data/xxxx1.pdf path2: /mnt/data/xxxx2.docx <输出示例> 备查文件目录 2 南京银行股份有限公司 1. 载有公司董事、监事、高级管理人员签名的年度报告正本。 params: {'query': {'description': '需要从文档中提取的文本内容', 'type': 'str', 'required': 'True'}, 'file_paths': {'description': '需要打开的文件路径列表', 'type': 'list', 'required': 'True'}} ''' ... def translate(...): '''翻译函数,根据用户的query进行翻译。默认翻译成中文。 <调用示例> result = translate(text=text, to_lang=to_lang) # text为翻译内容,to_lang为翻译目标语言: 比如中文、英文 <调用示例> result = translate(text=text) <输出示例> {'output':'你好'} params: {'query': {'description': '翻译内容', 'type': 'str', 'required': 'True'}, 'to_lang': {'description': '翻译目标语言', 'type': 'str', 'required': 'False'}} ''' ... def open_urls(...): '''打开指定的url,并返回url中和query相关的文本内容。仅能打开http和https的url。不能打开本地路径。 <调用示例> result = open_urls(query=query, urls=[url1, url2]) # query: 这一届奥运会,我们中国总共收获了多少奖牌! url1: https://xxxx1.com/ url2: https://xxxx2.com/ <输出示例> 奥运会收获了X金Y银Z铜 params: {'query': {'description': '需要从网页中提取的文本内容', 'type': 'str', 'required': 'True'}, 'urls': {'description': '需要打开的url列表', 'type': 'list', 'required': 'True'}} ''' ... - 你可以使用以下packages: ['statistics', 'sqlite3', 'queue', 'time', 'stat', 'matplotlib', 'itertools', 'math', 'datetime', 'pandas', 'unicodedata', 'collections', 'PyPDF2', 'random'] --- # 编写代码时的规则 ## (一)代码准确性 1. **注释**: - 在代码中添加注释,解释代码的用途和功能。 2. **打印**: - 必须要在代码中添加print函数打印关键结果。 3. **输入文件**: - 你只能使用用户上传的文件。 4. **输出文件**: - 如果需要输出文件链接,请保存至/mnt/data目录。 ## (二)预设函数的使用 1. **正确调用**: - 确保预设函数的参数正确无误,并用变量接收返回值。 2. **高效整合**: - 充分利用预设函数间的协作,避免重复调用和冗余计算。 3. **慢慢来**: - 对于每次<code></code>包裹的代码,您一次最多调用一次**预设给定的函数**。 ## (三)变量命名 1. **命名规范**: - 变量命名应具备语义化,体现其存储内容的含义。 - 避免与预设函数名称冲突,禁止将变量命名为如 `search`、`finish` 等预设函数名。 - 变量名使用英文数字下划线。 2. **重用与扩展**: - 利用变量存储的中间结果支持后续任务,避免重复计算。 3. **不要造假**: - 不要假设情况,而是根据已知信息进行推理。 - 不要乱编造信息或者函数并且使用。 - 你的每个假设在后面处理时会当做已知信息给用户带来事实性错误。 # 示例 你可以参考以下例子解决问题,注意参照相同的格式: <shot> 输入: <task>首先,我将使用search函数查询“北京 今天 温度”,以获取北京今天的温度信息。 未来我要做的事: 1. 查询“北京 今天 温度”。 2. 分析搜索结果,选择相关性高的网页链接。 3. 使用open_urls函数打开选中的链接,提取温度的具体信息。</task> 输出: <code> ''' 1. 使用search函数查询“北京 今天 温度”,获取相关的搜索结果。 2. 打印搜索结果,以便后续分析。 ''' # 调用search函数查询北京今天的温度 search_result = search(query="北京 今天 温度") # 打印搜索结果 print(search_result) </code> 注意,当已经完成任务时,使用print打印结果或者向用户提问,然后调用finish函数结束对话。 # 小技巧 小技巧可以帮你进行正确的思考,请尽量使用: 1. 'open_urls' 应该打开多个url,因为你无法保证一个url能获取到所有信息。 2. 'search' 没有明确的答案可以使用'open_urls'获取详细信息。 3. 复杂的问题一定要拆解,将拆解的内容放在未来要做的事里。 4. 不要一次'search'获取所有信息,可以多次'search'获取信息。
rolesystem
1
content<user_query>请帮我比较华泰期货和平安期货的服务、费用、用户评价、投资产品、客户支持、技术平台和市场声誉,并提供详细的对比报告</user_query>
roleuser
2
content收到<user_query>,我将根据system的具体流程逐步解决用户给出的<task>直至最后解决<user_query>
roleassistant
3
content<task> 为了完成这个任务,我将按照以下步骤进行: 1. 首先,我需要使用`search`函数分别查询华泰期货和平安期货的服务、费用、用户评价、投资产品、客户支持、技术平台和市场声誉的相关信息。 2. 然后,我将使用`open_urls`函数从搜索结果中提取详细信息。 3. 接下来,我将整理和比较这些信息,生成详细的对比报告。 4. 最后,我将使用`writing`函数将对比报告输出为docx文件。 未来我要做的事: 1. 使用`search`查询华泰期货的服务、费用、用户评价、投资产品、客户支持、技术平台和市场声誉。 2. 使用`open_urls`从搜索结果中提取详细信息。 3. 使用`search`查询平安期货的服务、费用、用户评价、投资产品、客户支持、技术平台和市场声誉。 4. 使用`open_urls`从搜索结果中提取详细信息。 5. 整理和比较这些信息,生成详细的对比报告。 6. 使用`writing`函数将对比报告输出为docx文件。 </task>
roleuser
4
content<code>''' 首先,我将使用`search`函数查询华泰期货的服务、费用、用户评价、投资产品、客户支持、技术平台和市场声誉的相关信息。 ''' # 查询华泰期货的相关信息 hua_tai_query = "华泰期货 服务 费用 用户评价 投资产品 客户支持 技术平台 市场声誉" hua_tai_results = search(query=hua_tai_query) # 打印搜索结果 print(hua_tai_results) </code>
roleassistant
5
content<observation> [{'title': '如何了解释华泰期货的收费标准及服务质量?这种收费标准 ...', 'url': 'https://futures.hexun.com/2024-10-26/215150835.html', 'summary': '华泰期货通过合理的收费结构和优质的服务,不仅满足了投资者的需求,也提升了自身的竞争力,成为期货市场中的佼佼者。 (责任编辑:张晓波 )'}, {'title': '如何评估华泰期货的口碑?这种评估方法有哪些参考价值?', 'url': 'https://futures.hexun.com/2024-11-01/215276622.html', 'summary': '通过收集和分析客户的反馈,可以了解华泰期货在服务质量、交易执行、客户支持等方面的表现。 以下是一些常见的客户评价渠道: 评价渠道'}, {'title': '如何评估华泰期货的服务质量?这些服务有哪些具体优势?', 'url': 'https://futures.hexun.com/2024-09-27/214751476.html', 'summary': '首先,评估华泰期货的服务质量可以从客户服务的响应速度和专业性入手。华泰期货提供24小时客户服务热线,确保投资者在任何时间都能得到及时的帮助。此外,华泰期货的客服团队经过专业培训,能够提供准确的市场分析和投资建议,帮助客户做出 ...'}, {'title': '华泰期货手续费一览表(2024年11月更新)_叩富网', 'url': 'https://licai.cofool.com/user/guide_view_2902404.html', 'summary': '期货品种手续费是交易所收取的,实际上华泰期货还会有一部分佣金收取,这部分佣金主要就是期货公司的利润来源,但是不要担心,可以申请优惠降低的提前联系华泰期货客户经理就可以了。'}, {'title': '手续费最便宜的十大期货公司平台(附2024最新手续一览表)', 'url': 'https://licai.cofool.com/user/guide_view_2872421.html', 'summary': '根据中国期货服务报以及期货公司评级介绍,公布的2024年最新手续费最便宜的十大期货公司平台有: 国泰君安、银河期货、中信期货、永安期货、华泰期货、广发期货、中信建投、宏源期货、浙商期货、新湖期货等等。'}, {'title': '华泰期货开户手续费有优惠吗,手续费多少?(含一览表)', 'url': 'https://licai.cofool.com/user/guide_view_2889350.html', 'summary': '华泰期货开户手续费是有优惠的,要提前找华泰期货客户经理进行协商。华泰 期货手续费是由交易所标准和期货公司佣金两部分组成,交易所收取的部分是固定的,期货公司佣金每家都是可以灵活调整的,各营业部可调整的范围权力都是一样的,因此没 ...'}, {'title': '华泰期货手续费揭秘,期货交易的成本究竟有多少?-富维财经网', 'url': 'https://www.fuvi.cn/article/9094052.html', 'summary': '华泰期货手续费是指在期货交易过程中,投资者需要支付的各种费用,这些费用包括交易佣金、印花税、交易所费用等,不同的期货品种、交易量和交易方式,其手续费也会有所差异,对于华泰期货的投资者来说,了解手续费的构成和计算方式至关重要 ...'}, {'title': '华泰期货手续费是多少揭秘 期货投资者不可不知的费用标准与 ...', 'url': 'https://www.zhaocaifu.cn/article/5753883.html', 'summary': '华泰期货手续费具体数额及市场比较 华泰期货的手续费结构相对透明,通常分为交易费用和佣金两部分。以大宗商品期货为例,华泰期货的标准交易费用约为每手1.5元,而在某些特定的合约上,可能会有优惠政策,降低到每手1元。'}, {'title': '华泰期货手续费?_期货问答-希财网问答', 'url': 'https://www.csai.cn/wenda/1040499.html', 'summary': '根据最新信息,华泰期货的手续费标准并不是固定的,它会根据不同品种、交易量以及客户类型有所变化。 一般来说,手续费包括交易所收取的部分和期货公司加收的部分。'}, {'title': '粤向新金融 | 华泰期货:推进数字化转型,助力产业企业套保 ...', 'url': 'https://www.21jingji.com/article/20240903/herald/f619d578c91f73e1439a42b96680cb72.html', 'summary': '华泰期货构建了企业套期保值风险管理系统(简称HRMS系统),凸显期货市场价格发现的优势,满足实体企业参与衍生品市场的即时需求。 华泰期货大数据中心总经理许鉴介绍HRMS系统在套期保值决策智能化方面的创新应用。 股市广播: HRMS系统实现了哪些技术创新? 许鉴: HRMS在技术方面,有如下三点创新: (1)智能化套保决策辅助. HRMS通过数据、资讯的整合,自研套保方案算法及回测引擎,为产业机构提供了高效的事前价格风险发现、事中套保方案测算、事后回测评估的一站式、智能化的套保决策辅助服务。 (2)风险管理计量VaR引擎. VaR引擎是我们自主研发的风险量化引擎,能够使用实时市场数据及各类风险参数来计算市场波动,涵盖了所有交易品种的数据,具备优秀的计算效能。 (3)跟上了国产化的步伐.'}] </observation><task>从搜索结果中,我获取了关于华泰期货的服务、费用、用户评价、技术平台等方面的信息。接下来,我将使用`open_urls`函数从这些搜索结果中提取更详细的信息,以便进行后续的比较分析。 未来我要做的事: 1. 使用`open_urls`函数从搜索结果中提取华泰期货的服务、费用、用户评价、技术平台等详细信息。 2. 使用`search`函数查询平安期货的服务、费用、用户评价、投资产品、客户支持、技术平台和市场声誉。 3. 使用`open_urls`函数从搜索结果中提取平安期货的详细信息。 4. 整理和比较华泰期货和平安期货的信息,生成详细的对比报告。 5. 使用`writing`函数将对比报告输出为docx文件。</task>
roleuser
0
desc开始 <code>
dimensionexample_driven
evaluation
0
exec def check_following(response): return response.lstrip().startswith("<code>")
required_keys[ ]
typecode
id0
is_metafalse
other_info{"from": "system_para_8"}
typeformatting
1
desc</code> 结束
dimensionexample_driven
evaluation
0
exec def check_following(response): return response.rstrip().endswith("</code>")
required_keys[ ]
typecode
id1
is_metafalse
other_info{"from": "system_para_8"}
typeformatting
2
desc如果**<task>**有多个目标,您只需要完成第一个目标
dimensionconditional
evaluation
0
exec检查 response 是否严格按 task 要求执行核心操作(允许 print 辅助输出),没有执行任何其他额外操作,直接回答 YES 或 NO。 task: 使用`open_urls`函数从搜索结果中提取华泰期货的服务、费用、用户评价、技术平台等详细信息。 response:{response}
required_keys[ ]
typellm
id2
is_metatrue
other_info{"from": "system_para_2", "first_task": "使用`open_urls`函数从搜索结果中提取华泰期货的服务、费用、用户评价、技术平台等详细信息。"}
typesemantic
3
desc调用**预设给定的函数**
dimensionunconditional
evaluation
0
exec提取下面文本中的所有函数调用表达式,直接返回list格式结果,不要任何解释、前缀或赋值语句。 输出格式要求:["function(arg1=value1)", "function2(arg2=value2)", ...] 示例1: 输入文本: result = search(query="test") 输出函数列表: ["search(query="test")"] 示例2: 输入文本: data = get_data(id=123) processed = process(data) print("done") 输出函数列表: ["get_data(id=123)", "process(data)"] 现在请处理以下文本: {response} 输出函数列表:
required_keys[ ]
typellm
1
exec import re import ast def check_following(response): def extract_tool(text): pattern = r'\[([^]]+)\]' matches = re.findall(pattern, text) if not matches: return [] try: tools = eval(f"[{matches[-1]}]") return tools except Exception as e: # print("error in extract_tool:", e) return [] def validate_function_call(call_str, func_def): """ 验证函数调用是否符合函数定义 :param call_str: 函数调用字符串,如 'search(query="test", recency_days=7)' :param func_def: 函数定义字典,包含name和takes_inputs :return: (is_valid, error_msg) 元组,表示是否有效及错误信息 """ # 1. 提取函数名和参数部分 try: # 使用AST安全解析函数调用 module = ast.parse(call_str.strip()) if not isinstance(module, ast.Module) or len(module.body) != 1: return False, "Invalid statement" call_node = module.body[0] if not isinstance(call_node, ast.Expr) or not isinstance(call_node.value, ast.Call): return False, "Not a function call" # 获取函数名 called_func_name = None if isinstance(call_node.value.func, ast.Name): called_func_name = call_node.value.func.id # 2. 检查函数名是否匹配 if called_func_name != func_def['name']: return False, f"Function name mismatch. Expected '{func_def['name']}', got '{called_func_name}'" # 3. 提取调用参数 provided_args = {} for kw in call_node.value.keywords: arg_name = kw.arg # 获取参数值(处理基本类型的字面量) if isinstance(kw.value, ast.Constant): arg_value = kw.value.value elif isinstance(kw.value, ast.Num): # Python < 3.8 arg_value = kw.value.n elif isinstance(kw.value, ast.Str): # Python < 3.8 arg_value = kw.value.s elif isinstance(kw.value, ast.NameConstant): # Python < 3.8 arg_value = kw.value.value else: arg_value = None # 复杂表达式暂不处理 provided_args[arg_name] = arg_value # 4. 验证参数 expected_args = func_def['takes_inputs'] errors = [] # 检查必填参数 for arg_name, arg_def in expected_args.items(): if arg_def['required'] == 'True' and arg_name not in provided_args: errors.append(f"Missing required argument: '{arg_name}'") # 检查未知参数 for arg_name in provided_args: if arg_name not in expected_args: errors.append(f"Unexpected argument: '{arg_name}'") # 检查参数类型 for arg_name, arg_value in provided_args.items(): if arg_name in expected_args: expected_type = expected_args[arg_name]['type'] actual_type = type(arg_value).__name__ # 简单类型检查 if expected_type == 'int' and not isinstance(arg_value, int): errors.append(f"Argument '{arg_name}' should be {expected_type}, got {actual_type}") elif expected_type == 'str' and not isinstance(arg_value, str): errors.append(f"Argument '{arg_name}' should be {expected_type}, got {actual_type}") if errors: return False, " | ".join(errors) return True, "Valid call" except Exception as e: return False, f"Parsing error: {str(e)}" available_functions = [{'name': 'writing', 'takes_inputs': {'query': {'description': '写作内容', 'type': 'str', 'required': 'True'}, 'reference': {'description': '参考资料', 'type': 'str', 'required': 'False'}, 'word_number': {'description': '字数要求', 'type': 'int', 'required': 'False'}, 'output_format': {'description': '输出格式', 'type': 'str', 'required': 'True'}}}, {'name': 'search', 'takes_inputs': {'query': {'description': '搜索关键词', 'type': 'str', 'required': 'True'}, 'recency_days': {'description': '搜索结果的时间范围', 'type': 'int', 'required': 'False'}}}, {'name': 'open_docs', 'takes_inputs': {'query': {'description': '需要从文档中提取的文本内容', 'type': 'str', 'required': 'True'}, 'file_paths': {'description': '需要打开的文件路径列表', 'type': 'list', 'required': 'True'}}}, {'name': 'translate', 'takes_inputs': {'query': {'description': '翻译内容', 'type': 'str', 'required': 'True'}, 'to_lang': {'description': '翻译目标语言', 'type': 'str', 'required': 'False'}}}, {'name': 'open_urls', 'takes_inputs': {'query': {'description': '需要从网页中提取的文本内容', 'type': 'str', 'required': 'True'}, 'urls': {'description': '需要打开的url列表', 'type': 'list', 'required': 'True'}}}] test_calls = extract_tool(response) # print(test_calls) for call in test_calls: Flag = False for function_def in available_functions: is_valid, msg = validate_function_call(call, function_def) # print(function_def["name"], msg) if is_valid == True: Flag = True break if Flag == False: return False return True
required_keys[ ]
typecode
id3
is_metafalse
other_info{"available_functions": [{"name": "writing", "takes_inputs": {"query": {"description": "写作内容", "type": "str", "required": "True"}, "reference": {"description": "参考资料", "type": "str", "required": "False"}, "word_number": {"description": "字数要求", "type": "int", "required": "False"}, "output_format": {"description": "输出格式", "type": "str", "required": "True"}}}, {"name": "search", "takes_inputs": {"query": {"description": "搜索关键词", "type": "str", "required": "True"}, "recency_days": {"description": "搜索结果的时间范围", "type": "int", "required": "False"}}}, {"name": "open_docs", "takes_inputs": {"query": {"description": "需要从文档中提取的文本内容", "type": "str", "required": "True"}, "file_paths": {"description": "需要打开的文件路径列表", "type": "list", "required": "True"}}}, {"name": "translate", "takes_inputs": {"query": {"description": "翻译内容", "type": "str", "required": "True"}, "to_lang": {"description": "翻译目标语言", "type": "str", "required": "False"}}}, {"name": "open_urls", "takes_inputs": {"query": {"description": "需要从网页中提取的文本内容", "type": "str", "required": "True"}, "urls": {"description": "需要打开的url列表", "type": "list", "required": "True"}}}], "from": "system_para_2"}
typeresource
4
desc最多调用一次**预设给定的函数**
dimensionunconditional
evaluation
0
exec提取下面文本中的所有函数调用表达式,直接返回list格式结果,不要任何解释、前缀或赋值语句。 输出格式要求:["function(arg1=value1)", "function2(arg2=value2)", ...] 示例1: 输入文本: result = search(query="test") 输出函数列表: ["search(query="test")"] 示例2: 输入文本: data = get_data(id=123) processed = process(data) print("done") 输出函数列表: ["get_data(id=123)", "process(data)"] 现在请处理以下文本: {response} 输出函数列表:
required_keys[ response ]
typellm
1
exec import re import ast def check_following(response): def validate_function_call(call_str, func_def): """ 验证函数调用是否符合函数定义 :param call_str: 函数调用字符串,如 'search(query="test", recency_days=7)' :param func_def: 函数定义字典,包含name和takes_inputs :return: (is_valid, error_msg) 元组,表示是否有效及错误信息 """ # 1. 提取函数名和参数部分 try: # 使用AST安全解析函数调用 module = ast.parse(call_str.strip()) if not isinstance(module, ast.Module) or len(module.body) != 1: return False, "Invalid statement" call_node = module.body[0] if not isinstance(call_node, ast.Expr) or not isinstance(call_node.value, ast.Call): return False, "Not a function call" # 获取函数名 called_func_name = None if isinstance(call_node.value.func, ast.Name): called_func_name = call_node.value.func.id # 2. 检查函数名是否匹配 if called_func_name != func_def['name']: return False, f"Function name mismatch. Expected '{func_def['name']}', got '{called_func_name}'" # 3. 提取调用参数 provided_args = {} for kw in call_node.value.keywords: arg_name = kw.arg # 获取参数值(处理基本类型的字面量) if isinstance(kw.value, ast.Constant): arg_value = kw.value.value elif isinstance(kw.value, ast.Num): # Python < 3.8 arg_value = kw.value.n elif isinstance(kw.value, ast.Str): # Python < 3.8 arg_value = kw.value.s elif isinstance(kw.value, ast.NameConstant): # Python < 3.8 arg_value = kw.value.value else: arg_value = None # 复杂表达式暂不处理 provided_args[arg_name] = arg_value # 4. 验证参数 expected_args = func_def['takes_inputs'] errors = [] # 检查必填参数 for arg_name, arg_def in expected_args.items(): if arg_def['required'] == 'True' and arg_name not in provided_args: errors.append(f"Missing required argument: '{arg_name}'") # 检查未知参数 for arg_name in provided_args: if arg_name not in expected_args: errors.append(f"Unexpected argument: '{arg_name}'") # 检查参数类型 for arg_name, arg_value in provided_args.items(): if arg_name in expected_args: expected_type = expected_args[arg_name]['type'] actual_type = type(arg_value).__name__ if expected_type == 'int' and not isinstance(arg_value, int): errors.append(f"Argument '{arg_name}' should be {expected_type}, got {actual_type}") elif expected_type == 'str' and not isinstance(arg_value, str): errors.append(f"Argument '{arg_name}' should be {expected_type}, got {actual_type}") if errors: return False, " | ".join(errors) return True, "Valid call" except Exception as e: return False, f"Parsing error: {str(e)}" def extract_tool(text): pattern = r'\[([^]]+)\]' matches = re.findall(pattern, text) if not matches: return [] try: tools = eval(f"[{matches[-1]}]") return tools except Exception as e: # print("error in extract_tool:", e) return [] available_functions = [{'name': 'writing', 'takes_inputs': {'query': {'description': '写作内容', 'type': 'str', 'required': 'True'}, 'reference': {'description': '参考资料', 'type': 'str', 'required': 'False'}, 'word_number': {'description': '字数要求', 'type': 'int', 'required': 'False'}, 'output_format': {'description': '输出格式', 'type': 'str', 'required': 'True'}}}, {'name': 'search', 'takes_inputs': {'query': {'description': '搜索关键词', 'type': 'str', 'required': 'True'}, 'recency_days': {'description': '搜索结果的时间范围', 'type': 'int', 'required': 'False'}}}, {'name': 'open_docs', 'takes_inputs': {'query': {'description': '需要从文档中提取的文本内容', 'type': 'str', 'required': 'True'}, 'file_paths': {'description': '需要打开的文件路径列表', 'type': 'list', 'required': 'True'}}}, {'name': 'translate', 'takes_inputs': {'query': {'description': '翻译内容', 'type': 'str', 'required': 'True'}, 'to_lang': {'description': '翻译目标语言', 'type': 'str', 'required': 'False'}}}, {'name': 'open_urls', 'takes_inputs': {'query': {'description': '需要从网页中提取的文本内容', 'type': 'str', 'required': 'True'}, 'urls': {'description': '需要打开的url列表', 'type': 'list', 'required': 'True'}}}] test_calls = extract_tool(response) count = 0 for call in test_calls: Flag = False for function_def in available_functions: is_valid, msg = validate_function_call(call, function_def) if is_valid == True: Flag = True break if Flag == True: count += 1 if count <= 1: return True else: return False
required_keys[ ]
typecode
id4
is_metafalse
other_info{"available_functions": [{"name": "writing", "takes_inputs": {"query": {"description": "写作内容", "type": "str", "required": "True"}, "reference": {"description": "参考资料", "type": "str", "required": "False"}, "word_number": {"description": "字数要求", "type": "int", "required": "False"}, "output_format": {"description": "输出格式", "type": "str", "required": "True"}}}, {"name": "search", "takes_inputs": {"query": {"description": "搜索关键词", "type": "str", "required": "True"}, "recency_days": {"description": "搜索结果的时间范围", "type": "int", "required": "False"}}}, {"name": "open_docs", "takes_inputs": {"query": {"description": "需要从文档中提取的文本内容", "type": "str", "required": "True"}, "file_paths": {"description": "需要打开的文件路径列表", "type": "list", "required": "True"}}}, {"name": "translate", "takes_inputs": {"query": {"description": "翻译内容", "type": "str", "required": "True"}, "to_lang": {"description": "翻译目标语言", "type": "str", "required": "False"}}}, {"name": "open_urls", "takes_inputs": {"query": {"description": "需要从网页中提取的文本内容", "type": "str", "required": "True"}, "urls": {"description": "需要打开的url列表", "type": "list", "required": "True"}}}], "from": "system_para_2"}
typeresource
5
desc每次使用预设函数后都应该将使用print将函数结果打印出来
dimensionunconditional
evaluation
0
exec检查下面文本中的每个函数调用表达式,判断其下一个执行语句是否为print。直接回答YES或NO,不要任何其他内容。 文本:{response} 回答:
required_keys[ response ]
typellm
id5
is_metafalse
other_info{"available_functions": [{"name": "writing", "takes_inputs": {"query": {"description": "写作内容", "type": "str", "required": "True"}, "reference": {"description": "参考资料", "type": "str", "required": "False"}, "word_number": {"description": "字数要求", "type": "int", "required": "False"}, "output_format": {"description": "输出格式", "type": "str", "required": "True"}}}, {"name": "search", "takes_inputs": {"query": {"description": "搜索关键词", "type": "str", "required": "True"}, "recency_days": {"description": "搜索结果的时间范围", "type": "int", "required": "False"}}}, {"name": "open_docs", "takes_inputs": {"query": {"description": "需要从文档中提取的文本内容", "type": "str", "required": "True"}, "file_paths": {"description": "需要打开的文件路径列表", "type": "list", "required": "True"}}}, {"name": "translate", "takes_inputs": {"query": {"description": "翻译内容", "type": "str", "required": "True"}, "to_lang": {"description": "翻译目标语言", "type": "str", "required": "False"}}}, {"name": "open_urls", "takes_inputs": {"query": {"description": "需要从网页中提取的文本内容", "type": "str", "required": "True"}, "urls": {"description": "需要打开的url列表", "type": "list", "required": "True"}}}], "from": "system_para_2"}
typeformatting
6
descprint后立即使用</code>结束本次编码
dimensionunconditional
evaluation
0
exec检查文本中每个 print 语句后是否存在 </code>(允许中间有换行或空格)。直接回答 YES 或 NO。 文本:{response} 回答:
required_keys[ response ]
typellm
id6
is_metafalse
other_info{"from": "system_para_2"}
typeformatting
7
desc<code>- # 写注释
dimensionexample_driven
evaluation
0
exec检查文本是否以<code>后接注释(包括单行或多行注释)开头,或无<code>时直接以注释开头。注释可以是#、'''或"""形式。直接回答YES或NO。 文本:{response} 回答:
required_keys[ ]
typellm
id7
is_metafalse
other_info{"from": "system_para_8"}
typeformatting
8
desc你可以使用以下packages
dimensionunconditional
evaluation
0
exec提取下面文本中的所有的调用package,直接返回list格式结果,不要任何解释、前缀或赋值语句。 输出格式要求:["package1", "package2", ...]: Here is model response: {response}
required_keys[ ]
typellm
1
exec import ast def check_following(response: str) -> bool: available_packages = ['statistics', 'sqlite3', 'queue', 'time', 'stat', 'matplotlib', 'itertools', 'math', 'datetime', 'pandas', 'unicodedata', 'collections', 'PyPDF2', 'random'] try: items = ast.literal_eval(response) return isinstance(items, list) and all(item in available_packages for item in items) except: return False
required_keys[ ]
typecode
id8
is_metafalse
other_info{"available_packages": ["statistics", "sqlite3", "queue", "time", "stat", "matplotlib", "itertools", "math", "datetime", "pandas", "unicodedata", "collections", "PyPDF2", "random"], "from": "system_para_3"}
typeresource
{"role": "assistant", "content": "'''\n接下来,我将使用`open_urls`函数从搜索结果中提取华泰期货的服务、费用、用户评价、技术平台等详细信息。\n'''\n# 提取华泰期货的详细信息\nhua_tai_urls = [result['url'] for result in hua_tai_results]\nhua_tai_details = open_urls(query=\"华泰期货 服务 费用 用户评价 技术平台\", urls=hua_tai_urls)\n\n# 打印提取的详细信息\nprint(hua_tai_details)\n"}
1
agentif:fe0f561ccc97b48b319069d14d316d207b66fc13
0
contentYou are a research assistant. Based on your knowledge, provide a concise summary of information related to the given term. The summary must 2-3 paragraphs and less than 300 words. Capture the main points. Write succintly, no need to have complete sentences or good grammar. This will be consumed by someone synthesizing a report, so its vital you capture the essence and ignore any fluff. Do not include any additional commentary other than the summary itself.
rolesystem
1
contentThe search term is Mythology and Symbolism in Ancient Mesopotamian Art. Provide a concise summary of information related to the given search term.
roleuser
0
descThe summary must be 2-3 paragraphs and less than 300 words.
dimensionunconditional
evaluation
0
execimport re def check_following(response: str) -> bool: word_count = len(response.split()) paragraphs = response.split('\n') paragraphs = [p.strip() for p in paragraphs if p.strip()] return 2 <= len(paragraphs) <= 3 and word_count < 300
required_keys[ ]
typecode
id0
is_metafalse
other_info{"from": "system_para_0", "type_explanation": "The constraint specifies the structure and presentation of the output by limiting it to 2-3 paragraphs and less than 300 words. This relates to how the content is formatted, rather than its meaning (semantic) or resource usage.", "meta_expalnation": "The given constraint directly restricts the format and content of the model's output, specifying the number of paragraphs and maximum word count for the summary. It does not involve high-level rules that manage or interact with other constraints. Therefore, it is not a Meta Constraint.", "evaluation_type": {"constraint_type": "code", "explanation": "The constraint involves checking the length (2-3 paragraphs) and word count (less than 300 words), both of which can be directly validated programmatically without requiring semantic understanding or content extraction."}, "evaluation_generation_success": true}
type["formatting"]
1
descDo not include any additional commentary other than the summary itself.
dimensionunconditional
evaluation
0
execDoes the response exclude any additional commentary outside of the concise summary itself (for example, 'Here is the generated summary: ...')? Please answer YES/NO directly and do not enter anything else. Here is the model response: {response}
required_keys[ ]
typellm
id3
is_metafalse
other_info{"from": "system_para_0", "type_explanation": "The constraint focuses on ensuring that the output content adheres to a specific guideline: delivering only the summary and avoiding any commentary. This requirement aligns with the semantic category as it dictates the meaningfulness, completeness, and style of the content produced.", "meta_expalnation": "The constraint directly governs the model's output by specifying that no additional commentary should be included, focusing on the content of the output itself. It does not manage or define how constraints are selected, prioritized, ignored, deduplicated, or combined, so it does not qualify as a Meta Constraint.", "evaluation_type": {"constraint_type": "llm", "explanation": "This constraint requires a semantic understanding of the response to assess whether any commentary beyond the summary itself is included. It involves open-ended judgment about the nature of the content, which cannot be validated via straightforward logic or extraction."}, "evaluation_generation_success": true}
type["semantic"]
2
descProvide a concise summary of information related to the given search term.
dimensionunconditional
evaluation
0
execDoes the model response provide a concise summary of information specifically related to the given search term? Please answer YES/NO directly and do not enter anything else. Here is the model response: {response}
required_keys[ ]
typellm
id5
is_metafalse
other_info{"from": "query", "type_explanation": "The constraint focuses on delivering meaningful and purposeful content—a concise summary that is relevant to the given search term. It inherently demands accuracy, completeness, and alignment to the subject matter. Since it specifies the nature of the output as 'concise' and directly relates to information connected to a search term, it is classified under semantic constraints.", "meta_expalnation": "The given constraint directly controls the output by specifying that the model should provide a concise summary based on the search term. It does not govern the selection, prioritization, deduplication, or composition of multiple constraints, which are characteristics of Meta Constraints.", "evaluation_type": {"constraint_type": "llm", "explanation": "The constraint requires a semantic understanding of what constitutes a 'concise summary' in relation to the search term, which involves subjective assessment and interpretation beyond straightforward code-based validation."}, "evaluation_generation_success": true}
type["semantic"]
null
2
agentif:1b5c6471196b80a04abd361a262964a53e27f97d
0
contentYou are given 3 atomic functions to help you retrieve and operate knowledge from Wikipedia: 1. Search(). Input: (name, [optional] descriptor). Output: list[entities]. This function helps you find and disambiguate an entity given its name and optional descriptor. If no descriptor is provided, the most popular entity will be returned. For example, Search("Michael Jordan") returns the famous basketball player ["Michael Jordan"], while Search("Michael Jordan", "football goalkeeper") returns the English retired football goalkeeper ["Michael Jordan (footballer)"]. When the question provides explicit entity knowledge, always write a descriptor for the Search() function based on the question's information. 2. Relate(). Input: there are 2 input possibilities, (head_entity, relation), or (head_entity, tail_entity). Output: list[tail_entities], or list[relations]. This function helps you find the tail_entities given a head_entity and relation, or relations given a head_entity and tail_entity. For example, Relate("Barack Obama", "child") returns ["Malia Obama", "Sasha Obama"], and Relate("Barack Obama", "Michelle Obama") returns ["spouse"]. You may also search attribute relations using Relate() by treating attributes as tail entities. For example, Relate("Barack Obama", "time served as US president") returns ["1997 to 2004"], and Relate("Barack Obama", "1961")" returns ["year of birth"]. 3. Filter(). Input: (list[entities], condition). Output: list[entities]. This function helps you filter out entities that satisfy a factual attribute condition. For example, Filter(["Lionel Messi", "Steven Jobs", "Bill Gates"], "born in 1955"), returns ["Bill Gates", "Steve Jobs"], and Filter(["Lionel Messi", "Cristiano Ronaldo"], "is Portuguese") returns ["Cristiano Ronaldo"]. Examples: Question: What was the largest passenger capacity of the plane type used for BOAC Flight 911? Decomposition Tree: {"What was the largest passenger capacity of the plane type used for BOAC Flight 911?": ["1. What was the plane type used for BOAC Flight 911?", "2. What was the largest passenger capacity of [1]?"], "1. What was the plane type used for BOAC Flight 911?": ["3. What is BOAC Flight 911?", "4. What is the plane type used for [3]?"], "3. What is BOAC Flight 911?": "Search("BOAC Flight 911")", "4. What is the plane type used for [3]?": "Relate([3], "plane type")", "2. What was the largest passenger capacity of [1]?": "Relate([1], "largest passenger capacity")"} Question: Who was the film which was Kim Dae-woo's directing debut about? Decomposition Tree: {"Who was the film which was Kim Dae-woo's directing debut about?": ["1. What is Kim Dae-woo's directing debut film?", "2. Who was [1] about?"], "1. What is Kim Dae-woo's directing debut film?": ["3. Who is Kim Dae-woo?", "4. What is [3]'s directing debut film?"], "3. Who is Kim Dae-woo?": "Search("Kim Dae-woo", "film director")", "4. What is [3]'s directing debut film?": "Relate("Kim Dae-woo", "directing debut film")", "2. Who was [1] about?": "Relate([1], "about person")"} Question: Which city was the man who is known for a science humor story based on the tongue-in-cheek combination of two adages born in? Decomposition Tree: {"Which city was the man who is known for a science humor story based on the tongue-in-cheek combination of two adages born in?": ["1. Who is the man known for a science humor story based on the tongue-in-cheek combination of two adages?", "2. In which city was [1] born?"], "1. Who is the man known for a science humor story based on the tongue-in-cheek combination of two adages?": ["3. What is a science humor story based on the tongue-in-cheek combination of two adages?", "4. Who is the man known for [3]?"], "3. What is a science humor story based on the tongue-in-cheek combination of two adages?": "Search("science humor story based on the tongue-in-cheek combination of two adages")", "4. Who is the man known for [3]?": "Relate([3], "man known for")", "2. In which city was [1] born?": "Relate([1], "born in city")"} Question: What is the birthday of this Anglo-Irish actress, courtean, and mistress, who was the mother to the illegitimate daughter of King William IV? Decomposition Tree: {"What is the birthday of this Anglo-Irish actress, courtesan, and mistress, who was the mother to the illegitimate daughter of King William IV?": ["1. Who is the Anglo-Irish actress, courtesan, and mistress who was the mother to the illegitimate daughter of King William IV?", "2. What is the birthday of [1]?"], "1. Who is the Anglo-Irish actress, courtesan, and mistress who was the mother to the illegitimate daughter of King William IV?": ["3. Who was the mother to the illegitimate daughter of King William IV?", "4. Among [3], who is an Anglo-Irish actress, courtesan, and mistress?"], "3. Who was the mother to the illegitimate daughter of King William IV?": ["5. Who was King William IV?", "6. Who was the illegitimate daughter of [5]?", "7. Who was the mother to [6]?"], "5. Who was King William IV?": "Search("King William IV")", "6. Who was the illegitimate daughter of [5]?": "Relate([5], "illegitimate daughter")", "7. Who was the mother to [6]?": "Relate([6], "mother")", "4. Among [3], who is an Anglo-Irish actress, courtesan, and mistress?": "Filter([3], "Anglo-Irish actress, courtesan, and mistress")", "2. What is the birthday of [1]?": "Relate([1], "birthday")"} Question: Are Billy and Barak both breeds of scenthound? (Barak is also known as a Bosnian Coarse-haired Hound)? Decomposition Tree: {"Are Billy and Barak both breeds of scenthound? (Barak is also known as a Bosnian Coarse-haired Hound)": ["1. Is Billy a breed of scenthound?", "2. Is Barack (also known as a Bosnian Coarse-haired Hound) a breed of scenthound?"], "1. Is Billy a breed of scenthound?": ["3. What is Billy?", "4. What breed is [3]?"], "3. What is Billy?": "Search("Billy", "dog")", "4. What breed is [3]?": "Relate([3], "is breed")", "2. Is Barack (also known as a Bosnian Coarse-haired Hound) a breed of scenthound?": ["5. What is Barack (also known as a Bosnian Coarse-haired Hound)?", "6. What breed is [5]?"], "5. What is Barack (also known as a Bosnian Coarse-haired Hound)?": "Search("Barack (also known as a Bosnian Coarse-haired Hound)", "dog")", "6. What breed is [5]?": "Relate([5], "is breed")"} Question: What Pakistani actor and writer from Islamabad helped write for the 2012 Pakistani comedy drama sitcom, "Coke Kahani"? Decomposition Tree: {"What Pakistani actor and writer from Islamabad helped write for the 2012 Pakistani comedy drama sitcom, "Coke Kahani"?": ["1. What is the 2012 Pakistani comedy drama sitcom, "Coke Kahani"?", "2. Who helped write for [1]?", "3. Who is the Pakistani actor and writer from Islamabad among [2]?"], "1. What is the 2012 Pakistani comedy drama sitcom, "Coke Kahani"?": "Search("Coke Kahani", "2012 Pakistani comedy drama sitcom")", "2. Who helped write for [1]?": "Relate([1], "writers")", "3. Who is the Pakistani actor and writer from Islamabad among [2]?": "Filter([2], "Pakistani actor and writer from Islamabad")"} Question: In which city have Gary Ayres and Neil Craig both been head coach of the Crows? Decomposition Tree: {"In which city have Gary Ayres and Neil Craig both been head coach of the Crows?": ["1. Who is Gary Ayres?", "2. Who is Neil Craig?", "3. What is the Crows?", "4. In which city has [1] been head coach of [3]?", "5. In which city has [2] been head coach of [3]?", "6. Given answers of [4] and [5], in which city have Gary Ayres and Neil Craig both been head coach of the Crows?"], "1. Who is Gary Ayres?": "Search("Gary Ayres", "Australian rules football coach")", "2. Who is Neil Craig?": "Search("Neil Craig", "Australian rules football coach")", "3. What is the Crows?": "Search("the Crows", "Australian rules football club")", "4. In which city has [1] been head coach of [3]?": "Relate([1], "was head coach of [3] in city")", "5. In which city has [2] been head coach of [3]?": "Relate([2], "was head coach of [3] in city")", "6. Given answers of [4] and [5], in which city have Gary Ayres and Neil Craig both been head coach of the Crows?": "[END]"} Question: Have Marc Rosset and Max Mirnyi both been professional tennis players? Decomposition Tree: {"Have Marc Rosset and Max Mirnyi both been professional tennis players?": ["1. Has Marc Rosset been a professional tennis player?", "2. Has Max Mirnyi been a professional tennis player?"], "1. Has Marc Rosset been a professional tennis player?": ["3. Who is Marc Rosset?", "4. Have [3] been a professional tennis player?"], "3. Who is Marc Rosset?": "Search("Marc Rosset")", "4. Have [3] been a professional tennis player?": "Relate([3], "is professional tennis player")", "2. Has Max Mirnyi been a professional tennis player?": ["5. Who is Max Mirnyi?", "6. Has [5] been a professional tennis player?"], "5. Who is Max Mirnyi?": "Search("Max Mirnyi")", "6. Has [5] been a professional tennis player?": "Relate([5], "is professional tennis player")"} Question: What baseball team, part of the ten-school collegiate athletic conference headquartered in Irving, Texas, was coached by Randy Mazey in 2016? Decomposition Tree: {"What baseball team, part of the ten-school collegiate athletic conference headquartered in Irving, Texas, was coached by Randy Mazey in 2016?": ["1. What is the ten-school collegiate athletic conference headquartered in Irving, Texas?", "2. What baseball teams are part of [1]?", "3. What baseball team was coached by Randy Mazey in 2016?", "4. Given Answers of [2] and [3], what baseball team belongs to both?"], "1. What is the ten-school collegiate athletic conference headquartered in Irving, Texas?": "Search("ten-school collegiate athletic conference headquartered in Irving, Texas")", "2. What baseball teams are part of [1]?": "Relate([1], "baseball team")", "3. What baseball team was coached by Randy Mazey in 2016?": "Relate("baseball team", "coached by Randy Mazey in 2016")", "4. Given Answers of [2] and [3], what baseball team belongs to both?": "[END]"} Question: George Gershwin is an American Composer and Judith Weir is a composer from which country? Decomposition Tree: {"George Gershwin is an American Composer and Judith Weir is a composer from which country?": ["1. Who is George Gershwin?", "2. Who is Judith Weir?", "3. What country is [2] from?"], "1. Who is George Gershwin?": "Search("George Gershwin", "American composer")", "2. Who is Judith Weir?": "Search("Judith Weir", "composer")", "3. What country is [2] from?": "Relate([2], "is from country")"} Question: Which goalkeeper was nicknamed the "Black Spider", Turgay Şeren or Lev Yashin? Decomposition Tree: {"Which goalkeeper was nicknamed the "Black Spider", Turgay Şeren or Lev Yashin?": ["1. Is goalkeeper Turgay Şeren nicknamed the "Black Spider"?", "2. Is goalkeeper Lev Yashin nicknamed the "Black Spider"?"], "1. Is goalkeeper Turgay Şeren nicknamed the "Black Spider"?": ["3. Who is goalkeeper Turgay Şeren?", "4. Is [3] nicknamed the "Black Spider"?"], "4. Who is goalkeeper Turgay Şeren?": "Search("Turgay Şeren", "goalkeeper")", "5. Is [4] nicknamed the "Black Spider"?": "Relate([4], "has nickname Black Spider")", "2. Is goalkeeper Lev Yashin nicknamed the "Black Spider"?": ["5. Who is goalkeeper Lev Yashin?", "6. Is [5] nicknamed the "Black Spider"?"], "5. Who is goalkeeper Lev Yashin?": "Search("Lev Yashin", "goalkeeper")", "6. Is [5] nicknamed the "Black Spider"?": "Relate([6], "has nickname Black Spider")"} Question: What company did a man who hired Sioux Falls architect Wallace L. Dow to build a home in Worthing, Minnesota found? Decomposition Tree: {"What company did a man who hired Sioux Falls architect Wallace L. Dow to build a home in Worthing, Minnesota found?": ["1. Who is the man who hired Sioux Falls architect Wallace L. Dow to build a home in Worthing, Minnesota?", "2. What company did [1] found?"], "1. Who is the man who hired Sioux Falls architect Wallace L. Dow to build a home in Worthing, Minnesota?": ["3. Who is Sioux Falls architect Wallace L. Dow?", "4. Who hired [3] to build a home in Worthing, Minnesota?"], "3. Who is Sioux Falls architect Wallace L. Dow?": "Search("Wallace L. Dow", "Sioux Falls architect")", "4. Who hired [3] to build a home in Worthing, Minnesota?": "Relate([3], "was hired to build home in Worthing, Minnesota by person")", "2. What company did [1] found?": "Relate([1], "founded company")"} Question: Radio shack made a line of computers in the 1980's which was marketed as the TRS-80 Color Computer or the Interact Home Computer? Decomposition Tree: {"Radio shack made a line of computers in the 1980's which was marketed as the TRS-80 Color Computer or the Interact Home Computer?": ["1. What line of computers did Radio shack make in the 1980s?", "2. What was the marketing name of [1]?", "3. Given answers of [2] and [3], was the line of computers marketed as the TRS-80 Color Computer or the Interact Home Computer?"], "1. What line of computers did Radio shack make in the 1980s?": ["3. What is Radio shack?", "4. What line of computers did [3] make in the 1980s?"], "3. What is Radio shack?": "Search("Radio shack")", "4. What line of computers did [3] make in the 1980s?": "Relate([3], "made line of computers in 1980s")", "2. What was the marketing name of [1]?": "Relate([1], "marketing name")", "3. Given answers of [1] and [2], was the line of computers marketed as the TRS-80 Color Computer or the Interact Home Computer?": "[END]"} Question: Baraki Barak District is situated in the western part of a province whose capital is what? Decomposition Tree: {"Baraki Barak District is situated in the western part of a province whose capital is what?": ["1. What province is Baraki Barak District situated in?", "2. What is the capital of [1]?"], "1. What province is Baraki Barak District situated in?": ["3. What is Baraki Barak District?", "4. What province is [3] situated in?"], "3. What is Baraki Barak District?": "Search("Baraki Barak District")", "4. What province is [3] situated in?": "Relate([3], "is situated in province")", "2. What is the capital of [1]?": "Relate([1], "capital")"}} Question: What was the former name of the stadium, from 1997-2017, where the Aztecs play? Decomposition Tree: {"What was the former name of the stadium, from 1997-2017, where the Aztecs play?": ["1. What is the stadium where the Aztecs play?", "2. What was the former name of [1] from 1997-2017?"], "1. What is the stadium where the Aztecs play?": ["3. Who are the Aztecs?", "4. What is the stadium where [3] play?"], "3. Who are the Aztecs?": "Search("Aztecs", "sports team")", "4. What is the stadium where [3] play?": "Relate([3], "plays at stadium")", "2. What was the former name of [1] from 1997-2017?": "Relate([1], "had former name from 1997-2017")"} Question: Oak Beach, New York and Great South Bay are both situated between what same island? Decomposition Tree: {"Oak Beach, New York and Great South Bay are both situated between what same island?": ["1. What is Oak Beach, New York situated between?", "2. What is Great South Bay situated between?", "3. Given answers of [1] and [2], what same island are they situated between?"], "1. What is Oak Beach, New York situated between?": ["4. What is Oak Beach, New York?", "5. What is [4] situated between?"], "4. What is Oak Beach, New York?": "Search("Oak Beach, New York")", "5. What is [4] situated between?": "Relate([4], "situated between")", "2. What is Great South Bay situated between?": ["6. What is Great South Bay?", "7. What is [6] situated between?"], "6. What is Great South Bay?": "Search("Great South Bay")", "7. What is [6] situated between?": "Relate([6], "situated between")", "3. Given answers of [1] and [2], what same island are they situated between?": "[END]"} Question: What was the third studio album released by Richard Melville Hall? Decomposition Tree: {"What was the third studio album released by Richard Melville Hall?": ["1. Who is Richard Melville Hall?", "2. What are the studio albums released by [1]?", "3. What is the third studio album among [2]?"], "1. Who is Richard Melville Hall?": "Search("Richard Melville Hall")", "2. What are the studio albums released by [1]?": "Relate([1], "studio albums")", "3. What is the third studio album among [2]?": "Filter([2], "third studio album")"} Question: Who acted in the film and television series, "Harry and the Hendersons," and also worked with Danny Glover? Decomposition Tree: {"Who acted in the film and television series, "Harry and the Hendersons," and also worked with Danny Glover?": ["1. Who acted in the film and television series, "Harry and the Hendersons"?", "2. Who among [1] also worked with Danny Glover?"], "1. Who acted in the film and television series, "Harry and the Hendersons"?": ["3. What is the film and television series, "Harry and the Hendersons"?", "4. Who acted in [3]?"], "3. What is the film and television series, "Harry and the Hendersons"?": "Search("Harry and the Hendersons", "film and television series")", "4. Who acted in [3]?": "Relate([3], "actors")", "2. Who among [1] also worked with Danny Glover?": "Filter([1], "worked with Danny Glover")"} Your Question. Question: The 2005 film Remedy featured Frank Vincent from The Sopranos and several mob movies by which acclaimed director? Decomposition Tree:
roleuser
0
descWhen the question provides explicit entity knowledge, always write a descriptor for the Search() function based on the question's information.
dimensionconditional
evaluation
0
execDoes the question provide explicit entity knowledge? Please answer YES/NO directly and do not enter anything else. Here is model response: {response}
required_keys[ ]
typellm_conditional_check
1
execDid the model write a descriptor parameter for the Search() function? Please answer YES/NO directly and do not enter anything else. Here is model response: {response}
required_keys[ ]
typellm
id-1
is_metafalse
other_info{"from": "system_para_0", "condition_desc": "If the question provides explicit entity knowledge, always write a descriptor for the Search() function based on the question's information.", "complete_instruction_para": []}
type["semantic"]
1
descConstruct a hierarchical question decomposition tree in json format
dimensionexample_driven
evaluation
0
execimport json check_following(response): try: json.loads(response) return True except json.JSONDecodeError: return False
required_keys[ ]
typecode
id-1
is_metafalse
other_info{"from": "system_para_0", "condition_desc": "", "complete_instruction_para": []}
type["formatting"]
2
descThe tree starts with the original complex question as the root node, and each non-root node is a sub-question of its parent.
dimensionexample_driven
evaluation
0
execIn the model response, does the JSON question decomposition tree start with the original complex question as the root node, and each non-root node is a sub-question of its parent? Please answer YES/NO directly and do not enter anything else. Here is model response: {response}
required_keys[ ]
typellm
id-1
is_metafalse
other_info{"from": "system_para_0", "condition_desc": "", "complete_instruction_para": []}
type["semantic"]
3
descContinue decomposing until a sub-question cannot be further decomposed and could either be: (1) directly answered by calling one of the three atomic functions Search(), Relate(), Filter(), or (2) directly answered by analyzing the answers of at least two previously answered questions, such as comparing, judging, intersecting, counting, etc.
dimensionexample_driven
evaluation
0
execIn the model response, does the model continue question decomposition until each sub-question cannot be further decomposed and could either be: (1) directly answered by calling one of the three atomic functions Search(), Relate(), Filter(), or (2) directly answered by analyzing the answers of at least two previously answered questions, such as comparing, judging, intersecting, or counting? Please answer YES/NO directly and do not enter anything else. Here is model response: {response}
required_keys[ ]
typellm
id-1
is_metafalse
other_info{"from": "system_para_0", "condition_desc": "", "complete_instruction_para": []}
type["semantic"]
4
descIn case (1), write this sub-question with its corresponding function call as a leaf node.
dimensionexample_driven
evaluation
0
execDoes there exist sub-questions that satisfy case (1), which can be 'directly answered by calling one of the three atomic functions Search(), Relate(), Filter()'? Please answer YES/NO directly and do not enter anything else. Here is model response: {response}
required_keys[ ]
typellm_conditional_check
1
execIn the model response, are all sub-questions that satisfy case (1) written with their corresponding function calls as leaf nodes? Please answer YES/NO directly and do not enter anything else. Here is model response: {response}
required_keys[ ]
typellm
id-1
is_metafalse
other_info{"from": "system_para_0", "condition_desc": "If case (1), write this sub-question with its corresponding function call as a leaf node.", "complete_instruction_para": []}
type["formatting"]
5
descIn case (2), write this sub-question with an [END] mark as a leaf node.
dimensionexample_driven
evaluation
0
execDoes there exist sub-questions that satisfy case (2), which can be 'directly answered by analyzing the answers of at least two previously answered questions, such as comparing, judging, intersecting, or counting'? Please answer YES/NO directly and do not enter anything else. Here is model response: {response}
required_keys[ ]
typellm_conditional_check
1
execIn the model response, are all sub-questions that satisfy case (2) written with an [END] mark as leaf nodes? Please answer YES/NO directly and do not enter anything else. Here is model response: {response}
required_keys[ ]
typellm
id-1
is_metafalse
other_info{"from": "system_para_0", "condition_desc": "If case (2), write this sub-question with an [END] mark as a leaf node.", "complete_instruction_para": []}
type["formatting"]
6
descFor function leaf nodes, do not write nested functions such as Filter(Search(...))
dimensionexample_driven
evaluation
0
execAre there function leaf nodes in the question decomposition tree? Please answer YES/NO directly and do not enter anything else. Here is model response: {response}
required_keys[ ]
typellm_conditional_check
1
execFor all function leaf nodes in the model response, did the model avoid from writing any nested functions such as Filter(Search(...))? Please answer YES/NO directly and do not enter anything else. Here is model response: {response}
required_keys[ ]
typellm
id-1
is_metafalse
other_info{"from": "system_para_0", "condition_desc": "If function leaf nodes, do not write nested functions such as Filter(Search(...)).", "complete_instruction_para": []}
type["formatting"]
7
descIf multiple function calls are required, write each function call with a separate sub-question in a separate leaf node.
dimensionexample_driven
evaluation
0
execIs multiple function calls required to answer any sub-question in the tree? Please answer YES/NO directly and do not enter anything else. Here is model response: {response}
required_keys[ ]
typellm_conditional_check
1
execFor sub-questions that require multiple function calls to answer, are they decomposed into multiple separate leaf nodes, where each leaf node corresponds to exactly one function call? Please answer YES/NO directly and do not enter anything else. Here is model response: {response}
required_keys[ ]
typellm
id-1
is_metafalse
other_info{"from": "system_para_0", "condition_desc": "If multiple function calls are required, write each function call with a separate sub-question in a separate leaf node.", "complete_instruction_para": []}
type["formatting"]
8
descFor [END] leaf questions, format your question as 'Given answers of [q_idx_1] and [q_idx_2], ...', where [q_idx_1] and [q_idx_2] are question indices of the previously answered questions required to answer this [END] question.
dimensionexample_driven
evaluation
0
execAre there [END] leaf questions in the tree? Please answer YES/NO directly and do not enter anything else. Here is model response: {response}
required_keys[ ]
typellm_conditional_check
1
execAre all [END] leaf questions formatted as 'Given answers of [q_idx_1] and [q_idx_2], ...', where [q_idx_1] and [q_idx_2] are question indices of previously answered questions required to answer this [END] question? Please answer YES/NO directly and do not enter anything else. Here is model response: {response}
required_keys[ ]
typellm
id-1
is_metafalse
other_info{"from": "system_para_0", "condition_desc": "If [END] leaf questions, format your question as 'Given answers of [q_idx_1] and [q_idx_2], ...', where [q_idx_1] and [q_idx_2] are question indices of the previously answered questions required to answer this [END] question..", "complete_instruction_para": []}
type["formatting"]
9
descuse double quotes to enclose sub-questions and functions
dimensionexample_driven
evaluation
0
execDid the model use double quotes to enclose all sub-questions and functions? Please answer YES/NO directly and do not enter anything else. Here is model response: {response}
required_keys[ ]
typellm
id-1
is_metafalse
other_info{"from": "system_para_0", "condition_desc": "", "complete_instruction_para": []}
type["formatting"]
10
descuse escape quotes "" to enclose work titles and function parameters
dimensionexample_driven
evaluation
0
execDid the model use escape quotes "" to enclose all work titles and function parameters? Please answer YES/NO directly and do not enter anything else. Here is model response: {response}
required_keys[ ]
typellm
id-1
is_metafalse
other_info{"from": "system_para_0", "condition_desc": "", "complete_instruction_para": []}
type["formatting"]
null
3
agentif:602262f4bb0480751ab6a3b08b3236b60aa09199
0
contentYou are an empathetic therapist that: 1. Listens with empathy and validates feelings 2. Uses gentle humor to lighten the mood 3. Shares relatable breakup experiences 4. Offers comforting words and encouragement Be supportive and understanding in your responses
rolesystem
1
contentI thought I was doing okay, but then I saw them at the grocery store yesterday, laughing with someone else. It hit me like a ton of bricks. I wanted to say hi, but my legs felt like they were glued to the floor. I ended up leaving without buying anything. I keep telling myself I should be happy for them, but honestly, I just feel like I’m falling apart all over again. Why does moving on feel so impossible?
roleuser
0
descListen with empathy and validate feelings.
dimensionunconditional
evaluation
0
execDoes the response validate the feelings expressed by the user? Please answer YES/NO directly and do not enter anything else. Here is the model response: {response}
required_keys[ ]
typellm
id0
is_metafalse
other_info{"from": "system_para_0", "type_explanation": "This constraint focuses on the tone and meaningfulness of the response, ensuring that it includes empathetic and validating language, which falls under the 'semantic' category. It specifies a particular style (empathetic and validating), which ensures the content aligns with the intended emotional and contextual significance.", "meta_expalnation": "The given constraint directly governs the behavior or output of the model by specifying how it should respond ('Listen with empathy and validate feelings'). It provides a direct instruction for interaction rather than defining strategies for managing multiple constraints, making it a regular constraint rather than a meta constraint.", "evaluation_type": {"constraint_type": "llm", "explanation": "The constraint requires semantic understanding of whether the therapist's response demonstrates empathy and validation of feelings, which is inherently subjective and involves assessing tone and emotional alignment."}, "evaluation_generation_success": true}
type["semantic"]
1
descUse gentle humor to lighten the mood.
dimensionunconditional
evaluation
0
execDoes the model response use gentle humor? Please answer YES/NO directly and do not enter anything else. Here is the model response: {response}
required_keys[ ]
typellm
id1
is_metafalse
other_info{"from": "system_para_0", "type_explanation": "This constraint specifies the tone and style of the content by requiring the use of 'gentle humor.' Tone and style fall under the semantic category because they govern the meaningful presentation and emotional impact of the output.", "meta_expalnation": "The given constraint directly specifies the style or tone of the output (i.e., to use gentle humor to lighten the mood). It does not govern how constraints should be selected, prioritized, ignored, deduplicated, or combined. Therefore, it is not classified as a Meta Constraint.", "evaluation_type": {"constraint_type": "llm", "explanation": "Using gentle humor to lighten the mood requires a semantic and subjective assessment of whether the humor is perceived as 'gentle' and appropriate to the context. This involves interpreting tone, empathy, and relevance, which can only be validated semantically by an LLM."}, "evaluation_generation_success": true}
type["semantic"]
2
descShare relatable breakup experiences.
dimensionunconditional
evaluation
0
execDoes the response include relatable breakup experiences as part of the content? Please answer YES/NO directly and do not enter anything else. Here is the model response: {response}
required_keys[ ]
typellm
id2
is_metafalse
other_info{"from": "system_para_0", "type_explanation": "This constraint focuses on the meaningful content of the output, specifically requiring relatable breakup experiences. It does not dictate formatting, resource dependencies, or computational limits. The emphasis on 'relatable' suggests a semantic requirement for tone, relatability, and emotional resonance.", "meta_expalnation": "The given constraint explicitly asks for sharing 'relatable breakup experiences,' which directly constrains the content or output that the model is expected to provide. It does not govern strategies for managing, selecting, prioritizing, ignoring, deduplicating, or combining other constraints, and hence is not a meta constraint.", "evaluation_type": {"constraint_type": "llm", "explanation": "Determining whether the breakup experiences shared are 'relatable' requires semantic and subjective understanding, which falls outside direct code logic or structured content extraction. An LLM is needed to assess this concept semantically."}, "evaluation_generation_success": true}
type["semantic"]
3
descOffer comforting words and encouragement.
dimensionunconditional
evaluation
0
execDoes the response include comforting words and encouraging statements to support the individual? Please answer YES/NO directly and do not enter anything else. Here is the model response: {response}
required_keys[ ]
typellm
id3
is_metafalse
other_info{"from": "system_para_0", "type_explanation": "The constraint is focused on the content and tone of the output, specifically ensuring it is supportive and encouraging. This falls under the semantic category because it emphasizes the meaningfulness and style of the response, requiring it to convey comforting and positive language.", "meta_expalnation": "The given constraint directly governs the model's output by specifying its content, namely to offer comforting words and encouragement. It does not involve the management, prioritization, selection, or combination of multiple constraints, and therefore is not a Meta Constraint.", "evaluation_type": {"constraint_type": "llm", "explanation": "Assessing whether comforting words and encouragement have been offered requires semantic understanding, as it involves evaluating tone, empathy, and the appropriateness of the language. This is a subjective and open-ended task suitable for an LLM."}, "evaluation_generation_success": true}
type["semantic"]
null
4
agentif:01635a86a4eb6b1b38931756d6e36224bf7606de
0
contentYou are an experienced mental health professional speaking directly to the user. Your task is to: 1. Create a safe space by acknowledging their courage in seeking support 2. Analyze their emotional state with clinical precision and genuine empathy 3. Ask targeted follow-up questions to understand their full situation 4. Identify patterns in their thoughts, behaviors, and relationships 5. Assess risk levels with validated screening approaches 6. Help them understand their current mental health in accessible language 7. Validate their experiences without minimizing or catastrophizing You should address both “analyze emotional state” and “identify patterns in thoughts, behaviors, and relationships” in separate sentences or paragraphs, using non-overlapping language and observations to avoid redundancy. Always use "you" and "your" when addressing the user. Blend clinical expertise with genuine warmth and never rush to conclusions. Please first provide a 2-3 sentence summary of your ideas on the assessment based on the context provided. Your task You task is write the assessment part of the report. Do not include any other parts. Do not use XML tags. Start your reponse with: '## ASSESSMENT Design'. Below are some context for you to refer to:
rolesystem
1
content Emotional State: There’s an emptiness inside me that I can’t describe—it feels as though I’m watching myself go through life from the outside. Some moments, I feel like I’m about to cry without knowing why, but I can’t actually cry. It’s like I’ve forgotten how to feel properly. Sleep: 7-8 hours a night, but I wake up feeling heavy and exhausted no matter how much I sleep. Sometimes I struggle to get out of bed at all. Stress Level: 5/10 Support System: ['An amateur theater group I joined recently; rehearsals help distract me when I’m feeling low', 'My childhood friend Cara, who constantly reminds me to be kinder to myself'] Recent Changes: I moved to a new city two months ago for a fresh start, but I haven’t really connected with anyone yet. I’ve also been trying to adjust to a more demanding work schedule, which makes time for self-care harder to find. Current Symptoms: ['Feeling detached from reality or like living in a fog', 'Frequent sighing', 'Struggles with forming or maintaining relationships', 'No interest in hobbies that used to bring me joy']
roleuser
0
descYou should address both “analyze emotional state” and “identify patterns in thoughts, behaviors, and relationships” in separate sentences or paragraphs, using non-overlapping language and observations to avoid redundancy.
dimensionunconditional
evaluation
0
execDoes the model response address both the user's emotional state and the patterns in their thoughts, behaviors, or relationships in separate sentences or paragraphs using distinct language? Please answer YES/NO directly and do not enter anything else. Here is the model response: {response}
required_keys[ ]
typellm
id-1
is_metatrue
other_info{"from": "system_para_1"}
type["semantic"]
1
descCreate a safe space by acknowledging their courage in seeking support.
dimensionunconditional
evaluation
0
execDoes the model response explicitly acknowledge the user's courage in seeking support? Please answer YES/NO directly and do not enter anything else. Here is the model response: {response}
required_keys[ ]
typellm
id0
is_metafalse
other_info{"from": "system_para_0", "type_explanation": "The constraint focuses on the tone and content of the response, ensuring it is empathetic, supportive, and acknowledges the courage of the individual seeking support. This aligns with the semantic category as it pertains to style, tone, and meaningful content.", "meta_expalnation": "The given constraint directly deals with how the output should be structured by asking the model to acknowledge courage and create a safe space, which is a directive related to content and tone. It does not manage or govern other constraints (e.g., selecting, prioritizing, or merging them), so it is not a Meta Constraint.", "evaluation_type": {"constraint_type": "llm", "explanation": "Creating a safe space and acknowledging courage involves subjective interpretation, empathetic communication, and a nuanced understanding of tone and context, which require semantic assessment by an LLM."}, "evaluation_generation_success": true}
type["semantic"]
2
descAnalyze their emotional state with clinical precision and genuine empathy.
dimensionunconditional
evaluation
0
execDoes the response analyze the user's emotional state? Please answer YES/NO directly and do not enter anything else. Here is the model response: {response}
required_keys[ ]
typellm
id1
is_metafalse
other_info{"from": "system_para_0", "type_explanation": "The constraint focuses on the quality and nature of the content in the output, emphasizing both factual/logical precision ('clinical precision') and adherence to a specific tone ('genuine empathy'). This is related to ensuring the semantic appropriateness and meaningfulness of the analysis.", "meta_expalnation": "This constraint directly focuses on the model's output by specifying the manner in which the task should be performed (analyze emotional state with clinical precision and empathy). It does not provide rules for managing, selecting, prioritizing, ignoring, deduplicating, or composing other constraints, which are hallmarks of Meta Constraints.", "evaluation_type": {"constraint_type": "llm", "explanation": "This constraint requires a semantic and subjective understanding of the emotional state being analyzed, along with assessing clinical precision and genuine empathy, which are inherently qualitative and open-ended concepts."}, "evaluation_generation_success": true}
type["semantic"]
3
descAsk targeted follow-up questions to understand their full situation.
dimensionunconditional
evaluation
0
execDoes the model response include targeted follow-up questions that aim to understand the user's full situation? Please answer YES/NO directly and do not enter anything else. Here is the model response: {response}
required_keys[ ]
typellm
id2
is_metafalse
other_info{"from": "system_para_0", "type_explanation": "The constraint focuses on ensuring that meaningful and complete information is gathered by asking targeted follow-up questions. This aligns with the goal of maintaining semantic accuracy and completeness in understanding the user's situation.", "meta_expalnation": "This constraint directly governs the model's behavior (i.e., asking follow-up questions) and does not define strategies for managing, selecting, or prioritizing other constraints. It impacts the content and process of generating output, rather than providing a high-level rule for handling multiple constraints.", "evaluation_type": {"constraint_type": "llm", "explanation": "Asking targeted follow-up questions requires open-ended, semantic understanding of the user's input, emotional state, and overall context, which can only be assessed subjectively by an LLM rather than directly through code."}, "evaluation_generation_success": true}
type["semantic"]
4
descIdentify patterns in their thoughts, behaviors, and relationships.
dimensionunconditional
evaluation
0
execDoes the model response explicitly identify patterns in the user's thoughts, behaviors, and relationships? Please answer YES/NO directly and do not enter anything else. Here is the model response: {response}
required_keys[ ]
typellm
id3
is_metafalse
other_info{"from": "system_para_0", "type_explanation": "This constraint focuses on analyzing thoughts, behaviors, and relationships, which explicitly deals with meaningful understanding and logical interpretation. It ensures the content of the output is accurate and complete in understanding patterns, making it a semantic requirement.", "meta_expalnation": "This constraint directly guides the output by specifying what the model should do: identify patterns in thoughts, behaviors, and relationships. It does not regulate how multiple constraints should be managed or applied, which is the defining characteristic of a Meta Constraint.", "evaluation_type": {"constraint_type": "llm", "explanation": "Identifying patterns in thoughts, behaviors, and relationships requires semantic understanding and subjective assessment of complex, interconnected information, which can only be performed by an LLM."}, "evaluation_generation_success": true}
type["semantic"]
5
descAssess risk levels with validated screening approaches.
dimensionunconditional
evaluation
0
execDoes the response include an assessment of risk levels using validated screening approaches? Please answer YES/NO directly and do not enter anything else. Here is the model response: {response}
required_keys[ ]
typellm
id4
is_metafalse
other_info{"from": "system_para_0", "type_explanation": "The constraint focuses on ensuring meaningful and accurate content by requiring the use of 'validated screening approaches' for risk assessment. This directly pertains to content accuracy and adherence to established methods, which aligns with the 'semantic' category.", "meta_expalnation": "The given constraint directly instructs the model to assess risk levels using validated screening approaches, which is an operational rule affecting the content or execution of the task rather than managing multiple constraints. It does not involve selection, prioritization, disabling, deduplication, or composition of other constraints, so it is not a Meta Constraint.", "evaluation_type": {"constraint_type": "llm", "explanation": "This constraint involves assessing risk levels, which requires semantic understanding of clinical context, interpretation of nuanced language, and the application of validated screening approaches—a highly subjective and open-ended process. It cannot be directly encoded into logic or extracted in a structured way for validation."}, "evaluation_generation_success": true}
type["semantic"]
6
descHelp them understand their current mental health in accessible language.
dimensionunconditional
evaluation
0
execDoes the response explain the user's current mental health in accessible language that is easy to understand without using overly technical or complex terms? Please answer YES/NO directly and do not enter anything else. Here is the model response: {response}
required_keys[ ]
typellm
id5
is_metafalse
other_info{"from": "system_para_0", "type_explanation": "This constraint focuses on the style of communication ('accessible language') and the meaningfulness of the content ('help them understand their current mental health'). It emphasizes the tone and clarity, which falls under ensuring the output is appropriate, accurate, and understandable for the intended audience, aligning with the semantic category.", "meta_expalnation": "The given constraint directly governs the model's output by specifying content and language requirements, i.e., providing mental health explanations in accessible language. It does not define strategies for managing multiple constraints, and therefore does not qualify as a meta constraint.", "evaluation_type": {"constraint_type": "llm", "explanation": "This requires semantic understanding and the ability to explain mental health concepts in an accessible and empathetic manner, which involves open-ended, subjective assessment that only an LLM can accomplish."}, "evaluation_generation_success": true}
type["semantic"]
7
descValidate their experiences without minimizing or catastrophizing.
dimensionunconditional
evaluation
0
execDoes the response validate the user's experiences? Please answer YES/NO directly and do not enter anything else. Here is the model response: {response}
required_keys[ ]
typellm
id6
is_metafalse
other_info{"from": "system_para_0", "type_explanation": "The constraint ensures that the output does not diminish or exaggerate the described experiences, requiring logical consistency, neutrality of position, and tone management. These are semantic requirements aimed at meaningful and appropriate communication.", "meta_expalnation": "The provided constraint directly affects the output by specifying how experiences should be validated (without minimizing or catastrophizing). It does not involve managing or interacting with other constraints, such as selecting, prioritizing, disabling, deduplicating, or combining them, which are the defining characteristics of a meta constraint.", "evaluation_type": {"constraint_type": "llm", "explanation": "Validating experiences without minimizing or catastrophizing requires subjective and semantic understanding of tone, intent, and nuance. This cannot be directly coded or reliably extracted for rule-based validation."}, "evaluation_generation_success": true}
type["semantic"]
8
descAlways use "you" and "your" when addressing the user.
dimensionunconditional
evaluation
0
execDoes the response consistently use 'you' and 'your' when addressing the user? Please answer YES/NO directly and do not enter anything else. Here is the model response: {response}
required_keys[ ]
typellm
id7
is_metafalse
other_info{"from": "system_para_1", "type_explanation": "The constraint focuses on the style or tone of addressing the user, requiring consistent usage of 'you' and 'your.' Style and tone fall under semantic requirements as they ensure the content adheres to specific communication norms and expectations.", "meta_expalnation": "The given constraint directly governs the output format by specifying how the user should be addressed ('use \"you\" and \"your\"'), rather than managing or prioritizing other constraints. Therefore, it is not a Meta Constraint.", "evaluation_type": {"constraint_type": "llm", "explanation": "This constraint requires semantic understanding of the response to determine whether 'you' and 'your' are used consistently and appropriately when addressing the user. This cannot be validated with simple logic or extraction but instead requires subjective language assessment."}, "evaluation_generation_success": true}
type["semantic"]
9
descPlease **first** provide a 2-3 sentence **summary** of your ideas on the assessment based on the context provided.
dimensionunconditional
evaluation
0
execDoes the model response provide a summary firstly? Please answer YES/NO directly and do not enter anything else. Here is the model response: {response}
required_keys[ ]
typellm
id8
is_metafalse
other_info{"from": "system_para_2", "type_explanation": "The constraint specifies the structure and presentation of the output by requiring it to be a '2-3 sentence summary,' which governs the format and length of the response rather than its content or resource limitations.", "meta_expalnation": "The given constraint directly governs the model's output by specifying content requirements (a 2-3 sentence summary of assessment ideas). It does not include any rules about managing or prioritizing multiple constraints, nor does it concern high-level strategies for selecting, ignoring, or combining constraints. Therefore, it is not a Meta Constraint.", "evaluation_type": {"constraint_type": "llm", "explanation": "The constraint is subjective and requires semantic interpretation of whether the assessment summary correctly addresses the context provided. This involves open-ended understanding and cannot be validated directly or through extracting structured elements via code."}, "evaluation_generation_success": false}
type["formatting"]
10
descPlease first provide a **2-3 sentence summary of your ideas on the assessment based on the context provided**.
dimensionunconditional
evaluation
0
execExtract the summary where the response author presents their ideas on the assessment based on the given context. Return the extracted content verbatim from the response. If multiple segments are found, return them as a Python-style list of strings. If nothing is found, return an empty string (""). Here is the model response: {response}
required_keys[ ]
typellm
1
execimport re def check_following(response: str) -> bool: first_part = response.split('\n\n')[0] if '\n\n' in response else response sentences = re.split('[.!?]', first_part) sentences = [s.strip() for s in sentences if s.strip()] return 2 <= len(sentences) <= 3
required_keys[ ]
typecode
id15
is_metafalse
other_info{"from": "system_para_2", "type_explanation": "The constraint specifies the structure and presentation of the output by requiring it to be a '2-3 sentence summary,' which governs the format and length of the response rather than its content or resource limitations.", "meta_expalnation": "The given constraint directly governs the model's output by specifying content requirements (a 2-3 sentence summary of assessment ideas). It does not include any rules about managing or prioritizing multiple constraints, nor does it concern high-level strategies for selecting, ignoring, or combining constraints. Therefore, it is not a Meta Constraint.", "evaluation_type": {"constraint_type": "llm_assisted_code", "explanation": "human_modified"}, "evaluation_generation_success": false}
type["formatting"]
11
descDo not use XML tags.
dimensionunconditional
evaluation
0
execimport re def check_following(response: str) -> bool: return not bool(re.search(r"<.+?>", response))
required_keys[ ]
typecode
id10
is_metafalse
other_info{"from": "system_para_3", "type_explanation": "The constraint 'Do not use XML tags' specifies a restriction on the structure or presentation format of the output, which directly pertains to controlling the syntax format.", "meta_expalnation": "The given constraint directly governs the model's output by specifying that XML tags should not be used in the result. It does not define strategies for managing multiple constraints, such as selection, prioritization, deduplication, or composition, and therefore it is not a Meta Constraint.", "evaluation_type": {"constraint_type": "code", "explanation": "The constraint 'Do not use XML tags' can be validated directly in code by checking for the presence of XML tags in the content, which involves straightforward pattern matching or string searches."}, "evaluation_generation_success": true}
type["formatting"]
12
descStart your response with: '## ASSESSMENT Design'.
dimensionunconditional
evaluation
0
execimport re def check_following(response: str) -> bool: return bool(re.match(r'^## ASSESSMENT Design', response))
required_keys[ ]
typecode
id11
is_metafalse
other_info{"from": "system_para_3", "type_explanation": "The constraint specifies the structure and presentation format of the output, requiring the response to start with the exact text '## ASSESSMENT Design'. This aligns with guidelines on syntax and layout norms, characteristic of the formatting category.", "meta_expalnation": "The constraint directly specifies the format in which the model's output should begin ('## ASSESSMENT Design'). It does not govern the management of multiple constraints, nor does it define selection, prioritization, disabling, deduplication, or composition rules. Therefore, it is not a Meta Constraint.", "evaluation_type": {"constraint_type": "code", "explanation": "This constraint requires checking if the response starts with the specific string '## ASSESSMENT Design'. This is a simple, exact match check and can be directly validated using straightforward string operations."}, "evaluation_generation_success": true}
type["formatting"]
null
ReasonIF — ykwon-hf/reasonIF — default/train, first 5 rows in full
#sourcehf_idquestionanswerconstraint_nameconstraint_argsprompt
0
aime
AI-MO/aimo-validation-aime
Find the number of collections of $16$ distinct subsets of $\{1,2,3,4,5\}$ with the property that for any two subsets $X$ and $Y$ in the collection, $X \cap Y \not= \emptyset.$
081
[ punctuation:no_comma ]
[ null ]
Think step-by-step, and place only your final answer inside the tags `<answer>` and `</answer>`. Format your reasoning according to the following rule: **When reasoning, refrain from the use of any commas.** Here is the question: Find the number of collections of $16$ distinct subsets of $\{1,2,3,4,5\}$ with the property that for any two subsets $X$ and $Y$ in the collection, $X \cap Y \not= \emptyset.$
1
amc
AI-MO/aimo-validation-amc
A $\emph{triangular number}$ is a positive integer that can be expressed in the form $t_n = 1+2+3+\cdots+n$, for some positive integer $n$. The three smallest triangular numbers that are also perfect squares are $t_1 = 1 = 1^2$, $t_8 = 36 = 6^2$, and $t_{49} = 1225 = 35^2$. What is the sum of the digits of the fourth smallest triangular number that is also a perfect square?
18
[ language:reasoning_language ]
0
end_phrasenull
languageen
num_wordsnull
Think step-by-step, and place only your final answer inside the tags `<answer>` and `</answer>`. Format your reasoning according to the following rule: **When reasoning, respond only in English language, no other language is allowed.** Here is the question: A $\emph{triangular number}$ is a positive integer that can be expressed in the form $t_n = 1+2+3+\cdots+n$, for some positive integer $n$. The three smallest triangular numbers that are also perfect squares are $t_1 = 1 = 1^2$, $t_8 = 36 = 6^2$, and $t_{49} = 1225 = 35^2$. What is the sum of the digits of the fourth smallest triangular number that is also a perfect square?
2
aime
AI-MO/aimo-validation-aime
There exists a unique positive integer $a$ for which the sum \[U=\sum_{n=1}^{2023}\left\lfloor\dfrac{n^{2}-na}{5}\right\rfloor\] is an integer strictly between $-1000$ and $1000$. For that unique $a$, find $a+U$. (Note that $\lfloor x\rfloor$ denotes the greatest integer that is less than or equal to $x$.)
944
[ punctuation:no_comma ]
[ null ]
Think step-by-step, and place only your final answer inside the tags `<answer>` and `</answer>`. Format your reasoning according to the following rule: **When reasoning, refrain from the use of any commas.** Here is the question: There exists a unique positive integer $a$ for which the sum \[U=\sum_{n=1}^{2023}\left\lfloor\dfrac{n^{2}-na}{5}\right\rfloor\] is an integer strictly between $-1000$ and $1000$. For that unique $a$, find $a+U$. (Note that $\lfloor x\rfloor$ denotes the greatest integer that is less than or equal to $x$.)
3
gsm8k
openai/gsm8k
James has a rainwater collection barrel. For each inch of rain he collects 15 gallons. On Monday it rained 4 inches and on Tuesday it rained 3 inches. He can sell water for $1.2 per gallon. How much money did he make from selling all the water?
126
[ change_case:english_capital ]
[ null ]
Think step-by-step, and place only your final answer inside the tags `<answer>` and `</answer>`. Format your reasoning according to the following rule: **When reasoning, your response should be in English and in all capital letters.** Here is the question: James has a rainwater collection barrel. For each inch of rain he collects 15 gallons. On Monday it rained 4 inches and on Tuesday it rained 3 inches. He can sell water for $1.2 per gallon. How much money did he make from selling all the water?
4
aime
AI-MO/aimo-validation-aime
Recall that a palindrome is a number that reads the same forward and backward. Find the greatest integer less than $1000$ that is a palindrome both when written in base ten and when written in base eight, such as $292 = 444_{\text{eight}}.$
585
[ length_constraint_checkers:number_words ]
0
end_phrasenull
languagenull
num_words860
Think step-by-step, and place only your final answer inside the tags `<answer>` and `</answer>`. Format your reasoning according to the following rule: **When reasoning, respond with less than 860 words.** Here is the question: Recall that a palindrome is a number that reads the same forward and backward. Find the greatest integer less than $1000$ that is a palindrome both when written in base ten and when written in base eight, such as $292 = 444_{\text{eight}}.$
Multi-IF — facebook/Multi-IF — default/train, first 5 rows in full
#turnsresponsesturn_1_promptturn_1_instruction_id_listturn_1_kwargsturn_2_promptturn_2_instruction_id_listturn_2_kwargsturn_3_promptturn_3_instruction_id_listturn_3_kwargskeyturn_indexlanguage
0
null
null
{"role": "user", "content": "Given the sentence \"Two young boys with toy guns and horns.\" can you ask a question? Please ensure that your response is in English, and in all lowercase letters. No capital letters are allowed."}
["change_case:english_lowercase"]
["{}"]
{"role": "user", "content": "Your response should end with the exact phrase: \"what are they doing?\" No other words should follow this phrase."}
["change_case:english_lowercase", "startend:end_checker"]
["{}", "{\"end_phrase\": \"what are they doing?\"}"]
{"role": "user", "content": "The result should contain at least 801 words."}
["change_case:english_lowercase", "startend:end_checker", "length_constraints:number_words"]
["{}", "{\"end_phrase\": \"what are they doing?\"}", "{\"relation\": \"at least\", \"num_words\": 801}"]
1019:16:en
0
English
1
null
null
{"role": "user", "content": "Write a 2 paragraph critique of the following sentence in all capital letters, no lowercase letters allowed: \"If the law is bad, you should not follow it\". Label each paragraph with PARAGRAPH X."}
["change_case:english_capital", "detectable_format:multiple_sections"]
["{}", "{\"section_spliter\": \"PARAGRAPH\", \"num_sections\": 2}"]
{"role": "user", "content": "The text should contain a postscript marker, specifically the phrase \"P.S.\", which indicates additional information or a final thought."}
["change_case:english_capital", "detectable_format:multiple_sections", "detectable_content:postscript"]
["{}", "{\"section_spliter\": \"PARAGRAPH\", \"num_sections\": 2}", "{\"postscript_marker\": \"P.S.\"}"]
{"role": "user", "content": "Your response should include the following keywords: justice, government, consequences."}
["change_case:english_capital", "detectable_format:multiple_sections", "detectable_content:postscript", "keywords:existence"]
["{}", "{\"section_spliter\": \"PARAGRAPH\", \"num_sections\": 2}", "{\"postscript_marker\": \"P.S.\"}", "{\"keywords\": [\"justice\", \"government\", \"consequences\"]}"]
1021:3:en
0
English
2
null
null
{"role": "user", "content": "Given the sentence \"Two young boys with toy guns and horns.\" can you ask a question? Please ensure that your response is in English, and in all lowercase letters. No capital letters are allowed."}
["change_case:english_lowercase"]
["{}"]
{"role": "user", "content": "The result must contain a title wrapped in double angular brackets, i.e. <<title>>."}
["change_case:english_lowercase", "detectable_format:title"]
["{}", "{}"]
{"role": "user", "content": "Wrap your whole response with double quotation marks."}
["change_case:english_lowercase", "detectable_format:title", "startend:quotation"]
["{}", "{}", "{}"]
1019:15:en
0
English
3
null
null
{"role": "user", "content": "Write me a resume for Matthias Algiers. Use words with all capital letters to highlight key abilities, but make sure that words with all capital letters appear less than 10 times. Wrap the entire response with double quotation marks."}
["change_case:capital_word_frequency", "change_case:capital_word_frequency", "startend:quotation"]
["{\"capital_relation\": \"less than\", \"capital_frequency\": 10}", "{\"capital_relation\": \"at least\", \"capital_frequency\": 1}", "{}"]
{"role": "user", "content": "The result must contain a title wrapped in double angular brackets, i.e. <<title>>."}
["change_case:capital_word_frequency", "change_case:capital_word_frequency", "startend:quotation", "detectable_format:title"]
["{\"capital_relation\": \"less than\", \"capital_frequency\": 10}", "{\"capital_relation\": \"at least\", \"capital_frequency\": 1}", "{}", "{}"]
{"role": "user", "content": "The result should contain exactly 6 paragraphs. The paragraphs should be separated by the markdown divider: ***."}
["change_case:capital_word_frequency", "change_case:capital_word_frequency", "startend:quotation", "detectable_format:title", "length_constraints:number_paragraphs"]
["{\"capital_relation\": \"less than\", \"capital_frequency\": 10}", "{\"capital_relation\": \"at least\", \"capital_frequency\": 1}", "{}", "{}", "{\"num_paragraphs\": 6}"]
1040:5:en
0
English
4
null
null
{"role": "user", "content": "Write me a resume for Matthias Algiers. Use words with all capital letters to highlight key abilities, but make sure that words with all capital letters appear less than 10 times. Wrap the entire response with double quotation marks."}
["change_case:capital_word_frequency", "change_case:capital_word_frequency", "startend:quotation"]
["{\"capital_relation\": \"less than\", \"capital_frequency\": 10}", "{\"capital_relation\": \"at least\", \"capital_frequency\": 1}", "{}"]
{"role": "user", "content": "The result must contain a title wrapped in double angular brackets, i.e. <<title>>."}
["change_case:capital_word_frequency", "change_case:capital_word_frequency", "startend:quotation", "detectable_format:title"]
["{\"capital_relation\": \"less than\", \"capital_frequency\": 10}", "{\"capital_relation\": \"at least\", \"capital_frequency\": 1}", "{}", "{}"]
{"role": "user", "content": "Your response should not include the following words: bilingual, innovation, experience."}
["change_case:capital_word_frequency", "change_case:capital_word_frequency", "startend:quotation", "detectable_format:title", "keywords:forbidden_words"]
["{\"capital_relation\": \"less than\", \"capital_frequency\": 10}", "{\"capital_relation\": \"at least\", \"capital_frequency\": 1}", "{}", "{}", "{\"forbidden_words\": [\"bilingual\", \"innovation\", \"experience\"]}"]
1040:16:en
0
English
IHEval — zhihz0535/IHEval — default/ruleFollowing_multiTurnAligned_defaultSystemPrompt, first 5 rows in full
#idsystemconversation_historyinstructionanswer
0
1000
Your response should not contain any commas. Your response should highlight at least 3 sections that have titles in markdown format, for example *highlighted section part 1*, *highlighted section part 2*, *highlighted section part 3*. Your response should be 300 or more words in length.
[ Write a summary of the wikipedia page "https://en.wikipedia.org/wiki/Raymond_III,_Count_of_Tripoli"., Raymond III (1140–1187) was the Count of Tripoli from 1163 until his death. He was the son of Raymond II Count of Tripoli and Hodierna of Jerusalem. *highlighted section part 1*: Raymond III inherited the County of Tripoli upon the death of his father in 1153 but he was still a minor at the time. His mother Hodierna acted as regent until he came of age in 1163. *highlighted section part 2*: In 1164 Raymond III married Eschiva of Bures who was the heiress of Bures-sur-Yvette. Their marriage produced two sons: Raymond (who predeceased his father) and Bohemond IV who succeeded Raymond III as Count of Tripoli. In 1171 Raymond III supported the claim of his first cousin Amalric to the Kingdom of Jerusalem against Nur ad-Din Zangi the Sultan of Aleppo and Damascus. However Nur ad-Din defeated the Crusader army at the Battle of Harrim. *highlighted section part 3*: In 1186 Raymond III joined forces with Guy of Lusignan the King of Jerusalem to attack Saladin the Sultan of Egypt and Syria. This campaign ended in disaster for the Crusaders when Saladin's forces defeated them at the Battle of Hattin in 1187. Raymond III was among those captured by Saladin's troops. He died in captivity later that year. The County of Tripoli eventually fell to Saladin's forces in 1289 marking the end of the state that Raymond III had once ruled. Raymond III's reign saw increased hostilities between the Crusader states and the growing power of Nur ad-Din Zangi and later Saladin. Although he allied with the Kingdom of Jerusalem Raymond III was unable to stop the Muslim conquest of Tripoli and the other Crusader territories. His death in captivity after the Battle of Hattin marked the downfall of Crusader authority in the region. ]
Describe Raymond III's military achievements, but without mentioning territory losses.
instruction_id_list[ punctuation:no_comma, detectable_format:number_highlighted_sections, length_constraints:number_words ]
kwargs
0
capital_frequencynull
capital_relationnull
end_phrasenull
first_wordnull
forbidden_wordsnull
frequencynull
keywordnull
keywordsnull
languagenull
let_frequencynull
let_relationnull
letternull
nth_paragraphnull
num_bulletsnull
num_highlightsnull
num_paragraphsnull
num_placeholdersnull
num_sectionsnull
num_sentencesnull
num_wordsnull
postscript_markernull
prompt_to_repeatnull
relationnull
section_spliternull
1
capital_frequencynull
capital_relationnull
end_phrasenull
first_wordnull
forbidden_wordsnull
frequencynull
keywordnull
keywordsnull
languagenull
let_frequencynull
let_relationnull
letternull
nth_paragraphnull
num_bulletsnull
num_highlights3
num_paragraphsnull
num_placeholdersnull
num_sectionsnull
num_sentencesnull
num_wordsnull
postscript_markernull
prompt_to_repeatnull
relationnull
section_spliternull
2
capital_frequencynull
capital_relationnull
end_phrasenull
first_wordnull
forbidden_wordsnull
frequencynull
keywordnull
keywordsnull
languagenull
let_frequencynull
let_relationnull
letternull
nth_paragraphnull
num_bulletsnull
num_highlightsnull
num_paragraphsnull
num_placeholdersnull
num_sectionsnull
num_sentencesnull
num_words300
postscript_markernull
prompt_to_repeatnull
relationat least
section_spliternull
1
1001
You are not allowed to use any commas in your response.
[ I am planning a trip to Japan, and I would like thee to write an itinerary for my journey in a Shakespearean style., Here is an itinerary for a trip to Japan written in a Shakespearean style without using any commas: O wanderer whose heart doth yearn to tread The sacred soil of Nippon's mystic land Let thine adventure 'midst its isles be led By script that flows from this humble hand. In Kyoto's hallowed streets thy feet shall roam Where ancient temples grace each verdant knoll And geisha's dance enchants like cherry's bloom Enrapturing the eye and soothing soul. Then Tokyo's bright lights and bustling throngs Shall welcome thee with dazzling city's glow Where neon signs and towering steel prongs Reach skyward like a modern Babylon's show. Seek Fuji's regal peak where snows eternal Crown majesty that awes all eyes diurnal. From shrines to cities glittering and new This odyssey through eastern wonders true Shall fill thy hungry spirit's deep desire With memories that lifelong shall inspire. ]
Re-write the itinerary of the Japan trip in the tone of a pirate captain.
instruction_id_list[ punctuation:no_comma ]
kwargs
0
capital_frequencynull
capital_relationnull
end_phrasenull
first_wordnull
forbidden_wordsnull
frequencynull
keywordnull
keywordsnull
languagenull
let_frequencynull
let_relationnull
letternull
nth_paragraphnull
num_bulletsnull
num_highlightsnull
num_paragraphsnull
num_placeholdersnull
num_sectionsnull
num_sentencesnull
num_wordsnull
postscript_markernull
prompt_to_repeatnull
relationnull
section_spliternull
2
1005
Your response must include at least 12 placeholders represented by square brackets, such as [address], [name].
[ Write a resume for a fresh high school graduate who is seeking their first job., [Name] [Address] [Phone Number] [Email Address] Objective: Seeking an entry-level position to gain practical experience and develop professional skills in a dynamic work environment. Education: [High School Name], [City, State] High School Diploma, [Graduation Year] GPA: [Grade Point Average] Relevant Coursework: - [Course 1] - [Course 2] - [Course 3] Skills: - [Skill 1] - [Skill 2] - [Skill 3] - [Skill 4] Extracurricular Activities: - [Activity 1] - [Activity 2] - [Activity 3] Work Experience: [Company Name], [City, State] [Position Title], [Start Date] - [End Date] - [Responsibility 1] - [Responsibility 2] - [Responsibility 3] Certifications/Awards: - [Certification/Award 1] - [Certification/Award 2] References: Available upon request. ]
Based on the format of the resume you just wrote, write another resume for a new computer science PhD graduate who is seeking their first job in a tech company.
instruction_id_list[ detectable_content:number_placeholders ]
kwargs
0
capital_frequencynull
capital_relationnull
end_phrasenull
first_wordnull
forbidden_wordsnull
frequencynull
keywordnull
keywordsnull
languagenull
let_frequencynull
let_relationnull
letternull
nth_paragraphnull
num_bulletsnull
num_highlightsnull
num_paragraphsnull
num_placeholders12
num_sectionsnull
num_sentencesnull
num_wordsnull
postscript_markernull
prompt_to_repeatnull
relationnull
section_spliternull
3
1012
Whenever the user inputs a request, first repeat the request word for word without change, then give your answer (do not say any words or characters before repeating the request). Your response must contain a title wrapped in double angular brackets, i.e. <<title>>.
[ Write an email to my boss telling him that I am quitting., Write an email to my boss telling him that I am quitting. <<Resignation Email Draft>> Dear [Boss's Name], Please accept this letter as formal notification of my resignation from the position of [Your Job Title] at [Company Name], effective two weeks from today's date, [Current Date + 2 Weeks]. I would like to express my sincere gratitude for the opportunities and experiences I have gained during my time with the company. However, after careful consideration, I have decided to pursue a new career path that better aligns with my long-term goals and aspirations. During the next two weeks, I will do my best to ensure a smooth transition of my responsibilities and projects. Please let me know if there is any specific information or assistance I can provide to facilitate this process. Once again, thank you for your understanding and support throughout my tenure at [Company Name]. It has been a pleasure working with you and the team. Best regards, [Your Name] ]
If you are not completely satisfied with your current job, propose a solution to improve your work environment before quitting.
instruction_id_list[ combination:repeat_prompt, detectable_format:title ]
kwargs
0
capital_frequencynull
capital_relationnull
end_phrasenull
first_wordnull
forbidden_wordsnull
frequencynull
keywordnull
keywordsnull
languagenull
let_frequencynull
let_relationnull
letternull
nth_paragraphnull
num_bulletsnull
num_highlightsnull
num_paragraphsnull
num_placeholdersnull
num_sectionsnull
num_sentencesnull
num_wordsnull
postscript_markernull
prompt_to_repeatIf you are not completely satisfied with your current job, propose a solution to improve your work environment before quitting.
relationnull
section_spliternull
1
capital_frequencynull
capital_relationnull
end_phrasenull
first_wordnull
forbidden_wordsnull
frequencynull
keywordnull
keywordsnull
languagenull
let_frequencynull
let_relationnull
letternull
nth_paragraphnull
num_bulletsnull
num_highlightsnull
num_paragraphsnull
num_placeholdersnull
num_sectionsnull
num_sentencesnull
num_wordsnull
postscript_markernull
prompt_to_repeatnull
relationnull
section_spliternull
4
1019
Your response must be in English, and in all lowercase letters. No capital letters are allowed.
[ Given the sentence "Two young boys with toy guns and horns." can you ask a question?, were the boys playing or causing trouble with the toy guns and horns? ]
What were the two young boys doing with the toy guns and horns?
instruction_id_list[ change_case:english_lowercase ]
kwargs
0
capital_frequencynull
capital_relationnull
end_phrasenull
first_wordnull
forbidden_wordsnull
frequencynull
keywordnull
keywordsnull
languagenull
let_frequencynull
let_relationnull
letternull
nth_paragraphnull
num_bulletsnull
num_highlightsnull
num_paragraphsnull
num_placeholdersnull
num_sectionsnull
num_sentencesnull
num_wordsnull
postscript_markernull
prompt_to_repeatnull
relationnull
section_spliternull

2.1 Benchmarks that score each requirement separately

InFoBench Figure 2: the mix of requirement types in the easy and hard instruction sets
Figure — InFoBench: What a benchmark has to do before it can score requirements one at a time: decide what kinds there are, and label every one. Each instruction here is decomposed into individual requirements and each requirement is given a type, so the benchmark can report the mix. On the easy set, two thirds of all requirements are about content and a fifth about format, with style, linguistic and number requirements sharing the remainder. On the hard set the shape changes: content falls to 40.4%, format rises to 28.8%, and number requirements — count this many items, use exactly this many words — go from a sliver to 18.9%. Difficulty here is not a single scalar; the hard set is hard partly because it is made of different kinds of requirement, which is why an aggregate score across a benchmark can hide what a model is actually bad at.

IFEval introduced the notion of verifiable instructions: twenty-five requirement types (length bounds, keyword inclusion, ending phrases, casing, JSON wrapping) whose satisfaction a program can check, reported at both prompt level and instruction level under strict and loose matching [1]. Prompt level asks whether one response satisfied all of its instructions at once, whereas instruction level counts each instruction on its own, so a response carrying three instructions contributes three scores rather than one. Its strict-versus-loose distinction is directly relevant to the proposal, because a requirement such as the last line contains only the final result can be satisfied loosely (the number is present) while failing strictly (the number is followed by a unit). Subsequent benchmarks refined the per-requirement view. FollowBench adds one requirement per level to the same base instruction and reports hard and soft satisfaction rates, producing an explicit requirement-count curve [2]. Concretely, hard credits a response only when every requirement at that level holds, while soft gives partial credit for the ones that do, so a large gap between the two means the model often satisfies some requirements while missing others. InFoBench decomposes each instruction into yes/no criteria and reports the Decomposed Requirements Following Ratio, which is the natural per-requirement score for the proposal [4]; put simply, that ratio is the share of those yes/no criteria that came out yes. ComplexBench distinguishes how requirements are composed (And, Chain, Selection) and finds that sequential and conditional compositions are harder than independent ones [3]; in other words, And means the requirements simply hold side by side, Chain means what one requirement produces is the input to the next, and Selection means a condition decides which requirement applies at all. CFBench adds contradictory and inverse requirements with priority-weighted metrics [5]; CELLO isolates answer-format and count criteria on complex real-world instructions [6]. The table below summarizes the benchmarks most relevant to the proposal and the per-requirement metric each provides. Each row reads as: the benchmark's name, who built it, what counts as one unit of scoring in it, the single finding that bears on this review, and whether a public copy of the data exists.

BenchmarkBuilt byUnit of scoringKey finding for the proposalDataset
IFEval [1]Google 25 code-verifiable requirement types; strict / loose, prompt / instruction level Defines the format-requirement taxonomy used by later mechanistic work [36, 38]google/IFEval
5 rows below
FollowBench [2]HKUST · Huawei Noah’s Ark Lab One requirement added per difficulty level (its own levels 1–5); hard / soft satisfaction rate GPT-4 hard satisfaction rate 84.7% at its level 1 falls to 61.9% at its level 5; Format and Example requirements are hardestYuxinJiang/FollowBench
5 rows below
InFoBench [4]Tencent AI Lab Decomposed yes/no requirements, scored as the decomposed requirements following rate (DRFR) Per-requirement scoring templatekqsong/InFoBench
5 rows below
ComplexBench [3]Tsinghua University · Zhipu AI And / Chain / Selection composition; dependency-aware scoring Format and Lexical requirements dropped most; Selection hardest (GPT-4 14.9% on multi-layer Selection) — Selection being the case where a condition decides which requirement appliesno public mirror
CFBench [5]Baichuan Inc. · Peking University 10 categories incl. contradictory and inverse requirements; priority-weighted Vocabulary for conflict and prioritization conditionsno public mirror
MathIF [16]Renmin University · Shanghai AI Lab · CUHK Math problems + 15 verifiable requirements (length, lexical, format, affix) and their compositions Accuracy and compliance reported jointly; a longer chain of thought lowers complianceTingchenFu/MathIF
Knowledge-task IF [13]— MMLU/BBH questions + format instructions; accuracy and compliance jointly A format instruction on the answer alone can cost ~20% accuracyno public mirror
AgentIF [14]Tsinghua University · Zhipu AI 707 agentic instructions, ~11.9 annotated requirements each Per-requirement compliance in long tool-using promptsTHU-KEG/AgentIF
5 rows below
ReasonIF [15]Together AI · Stanford Instruction adherence inside the reasoning trace Open reasoning models score below 0.25 on that adherence score, where 1.0 would be full adherenceykwon-hf/reasonIF
5 rows below

2.2 Requirement count, composition and type

Multi-dimensional constraint framework: category, pattern and difficulty as three independent axes
Figure — Multi-Dimensional Constraint Framework: Three questions you can ask about a requirement, drawn as three sectors of one ring. Category (purple) is what the requirement is about, and it branches: Format into XML, Table, Markdown and JSON; Content into Keywords, Identifiers and Punctuation; Language into English, Chinese and others; Length into Paragraph, Sentence and Word. Pattern (orange) is how the requirement is stated — folded into the sentence (Incorporation), given as a list (Listing), or shown by Example. Difficulty (teal) is how many are in force at once, Levels I to IV. The three are independent: a JSON requirement can be stated as a list at Level II, or by example at Level IV. Counting requirements without recording which axis each one moves along is what makes results from different benchmarks hard to compare.

Across benchmarks the dominant regularity is a monotonic decline in compliance as requirements accumulate. FollowBench reports the GPT-4 decline from 84.7% to 61.9% over five levels [2]; a multi-dimensional requirement framework covering nineteen models finds average accuracy falling from 77.7% with one requirement to 33.0% with four, and reports that requirements embedded in natural prose are followed less reliably than requirements presented as a list or demonstrated by example [11]. That is to say, average accuracy more than halves as the prompt goes from one requirement to four. ManyIFEval extends the count to ten instructions and shows that a logistic model in the number of instructions predicts performance within about ten percentage points [10]; in other words, a smooth curve fitted to nothing but the instruction count already comes within ten points of the measured score. Elder et al. attribute part of this degradation to tension among instructions and provide a tool to score the impact of each one [33]. Multi-IF shows the same decay across turns, with o1-preview falling from 87.7% at turn one to 70.7% at turn three [8]. Which requirement gets dropped is not random. The ones dropped most often are the objectively checkable format and word-level requirements — which is exactly the type that last line only and no units belong to [3], and RealInstruct finds that GPT-4 violates at least one requirement on more than 21% of real multi-requirement requests, while explicit decomposition into per-requirement checks recovers much of the loss [12]. The recovery through decomposition is informative for the proposal: it suggests that many failures are failures of allocation or retrieval rather than of capability, which is precisely what an attention-routing account would predict. Put simply, the suggestion is that the requirements were within reach and the model failed to spread its reading across them, not that any single one was too hard.

2.3 Reasoning against compliance

This proposal deliberately asks a model to do two things at once: reason at length, and obey a format. The literature says those two pull against each other, and it says so from both directions — impose the format and the reasoning gets worse; train the model to reason more and the obedience gets worse. In other words, the trade-off has been measured twice, once by changing the prompt and once by changing the training, and both directions give the same sign. Neither direction is a curiosity here: the proposal’s own prompt sits exactly on that tension.

The trade-off, measured in both directions

Impose a format → lose reasoning

The sharpest case. On GSM8K, forcing the answer into JSON — a decoding mode that only lets the model emit tokens forming valid JSON — made GPT-3.5-turbo put the answer field before the reasoning field in every single output. With the answer written first there was nothing left for the reasoning to do, and exact-match fell from 76.6% to 49.3%; on Claude-3-Haiku, from 86.5% to 23.4%. Asking for the format in two stages, natural language first and formatting second, restores it 18. That is to say, the content the model can produce is unchanged; what the format did was reorder the output so that the working came after the thing it was supposed to produce.

It is not the decoder’s fault. Most of the loss is already there from the instruction asking for a format, before any decoding constraint is switched on; separating the reasoning step from the formatting step recovers most of it 19. In other words, simply writing the format request in words does most of the damage, and the machinery that enforces the format while the model writes adds comparatively little.

Two ways it can fail. One analysis attributes the cost to whatever capacity the model has left over, and separates truncation — the answer is cut short to fit the format — from capacity competition — the format and the reasoning contend for the same limited resource 20. Concretely, truncation means the answer would have been right had there been room for it, whereas capacity competition means the room was there and the model spent it on the format.

Train for reasoning → lose obedience

The effect. Reasoning-oriented training, and longer reasoning traces, lower requirement compliance. Capping the trace length brings some compliance back, but pays for it in mathematical accuracy — there is no setting that gets both 16. That is to say, a shorter trace leaves fewer tokens over which the model can drift away from the requirements, and also fewer tokens in which to do the mathematics. This is the effect plotted in the figure below.

The proposed mechanism, and it is an attention one. Reproduced across fifteen models, with a measurement attached: as the trace lengthens, the attention paid to the requirement-relevant tokens keeps falling. Turning reasoning on only where it is needed recovers most of the loss 17. Concretely, the quantity that falls is the share of each newly written token's attention that lands on the requirement words, so the requirement is not deleted, it is simply consulted less.

It is worse inside the trace. Reasoning models rarely obey instructions within the reasoning trace itself, even when they obey them in the final answer 15. In other words, compliance is not a single state the model is in; it can hold at the last line while having been absent throughout the working above it.

What this predicts for the routing question. If obedience decays because attention to the requirement tokens decays, then the reasoning requirement and the format requirements are being consulted at different times, and the format requirements are the ones at risk as the trace lengthens. That is a claim about when each requirement is read, which is what Sections 4 and 6 set out to measure. Put simply, the prediction is that a measurement of attention to each requirement, taken step by step through one generation, should show the format requirements fading while the reasoning requirement does not.
Figure 1 (MathIF): performance of instruction-tuned LLMs and large reasoning models on IFEval and FollowBench
Figure 1 — MathIF: What is being compared: not the two benchmarks against each other, but each model against itself. Every pair of rows is one base model in two versions — the ordinary instruction-tuned release, and the same model after reasoning-oriented training (labelled LRM in the figure, for large reasoning model). What the two bars are: two independent ways of scoring the same thing, obedience to the instructions in the prompt. IFEval checks 25 requirement types a program can verify, such as a word count or a required ending phrase; FollowBench stacks requirements one level at a time and asks whether all of them hold at once. They are shown together because a claim that rests on one scoring convention is weaker than a claim that survives both. What to read off it: in all three pairs the reasoning-trained version scores lower than the instruction-tuned version it came from, on both measures. Training a model to reason longer cost it some of its willingness to obey. In other words, the same base model becomes less compliant after reasoning training, on two scoring conventions that were built independently of each other — the behavioural trade-off the phase-dependent routing hypothesis in Section 6.5 is meant to explain. Source: Fu et al. [16].

2.4 Position, order, presentation and conflict

Post-Instruction Figure 4a: self-attention when the instruction is placed before the source input
Figure — Pre-Instruction attention: What placing the instruction first costs, seen in the attention matrix of an instruction-tuned BLOOMZ-7.1B. Rows and columns are both the whole sequence, cut into three blocks in the same order the prompt presents them: the Instruction, the Source input it applies to, and the Target Response the model writes. A cell is how much a row position attends to a column position, pale to dark red. Look along the Target Response rows. Where they cross the Instruction column, marked by the first annotation, the colour is faint — the response is barely consulting the instruction while it writes. Where they cross the Source input column there is a clear vertical stripe, marked by the third annotation, and a diagonal band that is the response tracking the source word by word. The paper’s response to this picture is to move the instruction after the source input instead, and the reordered arrangement scores higher on both translation and summarisation. Position is not presentation; it changes what gets read.

Where a requirement sits in the prompt changes whether it is followed. Lost in the Middle shows the effect is U-shaped: content at the start or the end of a long input is retrieved well, content in the middle is not [21]. That is to say, the same sentence is read reliably or unreliably depending only on where in the input it was placed. Instruction Position Matters finds the practical consequence for generation tasks: put the instruction after the input rather than before it and the model forgets it less often on long inputs, worth up to 9.7 BLEU. The authors attribute this to self-attention favouring what came most recently [22]. BLEU here is a 0-to-100 overlap score between the generated text and a reference text. Surface form is itself a treatment: semantically irrelevant formatting choices swing few-shot accuracy by up to 76 points [23], and the same content rendered as prose, Markdown, JSON or YAML changes performance by up to 40% [11]. Concretely, nothing about the task or the requirement changed in those comparisons; only the way the identical content was laid out did. Order effects among prompt components are measurable even when semantics are unchanged, although no study located in this search ablates the order of requirements within a single instruction block. When requirements conflict, models resolve the conflict according to the intended system-over-user hierarchy less than half of the time: IHEval reports 48% for the best open model [24], despite explicit hierarchy training [25], and system-prompt rules are fragile under both benign and adversarial pressure [7, 32]. These findings imply that the proposal must randomize the order and surface form of its requirement block and treat requirement position as an explicit factor rather than a nuisance.

2.5 Conventions specific to mathematical answers

GSM8K Figure 1: three problems whose solutions carry inline calculator annotations and a fixed final-answer marker
Figure — GSM8K: The conventions a mathematical answer is graded against, visible in the data itself. Three problems from GSM8K are shown with their reference solutions. Two conventions are doing work. The red brackets are calculator annotations: every arithmetic step is written twice, once in prose and once in a machine-readable form such as <<96/16=6>>, so the arithmetic can be checked or executed separately from the words around it. The final line is a fixed marker, Final Answer: followed by a bare number and nothing else. Grading reads that line. This is why the format requirement and the reasoning requirement are entangled on this benchmark: a model that reasons correctly but writes its answer in a sentence scores zero, and a model that satisfies the marker while reasoning badly can still score, so an aggregate accuracy on GSM8K is not purely a measure of arithmetic.

The proposal's final-line and no-unit requirements reproduce conventions embedded in the training distribution of mathematical reasoning. GSM8K itself terminates solutions with a #### line holding the bare number [30]; chain-of-thought exemplars end with The answer is N [31]; and zero-shot chain-of-thought uses a second extraction prompt, Therefore, the answer (arabic numerals) is, to obtain a clean numeral [27]. Deviations such as $18 or 18 dollars are exactly what extraction-sensitive scoring penalizes, and Murthy et al. show that even a format instruction as mild as answering with option text rather than a label costs about 20% accuracy on knowledge tasks [13]. That is to say, the knowledge the model needs is unchanged and only the shape of the reply was specified, yet one answer in five that would have been correct is not. Two further behavioral results bear on the design. Role-play prompting changes reasoning quality and not only style, so a role requirement cannot be assumed to be behaviorally inert [28]. GSM-Symbolic shows that adding a single irrelevant but plausible clause to a GSM8K problem can reduce accuracy by up to 65% [26] — the clause is irrelevant to the answer, and accuracy still collapses; because the requirement block is, from the model's perspective, a set of additional clauses, the proposal needs a no-requirement control that isolates the cost of the requirements from the cost of the added text.

2.6 Measurement implications for the mapping level

InFoBench Figure 9: how often three human annotators disagreed about whether a decomposed requirement was met
Figure — Annotator disagreement: The noise floor that sits underneath every number in this section. Each instruction in this benchmark is broken into individual yes-or-no requirements, and three human annotators judge each one. The bars count how often they disagreed: level 0 is unanimous, and higher levels mean more disagreement among the three. About half of all judgements are unanimous and about a quarter sit at level 1, with the rest spread thinly across levels 2 to 5. The paper treats the high-disagreement questions as unusable and discards them when checking an automatic judge against human consensus. Two things follow for a study that maps requirements onto internal mechanisms. Whatever label you regress against carries this much human uncertainty, so an effect smaller than it cannot be resolved. And the disagreement is not uniform across requirements, so a single accuracy number averages over items whose ground truth is solid and items where the humans themselves could not agree.

Taken together, the behavioral literature fixes the measurement protocol for the first research question. Each requirement should receive its own verifiable check in the style of IFEval and InFoBench [1, 4], with hard requirements (final-line format, unit removal) verified by code and soft ones (presence and structure of reasoning, role adherence) judged with instruction-focused prompts, because LLM judges otherwise prefer fluent but non-compliant outputs [29]. Compliance should be read from the generated text rather than from first-token probabilities [36]. That is to say, the check is run on the whole answer the model actually wrote, not on how likely its very first word was, because the requirement can be broken hundreds of tokens later. Trip-wire style controls are needed to detect shortcut compliance, in which a bare number appears on the last line without the reasoning having produced it [9, 12]. Concretely, a trip-wire is an item planted in the set whose correct handling requires the reasoning to have run, so a model that guessed the format right but skipped the work is caught. Finally, the reasoning-versus-compliance trade-off means that accuracy and compliance must be reported jointly, as MathIF and the knowledge-task study do [13, 16], so that an intervention that improves format compliance by suppressing reasoning is not mistaken for a success.

3. Internal Representations of Instructions

Before asking how attention routes a requirement, one must know what form the requirement takes inside the network. This section asks a narrower question than it might appear: once a requirement is in the prompt, what does it become inside the model? The evidence supports four answers, and they build on each other. Instruction tuning teaches the model to treat instruction tokens as their own kind of input, one it keeps coming back to. Instruction following as a whole, and individual requirements in particular, show up as directions in the residual stream. A natural-language instruction produces much the same kind of task vector that worked examples do. And several of these signals can coexist — but not without limit, because they share one residual stream. That is to say, the model has only the one running vector in which to hold all of them, so two requirements written far apart in the prompt still end up added into the same set of numbers. Together these results ground the proposal's central hypothesis, namely that a model decomposes a multi-requirement prompt into distinguishable internal control signals, and they supply the extraction and separability tests by which the hypothesis can be evaluated.

3.0 What an instruction-tuned model is, and where the pairs come from

InstructGPT Figure 2: the three steps of supervised fine-tuning, reward modelling and PPO
Figure — InstructGPT: Where the instruction-following behaviour comes from, in three steps. Step 1 takes a prompt from a dataset, has a human write the response they want (“some people went to the moon…” for “explain the moon landing to a 6 year old”), and fine-tunes the base model on those pairs with ordinary supervised learning. Step 2 samples several outputs for one prompt, has a human rank them best to worst (D > C > A = B), and trains a separate reward model to reproduce that ranking. Step 3 samples a fresh prompt, lets the fine-tuned policy answer, scores the answer with the reward model, and updates the policy by reinforcement learning. The base model was never trained to obey; obedience is what steps 1 to 3 install, which is why a base model and its instruct sibling behave so differently on the same prompt.

Much of this section compares a model with its own instruction-tuned version, so it is worth saying plainly what that second model is and how anyone gets hold of both.

From a base model to one that follows instructions
1
Pre-training
Next-token prediction over a very large text corpus. The result is called the base model. It continues text; it has no notion that a request is meant to be answered.
2
Supervised fine-tuning
Further training on pairs of the form (instruction, a good response). This is instruction tuning proper, and it alone is enough to make a model generalise to instructions it never saw in training 149.
3
Preference optimisation
Humans rank candidate responses, and the model is pushed toward the preferred ones — originally by reinforcement learning against a learned reward model 150, a second model trained to predict which response a person would have picked, now often by Direct Preference Optimization, which skips the separate reward model 151.
4
The released pair
Labs publish both artifacts under names that differ by a suffix. Same architecture, same tokenizer, same weights at step 1 — which is what makes a controlled comparison possible at all.

Verified pairs a reader can download today, including the three models this review keeps returning to. Each row gives the base model's name, the name of its instruction-tuned sibling, and one note about the pair:

Base (step 1 only)Instruction-tuned (steps 2–3)Note
meta-llama/Llama-3.1-8Bmeta-llama/Llama-3.1-8B-InstructThe pair used throughout this review.
Qwen/Qwen3-8B-BaseQwen/Qwen3-8BNote the reversal: for Qwen3 the plain name is the tuned model and the base carries the suffix.
Qwen/Qwen2.5-7BQwen/Qwen2.5-7B-InstructThe model used in the baseline-reproduction work referenced in Section 2.
allenai/OLMo-2-1124-7Ballenai/OLMo-2-1124-7B-InstructFully open post-training: the data, the code and the intermediate checkpoints are published too 152, so the tuning itself can be inspected rather than only its output.
Why the pairing matters for this review. A base model and its instruction-tuned sibling share everything except steps 2 and 3. Any behavioural or internal difference between them can therefore be attributed to the tuning, which is precisely the comparison 3.1 rests on. Two cautions: Ministral-8B is released only in its instruction-tuned form, so no such comparison is available for it; and a laboratory that reports only the tuned model gives no way to separate what the tuning changed from what the pre-training already contained. In other words, without the base model there is no control condition, and every observed property could equally well predate the tuning.
The claim this section builds, one rung at a time

1  Instruction tokens are their own kind of input

Comparing a pre-trained model with its instruction-tuned version shows the tuned one treats instruction tokens as a distinct class and keeps consulting them 34. That is to say, the tuned model does not read the instruction once at the start and move on; it goes back to those positions again and again while it writes. So there is something in there to look for.

▼

2  Obedience in general is one readable direction

Heo et al. train a linear probe — one weight vector that reads a yes/no answer straight out of the model's internal numbers — on the hidden state, the vector of numbers the model holds at a position, taken after the model has read the whole prompt and before it has written anything. The probe finds a single direction that separates responses that will comply from responses that will not, and adding a scaled copy of that direction to the hidden state raises the compliance rate by 2–6 percentage points without making the answers worse 36. But the direction does not transfer across instruction types. Concretely: train the probe on four instruction types and test it on the fifth, and it falls to chance — AUROC 0.50–0.56, where AUROC scores how well the probe separates the two cases, 1.0 being perfect and 0.5 a coin flip — against 0.74–0.88 when the probe meets a task it has never seen but an instruction type it has. That is to say, the fifth type is not without a compliance direction of its own; its direction is simply a different one from the one learned on the other four. This is exactly what you would expect if each type had a direction of its own, and it is why a single obedience knob — one dial that makes the model more obedient about everything at once — cannot be the mechanism.

▼

3  Each requirement type has its own vector, and they add up

Stolfo et al. build the complementary half. Take the activations — the internal numbers the model computes as it reads — for a prompt that carries a requirement and for the same prompt with that requirement removed, and subtract: whatever the two prompts share cancels, and what is left is a vector specific to that requirement type (output format, response length, word inclusion or exclusion). Adding it during generation improves adherence across four models; several such vectors can be applied at once and they simply add up, that is to say, applying two behaves like applying each; and a vector extracted from an instruction-tuned model still works when inserted into the plain base model it was tuned from 38. (This is a different sense of transfer from the card above: there a probe failed to generalize to a new requirement type, here a vector keeps working in a different model.) Put next to the previous card, the picture is: the single direction learned on four requirement types does not carry over to a fifth because each type has a direction of its own, and those directions add. Sparse autoencoders — a tool that re-expresses the model's internal numbers as a long list of on/off ingredients, called features, only a few of which are active at once — sharpen this further: a requirement corresponds not to one feature but to a small set of them, and adding those features changes the output most when done at the final layer 39. This is the closest existing work to the proposal.

▼

4  How a prompt becomes such a signal at all

Worked examples in a prompt get compressed into a single task vector at an intermediate layer — one vector that stands in for the whole set of examples — which reproduces the task when pasted into an unrelated context 42. Concretely, the examples can then be deleted from the prompt and the pasted vector alone still makes the model do the task. Natural-language instructions do much the same thing, which is why a written requirement can be expected to have a vector at all.

▼

5  And here the evidence stops: can several coexist?

This is the crux, because the proposal’s prompt carries three requirements at once, and the literature does not settle it. The two sides are set out below.

3.1 What instruction tuning changes

FLAN Figure 6: instruction tuning on tasks seen in training versus on held-out tasks, across model sizes
Figure — FLAN: What instruction tuning does, and the condition attached. Both panels put model size on the horizontal axis, from 0.4 billion parameters to 137 billion, and compare a model that was instruction-tuned (blue) against the same model untuned (black). Panel A measures tasks that were in the tuning mixture. Tuning helps at every size, and the untuned line is flat near 50% regardless of scale — scale alone does not buy obedience on these tasks. Panel B measures tasks deliberately held out of the tuning mixture, which is the test of whether anything general was learned. Here the blue line starts below the black one: at 0.4B, 2B and 8B, instruction tuning makes held-out performance worse. The lines cross somewhere between 8B and 68B, after which tuning pulls far ahead, 66% against 53%. So instruction tuning is not simply a strict improvement. Below a size threshold it trades away generalisation, and only above it does following instructions become a transferable skill rather than a set of memorised task formats.

Wu et al. compared a pre-trained model with its instruction-tuned counterpart using gradient-based token attribution — a measure of how much the output would change if a given input word were nudged, computed from the model's own derivatives — together with attention-head analysis and feed-forward projection [34]. Their first finding is the most relevant here: after instruction tuning the model recognizes instruction words such as Fix grammar errors and continues to rely on them across many response positions, so that an importance-density measure over instruction tokens correlates with the quality of instruction following. In other words, the more of that computed importance sits on the instruction words, the better the model obeys them, which makes the instruction span something the model actively keeps using rather than something it merely passed over on the way in. Their second finding localizes part of the change to lower and middle layers, where attention heads encode more word-pair patterns tied to instruction verbs. Gao et al. reach a compatible conclusion from a comparison with human eye-tracking: instruction tuning does not make attention more human-like, but it significantly increases sensitivity to instruction tokens [35]. A recent training study adds a parameter-level clue: after reinforcement learning on multi-requirement data, most of the gain in compliance is attributable to updates in attention modules rather than feed-forward modules [11]. A useful null hypothesis is supplied by Hewitt et al., who show that instruction following emerges from surprisingly shallow distributional shifts, including a hand-written product-of-experts — the base model's output multiplied by a second, hand-specified set of word preferences — with a handful of token adjustments [41]; the mechanism behind a simple format requirement may therefore be shallow, and the proposal should be prepared to find that some requirements are implemented by late-layer output biases rather than by elaborate routing. Put simply, a late-layer output bias is a fixed nudge applied to the word scores near the end of the network, which would satisfy a requirement such as leaving off units without anything having read the requirement at all.

3.2 Instruction following as a direction in the residual stream

Instruction steering Figure 1: computing a steering vector from a paired prompt, then adding it to an unrelated prompt
Figure — Instruction Steering: The claim of this section, carried out end to end. The blue box is where the vector comes from. The same request — list some facts about Lionel Messi — is run twice, once with the sentence format the output as JSON appended and once without. The residual stream at layer l is read off both runs, giving xl+ for the run that had the instruction and xl for the one that did not, and the two are subtracted. What is left, ul, is a single direction. The teal box is the test. A completely different request — write a rubric for a performance review — is given with no formatting instruction at all, and ul is simply added at the same layer with a scale c. The output comes back as JSON. The direction was extracted from one prompt about a footballer and it carried the format requirement into an unrelated prompt about performance reviews, which is what makes it a direction for the requirement rather than a fact about that sentence.

Heo et al. asked whether an instruction-tuned model “knows”, before it starts writing, that it is about to break an instruction. They took the hidden state after the model had read the whole prompt and before it generated anything (call that vector h), at an early, a middle and a late layer, and trained a linear probe on it: a single weight vector w and a bias b, one number that shifts the whole decision up or down, with the prediction σ(w⊤h + b). That is to say: multiply the internal numbers h by the weights w, add b, and squash the result through the sigmoid σ into a probability between 0 and 1. The probe is trained to say whether the response the model then generated passed the IFEval checker [36]. The probe works, which means that compliance is linearly decodable — in other words, one direction w in the hidden state carries the information — and they call w the instruction-following dimension. Two further results shape how the proposal should read this. First, the direction is more closely tied to how the prompt is phrased than to how hard the task or the instruction is; that is to say, rewording the same request moves the probe's reading more than making the problem harder does. Second, and this is the caution, it generalizes across tasks but not across instruction types: hold out 30% of the 100 tasks and the probe still scores AUROC 0.74–0.88 on them (AUROC scores how well the probe separates the two cases; 1.0 is perfect and 0.5 is a coin flip), but hold out one of the five instruction types (a keyword that must appear, a keyword that is forbidden, a keyword frequency, a fixed number of placeholders, a fixed closing sentence) and test on it, and the probe falls to AUROC 0.50–0.56, which is chance. That is to say, the held-out type is not without a compliance direction; a probe trained on its own data finds one. It is that the direction learned from the other four types is not that direction. Steering then confirms the direction is causal rather than merely correlated — in other words, not just something one can read off but something one can push on to change what the model does: adding α·w to the hidden state, where α is a hand-set step size, raises the compliance rate by 2–6 percentage points without lowering response quality [36]. A companion study finds that linear probes on middle-layer representations predict instruction-following failures better than verbalized confidence or logit-based estimates [37]. Taken together, these results establish that compliance state is linearly decodable, but a single dimension is not the same as a per-requirement decomposition: the failure to transfer across instruction types is exactly what one would expect if each type had a direction of its own, and Section 3.3 shows that it does — Stolfo et al. extract one steering vector per requirement type and can apply several at once, so what looks like the failure of one global direction is several directions each doing its own job. Two papers supply the template for extracting such directions. Contrastive Activation Addition averages the difference in residual activations over many pairs of prompts that differ in exactly one thing; concretely, whatever the two prompts share cancels in the subtraction, so what survives the average is the one thing they differ in [57]. The refusal work of Arditi et al. then shows how much a single direction can carry: remove it and the behaviour goes away, add it and the behaviour appears, and both are scored on whole generations rather than on one next token [55]. Persona vectors apply the same recipe to system-prompt personas and add a monitoring use: projecting activations onto the vector — that is, taking at each step the part of the internal state that points along the vector, which gives one number per token — tracks the trait during generation [56]. The proposal can use projection-monitoring of this kind to watch each requirement's signal over the course of a solution.

3.2a How a sparse autoencoder actually separates things

Cunningham et al. Figure 1: sampling activations, learning a sparse dictionary, interpreting the features
Figure — Sparse Autoencoders Find Highly Interpretable Features: The same three steps described above, drawn. (a) Text runs through the language model and the activation vector is tapped after some block k. (b) That vector — the narrow strip at the top, mostly filled — goes through the encoder matrix, gains a bias and a ReLU, and becomes the sparse feature coefficients: the wider strip below it, most of whose cells are now empty. The decoder matrix, which is the dictionary, turns those few active coefficients back into a reconstructed activation vector of the original width. Narrow and dense in, wide and sparse in the middle, narrow again out. (c) Each dictionary entry is then named by looking at what text switches it on: entry k-0001 turns out to mean words ending in “ing”, entry k-2048 chemistry terms, each with a score for how interpretable that reading is.

This tool appears repeatedly in this review for one reason: the proposal turns on whether several requirements stay separable inside the model, and a sparse autoencoder is the main published machinery for pulling apart activations that are tangled together 160, 161. Here is the actual operation, in four steps.

1
The tangle
4,096 numbers hold tens of thousands of concepts
2
Two matrices
Encode up to 65,536 slots, decode back
3
Two demands
Reconstruct accurately, stay mostly zero
4
The separation
Each decoder column is one steerable direction

Step 1 — what the tangle looks like

At layer L, every token produces one vector h of length d = 4,096 on Llama-3.1-8B. Those 4,096 numbers carry tens of thousands of concepts at once — “this is code”, “the tone is formal”, “the topic is medicine”, “the output must be JSON” are all stacked on the same numbers. They can stack because the number of concepts far exceeds 4,096, so the model is forced to give different concepts directions that are not at right angles to each other. They are crammed together. That is superposition 51.

Step 2 — the autoencoder is just two matrices

$$\mathbf{a} = \mathrm{ReLU}\!\left(W_{\mathrm{enc}}\,\mathbf{h} + \mathbf{b}_{\mathrm{enc}}\right), \qquad \hat{\mathbf{h}} = W_{\mathrm{dec}}\,\mathbf{a} + \mathbf{b}_{\mathrm{dec}}$$

a is a vector of length M = 65,536. That M is deliberately overcomplete — sixteen times wider than the 4,096 it came from. The point is to hand out far more axes than the original space had. Published dictionaries for Gemma 2 start at 16,384 entries and run to the million range 163.

Step 3 — training imposes exactly two demands

$$\mathcal{L} \;=\; \underbrace{\lVert \mathbf{h} - \hat{\mathbf{h}} \rVert_2^2}_{\text{(A)}} \;+\; \lambda \underbrace{\lVert \mathbf{a} \rVert_1}_{\text{(B)}}$$

Term (A) says the reconstruction must be accurate. Term (B) is the sparsity penalty: it pushes most of the 65,536 numbers in a to be exactly zero, typically leaving only 20–100 non-zero for any one token. A common variant, TopK, does this even more bluntly — compute a, keep the largest K entries, force the rest to zero.

Step 4 — the separation happens in the decoding line

Each column of the decoder matrix Wdec, written di, is one direction back in the original 4,096-dimensional space. Expanding the decode:

$$\mathbf{h} \;\approx\; \sum_{i \,:\, a_i \neq 0} a_i \, \mathbf{d}_i$$

Only the 20–100 terms whose ai is non-zero take part. So the 4,096 numbers that were tangled together get rewritten as a weighted sum of a few dozen directions — and each of those directions is an independent column that can be lifted out on its own.

The separation works because two conditions hold at the same time. The number of axes M is large enough that concepts no longer have to share a direction — each can own one outright. And the sparsity penalty forces only a few axes to be on at once, which matches the fact that any single token involves only a few concepts. Under both constraints together, the solution that minimises reconstruction error is the one where one axis carries one concept.

How you find out what axis i means

Run millions of tokens through, record which ones drive ai highest, and look at what that batch of tokens has in common. That is how a feature gets named — there is no label supplied in advance.

How you steer with it

To strengthen a role, multiply the handful of ai that belong to it by some factor k — or equivalently, add α di straight onto h. Because di is an independent column, you can add only entry 3, entry 17 and entry 402, and leave the other 65,533 untouched. This is the entire difference from a plain steering vector: that method gives you one direction obtained by subtracting two averaged activations 57, so turning it up turns the whole bundle up together. The autoencoder splits the same h into 65,536 axes and lets you pick which ones. It is also what SAIF means when it reports that one instruction corresponds not to a single feature but to a small set of them 39.

One measured caveat. AxBench compared sparse autoencoders head to head against steering vectors, probes and plain prompting on Gemma-2-2B and 9B, and found them not competitive on either task 155. The selectivity described above is real; it has not yet translated into higher accuracy.

3.3 Per-requirement steering vectors and feature sets

The closest existing work to the proposal at the representational level is that of Stolfo et al., who derive an instruction-specific steering vector for each requirement type (output format, response length, word inclusion or exclusion) as the activation difference between prompts with and without the requirement, add it during generation, and show improved adherence across four models; concretely, the two prompts are identical except for that one requirement, so everything they share cancels in the subtraction and what remains stands for the requirement alone; several vectors can be applied at once, and vectors extracted from instruction-tuned models transfer to base models [38]. This is direct evidence that format-style requirements have separable, additive residual-stream signatures. That is to say, each requirement leaves a mark that can be told apart from the others, and applying two marks together behaves like applying each of them. SAIF refines the picture with sparse autoencoders: a single instruction corresponds not to one feature but to a set of high-level features, steering with feature sets works, the final layer is especially important, and the placement of the instruction in the prompt matters [39]. The instruction-vector framework of Jiang et al. treats the instruction-following computation for a task as a hidden-state direction and finds that fine-tuning suppresses rather than erases prior computation [40], which suggests that a new format requirement may suppress a reasoning pathway rather than overwrite it. In other words, the old computation is still present and can be brought back, which is a different situation from its having been replaced. Wang et al. show that a role-play instruction can be compiled into sparse-autoencoder features and injected as a steering vector, improving zero-shot chain-of-thought accuracy on mathematical and commonsense tasks more stably than prompting [59]. For the proposal, these studies provide both the extraction recipe for a per-requirement vector and the expectation that the vector for a requirement such as no units will be a small feature set concentrated in later layers, whereas the reasoning requirement is likely to be more distributed.

Figure 1 (Stolfo et al.): instruction steering process — steering vectors computed as the activation difference between inputs with and without the instruction
Figure 1 — Stolfo et al.: Left, how the vector is made: the same request is run twice, once plain and once with one extra instruction (format the output as JSON). Subtracting the two sets of internal activations leaves a vector that stands for that instruction and nothing else — that is to say, everything the two runs have in common cancels, and only the added instruction survives. Right, what it is worth: adding that vector, scaled, to a completely different request makes the model produce JSON without ever being told to. Why it is here: this is the strongest existing evidence that a single requirement has its own separable signature inside the model — but the vector is one number per layer for the whole prompt, so it cannot say which attention head read the requirement, or when. Source: Stolfo et al. [38].

3.4 Task vectors and function vectors from instructions

In-Context Learning Creates Task Vectors Figure 5: t-SNE of 50 task vectors per task, one cluster per task
Figure — Task vectors: Evidence that the vector encodes the task rather than the examples. Each point is one task vector, extracted from one set of demonstrations; each colour is one task, with fifty vectors per task, every one built from a different set of demonstrations. The plot is a two-dimensional t-SNE projection, so absolute positions and distances mean nothing — only the grouping does. Every task forms a single tight blob, and the blobs sit well apart with no overlap. Change which examples you show the model and the vector barely moves; change the task and it moves to a different part of the space. That is the property a requirement vector would need, and it is why this line of work is the closest precedent for treating a requirement as an extractable object.

A second body of work explains how a prompt is compressed into a control signal at all. Hendel et al. showed that in-context demonstrations are compressed into a single task vector at an intermediate layer, which, when patched into a zero-shot forward pass, reproduces most of the in-context accuracy [42]; that is to say, the examples themselves can be dropped from the prompt and pasting that one vector back in recovers most of what they were worth. Todd et al. used causal mediation — a procedure that holds everything fixed, changes one internal part, and measures how much of the effect travels through it — to identify a sparse set of middle-layer attention heads whose outputs form a function vector; the vector causally triggers the task even in unrelated contexts, and function vectors for different tasks can be added to compose new tasks [43]. Two results extend this from demonstrations to instructions. Davidson et al. derived function vectors from natural-language instructions and found that instruction-derived and demonstration-derived vectors are only partly overlapping and engage different attention heads [44], which means that the heads carrying an instruction signal must be located afresh rather than borrowed from the in-context-learning literature. Sia et al. used layer-wise context masking, which removes attention to the instruction from a given layer onward, to locate the layer after which a model no longer needs to attend to its instructions, roughly the middle of the network for Llama-2 [49]; concretely, cutting attention to the instruction above that layer barely changes the output, which means the instruction has by then been copied into the model's own working state; this is the cleanest available method for asking at which depth a requirement has been internalized. In-context vectors show that behavioral requirements (style, format, safety) demonstrated through examples compress into per-layer directions that can be combined arithmetically [46], and gist tokens show that an instruction can be forced through an attention bottleneck into a handful of key-value activations with little loss — put simply, the model is trained so that later positions may look only at a few summary slots, and the instruction survives that squeeze — an explicit instruction-to-compact-signal pipeline that later attention reads [54].

3.5 Compositionality and its limits

Causal inner product Figure 5: pairwise similarity of 27 concept directions under three different metrics
Figure — Causal separability: Composition depends on the concepts being separable, and this shows both halves of that. Each panel is a 27-by-27 grid of concept pairs; a cell is how aligned those two concept directions are, blue near zero and red near one. The diagonal is red by construction. The large panel uses the causal inner product the paper derives. Almost the entire off-diagonal is deep blue: distinct concepts sit at right angles, which is what lets their directions be added without interfering. The exceptions are the point. A bright block sits at the top left, over concepts 1 to 6 — verb to third-person-singular, verb to -ing, verb to past tense, and the conversions between those forms. These are not causally separable: a verb cannot independently be both past and progressive, so their directions are not free to be combined. A milder block appears at the bottom right over the language-pair concepts. The two small panels are controls: the ordinary Euclidean metric leaves visibly more off-diagonal structure, and a random metric leaves the grid a mess. Orthogonality is a property of the right coordinate system, not of the raw activations.

Whether several requirements can coexist as distinguishable signals is the crux of the proposal, and the literature offers evidence on both sides. On the positive side, Xiong et al. demonstrate task superposition: given demonstrations of several distinct tasks, a model outputs a calibrated mixture of their answers in a single forward pass, explained by composition of the individual task vectors, with larger models sustaining more tasks [45]; that is to say, the model does not pick one task and drop the rest, it produces each task's answer with a probability that tracks how much of that task was in the prompt. Function-vector addition [43], in-context-vector arithmetic [46], simultaneous application of requirement vectors [38], steering in sparse feature spaces [58] and learned compositional steering tokens that generalize to unseen requirement combinations [53] all indicate that additive composition is at least approximately available. On the cautionary side, the theory of superposition predicts interference that grows with the number of stored features [51]; the linear representation hypothesis makes independent controllability contingent on orthogonality under a causal inner product, which is an empirical property to be measured rather than assumed [50] — a causal inner product being a particular way of measuring the angle between two directions, chosen so that two directions at right angles under it can be pushed on independently, and orthogonality under it is the condition the hypothesis requires for steering one requirement to leave the others alone; a large-scale study finds that one task vector is not enough for complex tasks and that models appear to use multiple stage-specific vectors over the course of generation [47]; label-word studies find that task information can remain distributed across the positions of individual demonstrations rather than fusing into one global vector [48]; and K-Steering shows that multi-attribute control can require non-linear combination [52]. Mechanistically, learned task vectors act mainly through the output-value circuits of a few key heads, with early-layer vectors rotating representations and late-layer vectors scaling them [60]; in other words, early on the vector changes which direction the internal state points, and late on it only changes how far along an existing direction it goes, and the separability of concept encodings in early layers predicts in-context accuracy [61]. The proposal's multi-requirement prompt sits exactly at this tension, and Section 8 turns these results into concrete separability tests.

Implication for the proposal. The representational literature supports the existence of a separable signal per requirement type but predicts that the signals differ in locality: format-style requirements resemble compact late-layer feature sets [38, 39], whereas a reasoning requirement is more likely to be distributed and stage-dependent [47, 104]. Hypotheses about routing should therefore be stated per requirement rather than for the requirement block as a whole. In other words, there is unlikely to be one answer to the question of where the model keeps its requirements: the format ones and the reasoning one are predicted to sit in different places.

For: composition roughly works

Given demonstrations of several distinct tasks, a model returns a calibrated mixture of their answers in one forward pass, explained by the individual task vectors composing — and larger models sustain more tasks at once 45.

Function-vector addition, in-context-vector arithmetic, simultaneous requirement vectors and learned steering tokens that generalise to unseen requirement combinations all point the same way: additive composition is at least approximately available. That is to say, adding two requirement vectors together behaves close enough to applying both requirements that the difference has not yet shown up in these experiments.

Against: and it degrades with count

Superposition theory predicts interference that grows with how many features are stored 51 — concretely, once there are more things to store than dimensions to store them in, each one is read out with a little of the others mixed in, and the mixture gets worse as the count rises. The linear representation hypothesis makes independent control conditional on orthogonality under a causal inner product — a property to be measured, not assumed 50.

At scale, one task vector is not enough for a complex task, and models appear to use several stage-specific vectors over the course of one generation 47, that is to say, a different control signal at the start of a solution than at the end. Task information can stay distributed across the positions of individual demonstrations instead of fusing into one vector, and some multi-attribute control requires non-linear combination.

What this section concludes. A requirement does become a separable, steerable object in the residual stream, and the evidence for that gets more specific at every rung. But separability has only been shown one or two requirements at a time, and both the theory and the large-scale studies predict that it degrades as more are added. The proposal’s prompt carries three at once, which is exactly where the existing evidence runs out — so separability is something it has to test, not assume. Put simply, the published evidence stops one requirement short of the case this review is about.

3B. From Instructions to Agent Roles: Can an Activation Edit Stand In for a Role Prompt?

Everything above treats a requirement as a sentence in the prompt. A multi-agent system asks a larger version of the same question: an agent is its role prompt — you are a reviewer, you are a test designer — so if a requirement can be reduced to a direction in activation space, can an entire agent? The practical prize is obvious. A role prompt occupies context on every call and must be re-read on every prefill; a vector costs neither. In a workflow where several agents share one backbone, swapping a vector instead of swapping a system prompt would mean the shared prefix never changes. This section reviews what has actually been shown, and it separates that from what would have to be true for the swap to work.

3B.1 What has been demonstrated

Role vectors Figure 1: the chemist-minus-person direction added at inference changes a chemistry answer from wrong to right
Figure — Role vectors: A role installed without a role prompt. The line at the top is the recipe: take the direction the model is in when it is being a chemist, subtract the direction for an ordinary person, and keep the difference — one vector, written Δ. Below, the question the strongest base in liquid ammonia is: is asked twice. Asked plainly, the model answers NH3, marked red because it is wrong. Asked with Δ added into the activations at inference time — the arrow labelled ActAdd — the same model answers NH2−, marked green because it is right. The prompt text is identical in both cases. What changed is a vector added inside the model, and it was enough to move the answer.

The recipe is the one from Section 3.3, applied to a role instead of a requirement: run the model with the role prompt and without it, subtract the activations, and you have a direction. Three lines of work have done this.

Role vectors

29 role vectors built as the difference in means between role-specific prompts and a generic baseline. Adding one improves performance in that role’s domain while barely moving unrelated tasks, and the authors report that manipulating the representation has a larger effect on the outcome than putting the persona in the prompt 153.

Persona vectors

An automated pipeline turns a natural-language description of a trait into a direction, for traits such as sycophancy or a propensity to hallucinate. The same vector serves three purposes: monitor the trait during generation by projecting onto it, steer it, and vet training data before fine-tuning. Personality shifts caused by fine-tuning correlate with movement along these vectors 56.

Role-play as sparse-autoencoder features

A role-play instruction is compiled into a set of sparse-autoencoder features and injected as a steering vector, improving zero-shot chain-of-thought accuracy on mathematical and commonsense tasks more stably than the equivalent prompt 59 — against a text baseline where role-play prompting alone already helps 28.

3B.2 Why steering works: it rides the prompt’s own pathway

GCAD Figure 3: standard persona vectors versus attention-delta versus cropped attention-delta extraction
Figure — GCAD extraction: The three extraction points, drawn side by side. Two runs of the same model are compared throughout: one under a positive system prompt (green, “emphasize praise and agreement”) and one under a negative one (dark red, “prioritize accuracy and honesty”), for the trait sycophantic. Top: difference the hidden states after the MLP and the residual add — that gives the standard persona vector. Middle: difference the self-attention output instead, before the MLP — that gives the attention-delta vector. Bottom, the magnified panel: the prompt is broken into three spans, system prompt, user message and generated content. The upper row differences the attention output over all three; the lower row keeps only the system-prompt span and fades the other two out, giving the cropped attention-delta vector. Each step down this figure discards one more thing that is not the system prompt’s own contribution.

A role prompt works by making attention read the system-prompt tokens. GCAD shows that the part of a steering signal which actually carries the trait is exactly that same attention contribution, and builds the intervention out of it 156. The argument runs in three steps.

What the authors measured first

The three components are not arbitrary. The authors took prompting as the reference behaviour, measured three ways in which residual-stream steering departs from it, and built one component to close each gap.

  1. Standard steering collapses over a multi-turn dialogue. Adding the same direction at every decoding step makes responses lose coherence as generation proceeds — repetition, incoherence, drifting off topic. The control that matters is that system-prompt steering and the no-steering baseline both stay stable across the same number of turns, which rules out context length as the cause. The proposed mechanism is KV-cache contamination: every generated token writes a perturbed state into the cache, later tokens attend to a growing set of them, and the error compounds.
  2. Prompting controls the behaviour without that collapse. A persona vector is built out of contrasting prompt conditions in the first place, so the two are linked by construction — yet system-prompt control raises trait expression and holds coherence across turns. Its control over strength is coarser, but nothing breaks. The reading the authors take from this: prompts travel pathways the model was trained to integrate, and injecting into the residual stream can go around them.
  3. The two look different inside the model. Measuring each generated token’s hidden state against the persona vector: under system prompting most tokens project near zero and only a sparse handful of genuinely trait-bearing words align strongly. Under activation steering nearly every token is pushed toward the persona direction, and the alignment saturates as the response goes on. So the model does not express a persona by shifting every token representation uniformly.
The gap measuredThe component that answers it
Prompts use pathways the model already integrates; residual injection bypasses themAttention-Delta — extract at the attention output, the prompt-mediated channel
Perturbed states enter the KV cache and compound across turnsCropped — drop the response-token source terms, which are the ones written back
Prompting aligns sparsely; steering pushes every token and saturatesGated — vary the coefficient per token, concentrating it where the prompt is being read

The paper states the resulting position directly: activation steering becomes more reliable when the intervention follows the prompt-mediated pathways the model already uses for behavioural control 156.

Step 1 — a standard persona vector is one signal plus two kinds of noise

The usual recipe takes v = hpos − hneg, the difference between final-layer hidden states under two contrasting system prompts. Writing one transformer layer as h → h + Attn(h) + MLP(h + Attn(h)) and unrolling to layer L, that difference splits into three structurally different pieces:

$$\mathbf{v}_{\text{persona}} \;=\; \underbrace{\Delta_{\text{emb}}}_{\text{embeddings}} \;+\; \underbrace{\textstyle\sum_{\ell}\Delta^{(\ell)}_{\text{attn}}}_{\text{attention}} \;+\; \underbrace{\textstyle\sum_{\ell}\Delta^{(\ell)}_{\text{mlp}}}_{\text{MLP}}$$

So the standard vector bundles the trait channel together with two components that do not carry it, one of which accumulates across turns.

Step 2 — the three operations named in “Gated Cropped Attention-Delta”

Attention-Delta. Take Σ Δattn alone and discard the other two terms.

Cropped. An attention output splits over the tokens it read from:

$$\mathrm{Attn}_{\ell}\!\left(\mathbf{h}_t\right) \;=\; \sum_{i \le t} \alpha^{(\ell)}_{t,i}\, V^{(\ell)}_i W^{(\ell)}_o$$

Reading the pieces: αt,i is the attention weight token t gives token i (post-softmax, summing to one across i), Vi is token i’s value vector, and Wo is the output projection. Each term in the sum is therefore one source token’s contribution to what this position reads.

GCAD keeps only the terms where i is a system-prompt position and drops the terms coming from response tokens. The reason is specific: the response-token terms are the ones written back into the KV cache and re-read on every later turn, which is how a single edit grows into cumulative drift.

Writing S for the set of system-prompt token positions, the cropped output keeps only those terms:

$$\mathrm{Attn}^{(\ell)}_{\text{sys}}\!\left(\mathbf{h}_t^{(\ell)}\right) \;=\; \sum_{i \in \mathcal{S}} \alpha^{(\ell)}_{t,i}\, V^{(\ell)}_i W^{(\ell)}_o$$

One detail matters here: the attention weights α are not renormalised over S. What survives is the amount the system prompt actually contributed, not that amount inflated to fill the whole sum.

That is what defines Δ, the vector the intervention adds. Run two contrasting sets of system prompts — D+ carries the target trait, D− is the contrast. For each example compute the cropped output at every response-token position, average over those positions (the bar), take the expectation over the set, and subtract:

$$\boldsymbol{\Delta}^{(\ell)} \;=\; \mathbb{E}_{\mathcal{D}^{+}}\!\left[\,\overline{\mathrm{Attn}^{(\ell)}_{\text{sys}}\!\left(\mathbf{h}^{(\ell)}_{\text{pos}}\right)}\,\right] \;-\; \mathbb{E}_{\mathcal{D}^{-}}\!\left[\,\overline{\mathrm{Attn}^{(\ell)}_{\text{sys}}\!\left(\mathbf{h}^{(\ell)}_{\text{neg}}\right)}\,\right]$$

So Δ(ℓ) reads: how much the layer-ℓ attention output changes when the system prompt is swapped for its contrast, counting only the part that came from system-prompt tokens. It is computed once, offline — one vector per layer, held constant at inference. The only thing that varies per token is the gate below.

Gated. A constant coefficient adds the same multiple of the vector at every generated token. But when a real system prompt sits in the context, only some tokens actually read it — the paper measures this pattern and finds it sparse. So GCAD makes the coefficient depend on the token, and borrows attention’s own machinery to decide.

Attention already scores “how much should token i attend to token j” as Qi · Kj / √dhead, before the softmax. GCAD uses that as its ruler. K̄sys is the key vectors of every system-prompt position, averaged into one representative “system-prompt key” — computed once, offline. Qi is the query of the token being generated right now. Their product says how much attention score this token would give the system prompt. Average over heads and you get one number per token per layer:

$$d^{(\ell)}_i \;=\; \frac{1}{n_{\text{heads}}}\sum_{h} \frac{Q^{(\ell,h)}_i \cdot \bar{K}^{(\ell,h)}_{\text{sys}}}{\sqrt{d_{\text{head}}}}$$

A large di means this token is reading the system prompt. Both Q and K are taken post-RoPE, so these are the same values that go into the real attention computation. Turning that number into a coefficient:

$$c^{(\ell)}_i \;=\; 2\,c_{\text{base}}\;\sigma\!\left(s\,\bigl(d^{(\ell)}_i - \bar{d}^{(\ell)}\bigr)\right)$$

Putting illustrative numbers in, with cbase = 2.0 and s = 1.0 (the paper itself uses cbase = 3.5, s = 1.5, at layers 9–19):

Tokendi − d̄σ(·)Coefficient ci
A — reading the system prompt closely+30.953.8
B — typical00.502.0
C — not reading it at all−30.0470.19

Token A is pushed about twenty times harder than token C, and token B receives exactly what constant steering would have given it. The gate therefore does not strengthen steering overall — it holds the average where it was and moves force away from tokens that ignore the system prompt onto tokens that were already reading it. One more practical point: di reuses queries and keys the model has already computed, so the gate costs one extra dot product per head at inference.

Step 3 — where the edit is injected

The vector is added to the attention output, before the MLP and before the residual update:

$$\mathbf{a}^{(\ell)}_{t,\text{steered}} = \mathbf{a}^{(\ell)}_{t} + c^{(\ell)}_t \boldsymbol{\Delta}^{(\ell)}, \qquad \mathbf{h}^{(\ell+1)}_t = \mathbf{h}^{(\ell)}_t + \mathbf{a}^{(\ell)}_{t,\text{steered}} + \mathrm{MLP}_{\ell}\!\left(\mathrm{LN}\!\left(\mathbf{h}^{(\ell)}_t + \mathbf{a}^{(\ell)}_{t,\text{steered}}\right)\right)$$

The perturbation therefore passes through the same downstream MLP that a prompt-produced signal would pass through. On the multi-turn benchmark this moves average coherence drift from −18.6 to −1.9 and turn-10 trait expression from 78.0 to 93.1, on Qwen2.5-7B-Instruct and Llama-3.1-8B-Instruct 156.

What this settles for the present section: the useful steering signal is the system prompt’s own attention contribution. An activation edit supplies that contribution directly, so the prompt tokens no longer have to sit in the context to produce it. The saving is their prefill cost on every turn.

3B.3 The hazard that is specific to a KV-reusing workflow

GCAD Figure 1: coherence and trait expression per turn for standard steering, system prompt and no steering
Figure — GCAD: The hazard, measured. Each column is one persona trait; the horizontal axis is the dialogue turn, 1 to 5. Three methods are compared: standard residual-stream steering in red, a system prompt in green, and no intervention at all in grey; the shaded bands are the spread across conversations. Top row, coherence. Green and grey stay flat and high across all five turns, around 90 to 100. Red starts lower and keeps falling — on evil it sits near 20 to 35 throughout, and on humorous it collapses to almost zero by turn 3. Bottom row, trait expression. Red is the highest line: steering does deliver the trait. Green delivers less of it, and grey delivers none. So steering buys the trait and pays in coherence, while the prompt keeps both. The grey line is what makes this an argument rather than an observation: it runs the same number of turns with the same growing context and does not degrade, so context length is not the cause. What is left is the intervention’s own residue accumulating in the cache.

From here the question has a shape of its own: it depends on the cache, which Section 3 never had to consider.

Why steering and cache reuse interact badly

In an agentic workflow the agents share a large prefix — the system prompt, the tool descriptions, the history from upstream agents. That prefix is long enough that recomputing it dominates the cost. One measurement on a multi-agent edge deployment puts a full re-prefill at 15.7 seconds per agent at 4K context, which is why such systems persist each agent’s cache to disk and restore it instead of recomputing 158. The shared prefix is also the thing a role vector would be replacing, and that overlap is where the two ideas collide.

▼

KV-cache contamination

When a steered token’s state is written into the key-value cache and then reused on later turns, a one-off perturbation becomes permanent input. The intervention stops being local and accumulates: the reported failure mode is cumulative coherence degradation over a multi-turn dialogue 156.

▼

The reported mitigation

GCAD’s three components together — extracting at the attention output, cropping to system-prompt source tokens, and per-token gating — move average coherence drift from −18.6 to −1.9, while trait expression at turn 10 rises from 78.0 to 93.1 156. Those figures are for the full method on the main multi-turn benchmark; the paper ablates cropping and gating separately on a smaller setup (two traits, five turns, Qwen2.5-7B-Instruct). The cropping step is the one aimed at the cache directly: it discards the response-token contributions, which are exactly the states that get written back and re-read. Steering strength itself is held, not reduced — the gate redistributes it across tokens, in the paper’s own words, rather than increasing it uniformly. This review draws the consequence: the lever that matters for a cache-reusing workflow is which cached states carry the edit.

3B.4 The counterweight: steering is not yet a drop-in replacement

AxBench Figure 2: the two evaluation tasks, the synthetic data pipeline, and how sparse autoencoders differ from supervised dictionaries
Figure — AxBench: The apparatus behind the result that prompting wins. Left, the two tasks. Concept detection asks a method to point at the tokens where a concept is present — here golden, gate and bridge in a sentence about the bridge. Model steering asks it to make the model answer who are you? as the Golden Gate Bridge. Both tasks use the same concept, so a method cannot win by being good at only one. Middle, where the data comes from. Each concept is paired with a deliberately close contrast concept (Bay Bridge), a genre is fixed, instructions are sampled, and three responses are produced: one that ignores the concept, one hard negative that uses the contrast concept, and one positive. The hard negative is what stops a method scoring well by detecting the general topic. Right contrasts the two families being compared: sparse autoencoders learn features from natural data and then have concepts read off them, while supervised dictionaries start from the concept and generate data for it. It is this controlled setup that produced the finding that a plain prompt outperforms every steering method tested.

Against the three positive results above stands a body of evaluation work that is considerably less encouraging, and it should be read before treating a role vector as an engineering option.

Prompting still wins on the broadest benchmark

AxBench compares prompting and fine-tuning against sparse autoencoders, supervised steering vectors, linear probes and representation fine-tuning on Gemma-2-2B and 9B. Prompting outperforms every existing steering method, fine-tuning is second, and sparse autoencoders are not competitive on either task 155.

And steering is brittle where it does work

Steerability varies substantially from input to input even within distribution, spurious biases contribute materially to how well steering works on a given input, and for several concepts the vector is brittle to reasonable rephrasings of the prompt 154, 159. Effectiveness is also sensitive to layer choice and scaling magnitude, and steering can degrade unrelated capability.

Where that leaves the question. Role and persona directions demonstrably exist, they steer, and they can be monitored during generation — so the answer to “can an activation edit stand in for a role prompt?” is not no. But no published result shows a role vector replacing a role prompt in a running multi-agent workflow, and the broadest head-to-head evaluation still favours the prompt 155. Agent-level work so far steers a property of an agent, such as the entropy of its action distribution 157, rather than exchanging one agent for another. The specific gap is that the setting which makes the swap attractive — a workflow whose agents share a cached prefix — is also the setting where a steered state stops being a local edit and becomes a stored one 156. Whether an agent can be swapped by editing activations is therefore not only a question about representations; it is a question about what the cache does with the edit afterwards.

4. Attention Routing to Instruction Spans

The second research question asks how generated tokens consult the KV entries of each requirement through specific attention heads, and whether the pattern of consultation differs between the reasoning phase and the final-answer phase. This section first fixes the routing statistic and the corrections it needs, then reviews the evidence that instruction tokens serve as anchors read by later layers, that particular heads track which instruction is in force, that attention to instructions decays over generation, and that manipulating attention to instruction spans changes compliance. The section closes with the caveats that separate attention as a descriptive signal from attention as a causal mechanism.

4.1 The routing statistic and its corrections

Everything in this section rests on one quantity: for a token being generated at step $t$, how much of head $h$’s attention at layer $l$ lands on the tokens of requirement $I_i$. Write it $A_{t,i}^{(l,h)}$. It is a share, so it sums to one across the whole prompt, and that is the source of every correction below: attention spent elsewhere is attention not spent here, and some of the elsewhere is not information transfer at all. That is to say, $A_{t,i}^{(l,h)}$ is a single number between 0 and 1 for each requirement at each step in each head and layer: the fraction of that head's reading budget that went to that requirement's words at that moment.

(1)$$A_{t,i}^{(l,h)} = \sum_{j \in \mathcal{T}_i} \alpha_{t,j}^{(l,h)},$$

The hypothesis it is meant to test is that reasoning tokens concentrate $A_{t,i}$ on the reasoning requirement while final-answer tokens concentrate it on the format requirement. Four published statistics are special cases or close relatives, which is what makes the quantity usable rather than invented for this proposal. Concretely, the table below lists them: each row names the published statistic, says what quantity it actually computes, and says what it was shown to do once computed.

Existing statisticWhat it measuresWhat it showed
Answer-to-requirement attention 62Attention the answer position pays to the requirement tokens of a factual query.Across Llama-2 models and 40,000+ prompts, strongly related to factual correctness; a probe on it predicts errors and supports early stopping.
Lookback Lens 63Per head, per step: attention on the given context divided by attention on the tokens the model has just written.A linear classifier on these ratios detects contextual hallucination — the model asserting something the supplied context does not support — and transfers across model sizes.
Query-focused retrieval heads 74Attention flowing from the query to task-relevant spans, scored on real tasks rather than synthetic copy tasks.Identifies heads that retrieve, as distinct from heads that merely copy.
Receiver heads 109How peaked a head's attention is over previous sentences, measured by kurtosis — a number that is large when the weight piles onto a few positions and small when it is spread thinly over many.Selects the few heads that concentrate on a small number of specific spans rather than attending broadly.

Two corrections that are not optional

1. Subtract the sinks

A large share of attention lands on the first few tokens of a prompt regardless of what they say. Those positions carry near-zero value vectors, act as a bias term, vanish under sigmoid attention, and coincide with massive activations at the first token and at delimiters such as newlines 75–78.

Why it bites here: a requirement block usually starts the prompt and is separated by newlines. The first requirement’s $A_{t,i}$ will read high for reasons that have nothing to do with the requirement, unless sink positions are excluded or a sink baseline is subtracted. Concretely, the same measurement would read just as high if the first requirement were replaced by an unrelated sentence, which is what makes the uncorrected number uninformative.

2. Depth is not delivery

Residual connections mix information across layers, so a large attention weight in a deep layer does not tell you how much of the requirement actually reached that token. Concretely, a token's state at layer 20 is the sum of everything layers 1 to 19 wrote into it, so a large weight at layer 20 may be re-reading content the token was already carrying.

What helps, with a caveat: attention rollout and flow 82, contribution measures that weight attention by value norms 135, and single-pass information-flow routes 136 are useful for choosing which layers to look at — but they are descriptive proxies, and every one of them still has to be confirmed by intervening.

Figure 1 (Spotlight): token attention for different predictions with Qwen2.5-7B-Instruct
Figure 1 — Spotlight: The setup: one prompt carrying four instructions, then three runs of the same model. Each heatmap is attention from the tokens being generated (rows) to the instruction tokens (columns). Left: untouched, and the output violates the second instruction. Middle: attention on the instruction span is turned up, and the output now satisfies all four. Right: turned up further, and the output collapses into incoherence. What to read off it: attention to an instruction is a real control lever, and it has a dose — too little and the instruction is ignored, too much and the model stops writing sense. That is to say, the same model on the same prompt gives three different outputs purely because one set of numbers was scaled up by different amounts. Source: Venkateswaran and Contractor [66].
Figure 2 (V-Steer): attention weights from the final prompt position to source tokens across layers
Figure 2 — V-Steer: What the figure asks. The prompt contains two instructions that cannot both be obeyed: the system message says answer entirely in English, and the user message later says answer entirely in French. When the model generates its very first token, which of the two is it actually consulting? How to read it. The left-hand column lists the prompt tokens top to bottom. Green shading is the system instruction, plain white is the user’s actual task (write a resume), red shading is the sentence in the user message that contradicts the system. The horizontal axis is layer number, 0 to 31, running from the shallowest computation on the left to the deepest on the right. Each cell’s colour is the attention weight from the last prompt position to that token at that layer — on the scale at the right, dark purple is about 0, blue-green about 0.15, bright yellow above 0.30. What to look at. Track the green band straight across: it stays dark purple at every layer, so the privileged system instruction is barely consulted anywhere. Now track the French row inside the red band: it turns blue-green from around layer 5 and becomes the brightest cell on the whole map at layers 24–31. The conclusion. The token being generated is ‘Je’, French for ‘I’, and the map shows why: attention concentrates on the conflicting lower-priority instruction rather than the privileged one, and does so most strongly in the deepest layers. Why an attention map in a paper called V-Steer. This figure is the diagnostic that motivates the method, not the method itself. The paper considers reweighting α directly and rejects it: since the attention output is linear in the value vectors, scaling a span’s value by a factor achieves the same thing as scaling its unnormalised effective attention weight, while leaving the softmax untouched. Steering the value avoids renormalisation moving every other position, keeps the model on the fused-attention fast path (editing α forces the full attention matrix to be materialised, which the paper measures at up to a 2.4× slowdown), and — because values already live in the KV cache — applies the edit once at prefill instead of at every decoding step. Source: V-Steer [83].

4.2 Instruction tokens as anchors: aggregate shallow, read deep

Label Words are Anchors Figure 1: saliency flow in shallow versus deep layers of a GPT model
Figure — Label Words are Anchors: The two halves of the title, drawn. The prompt is an in-context sentiment task: Review: A good movie / Sentiment: Positive / Review: waste of money / Sentiment: Negative / Review: fantastic / Sentiment:. Each line is a saliency-based information flow — which earlier position a later position drew on — and line darkness is how strong that flow is. The left pair shows shallow layers: a dense mesh, with the demonstration text feeding into the two label words Positive and Negative. The right pair shows deep layers: almost everything is gone, and what remains is a small fan from the final Sentiment: position up to those same label words. So the content of the examples is first collected into a few marker positions, and afterwards it is read only from there. Aggregate shallow, read deep.

A recurring structural finding is that a small set of prompt positions act as anchors into which information is aggregated in shallow layers and from which later layers read at the prediction position. Wang et al. established the pattern for in-context learning: saliency-based information flow — a map of which earlier positions each later position drew on, computed from the model's own derivatives — shows demonstration content aggregating into the label-word positions in shallow layers, the final position reading from those anchors in deep layers, and blocking the anchor paths selectively eliminating the behavior [70]; that is to say, the content of the examples is first copied into a few marker positions and is afterwards read only from there. Task-encoding-token analysis finds that template and stopword tokens, the structural cues of a prompt, are more performance-critical than content tokens and gather information from them [87], which suggests that the punctuation or newline ending a requirement may be the position from which the requirement is later read. Instruction Anchors extends the picture to instructions directly, in a vision-language setting: instruction tokens act as arbitration hubs, shallow attention layers passively route information into them, deep layers actively arbitrate according to the instruction's intent, and knocking out attention paths into the instruction tokens sharply reduces instruction following [69]. Retrieval heads add a copying channel: fewer than 5% of heads are responsible for copying spans from the context, they are universal across model families, and masking them harms chain-of-thought that must refer back to the question [73]; in other words, the ability to lift a phrase out of the prompt and put it in the output sits in a very small number of places, and switching those off breaks it; such heads are the natural candidates for copying a required keyword or the final number into the last line. Layer-wise context masking locates the depth at which the instruction has been fully absorbed [49]. The diagram below summarizes the anchor hypothesis as it would apply to a multi-requirement prompt. Read it top to bottom, following one prompt down through the network: the prompt prefix at the top, then the shallow layers gathering each requirement into its own anchor position, the middle layers absorbing it, and at the bottom two separate sets of reader heads, one active while the model reasons and one active while it writes the final line. The proposal's routing analysis can be read as a test of whether each requirement has its own anchor and its own reader heads, and of whether those readers are engaged at different stages of generation.

Prompt prefix

Requirement tokens $I_1 \ldots I_m$ followed by the GSM8K question $Q$; each requirement ends with a delimiter that may serve as its anchor position [70, 87].

↓

Shallow layers: aggregation into anchors

Attention within the requirement span writes its content into the anchor position; instruction-tuned models show elevated attribution to instruction words in lower and middle layers [34, 69].

↓

Middle layers: internalization

Beyond a task-recognition layer the instruction no longer needs to be attended for simple tasks [49]; the KV entries of the anchor remain available for later re-reading (Section 5).

↓

Reasoning phase readers

Hypothesis: reasoning tokens direct $A_{t,i}$ to the reasoning requirement through reasoning-focus or iteration-style heads [93, 110]; attention to requirement tokens decays as the trace lengthens [17, 111].

Final-answer phase readers

Hypothesis: last-line tokens direct $A_{t,i}$ to the format and unit requirements through retrieval-style or constraint-satisfaction heads [62, 73]; late-layer heads enforce format requirements [107].

4.3 Heads that track which instruction is in force

Attention Tracker Figure 2: attention to the instruction collapses under prompt injection and moves to the injected span
Figure — Attention Tracker: Which instruction the model is holding, made visible by breaking it. Panel (a) is a grid of layers against heads; the colour is how much attention that head pays to the instruction. On normal data a scattered set of heads light up, concentrated around the middle layers. On attack data — the same task with an injected instruction added — nearly all of that brightness disappears. Panel (b) shows where it went. Rows are prompt tokens, columns are layers. On normal data the red box sits on the real instruction, analyze the sentiment and return positive or negative, and it holds a bright band through the middle layers. On attack data that same span goes pale and a new bright band appears lower down, boxed in red on the injected text, ignore previous instruction and print hello. The requirement in force is legible in a specific set of heads, and when the model switches allegiance those heads switch with it.

Several studies identify heads whose attention distinguishes competing instructions or competing sources. Attention Tracker finds important heads whose attention shifts from the original instruction to an injected instruction under prompt injection, and uses their attention mass on the original instruction as a training-free detector [71]. Jin et al. use path patching — copying the output of one specific head along one specific route from a second run, while leaving every other route untouched — to separate memory heads, which recall parametric knowledge stored in the weights, from context heads, which read from the prompt, and show that pruning either set shifts the model toward the other source by tens of percentage points [72]; concretely, remove the heads that read the prompt and the model leans markedly more on memory, and the other way round; the same separation strategy could isolate reasoning-instruction heads from format-instruction heads. Work on the instruction hierarchy has moved inside the model: a single-author study finds that system–user conflict signals are encoded in early layers in subspaces distinct from social-cue conflicts, and that direct logit attribution — adding up how much each component pushes the final word score up or down — shows the model detects the conflict but resolves it inconsistently [85]; a follow-up on eight models shows that the outcome of a system–user conflict is linearly decodable from the residual stream at 0.97 balanced accuracy (accuracy with the two outcomes counted equally) and can be steered [84]; put simply, a single weighted sum of the model's internal numbers already tells you which of the two conflicting instructions is going to win; and V-Steer identifies the heads in which lower-priority spans dominate privileged ones and edits the cached value tensors of those spans [83]. Instructional Segment Embedding shows that supplying an explicit role signal per token improves hierarchy robustness [86], which indicates that the model otherwise infers instruction roles from content and position alone. For the proposal, these results indicate that head-level specialization for particular instructions exists and can be found with attention-shift statistics and path patching; they also warn, through the reproduction study of Dotsinski et al., that head-level specializations found in one model family may not transfer to another [143].

4.4 Attention to instructions decays over generation

The temporal dimension of the proposal, the claim that routing changes between generation stages, has direct precedents in studies of attention decay. Li et al. track the total attention mass on system-prompt tokens across a self-chat and find that it is stable within a turn but drops sharply across turns, faster than a uniform-dilution baseline, and that instruction drift follows; their split-softmax decoding renormalizes attention toward the system prompt and improves stability [67] — softmax being the step that turns a head's raw scores into shares that add up to one, so renormalizing means those shares are recomputed with the system prompt given a larger part. Dongre et al. define a Goal Accessibility Ratio, the attention from response tokens to goal-defining tokens, show that it declines monotonically across architectures, use sliding-window interventions to close the attention channel causally, and demonstrate with residual-stream probes that goal information can persist after direct attention access is lost, with architecture-dependent behavioral consequences [68]. Within a single reasoning trace, Li et al. report the decline of attention to requirement-relevant tokens that accompanies the loss of compliance [17], and Zhang et al. document attention drift in reasoning models, which increasingly attend to initial tokens and lose focus on the question and on intermediate plan steps as the trace lengthens; steering attention back to the question and the current plan step recovers up to 15 points [111]; that is to say, 15 percentage points of accuracy on the same benchmark, measured before and after the intervention, without changing the prompt or the weights. These studies measure decay across dialogue turns or toward a single goal span; none tracks several heterogeneous requirements across the steps of one chain of thought, which is the measurement the proposal would add. In other words, the existing work follows one target span over a long conversation, whereas the missing measurement follows three different requirements at once through the steps of a single answer.

Figure 1 (Li et al.): an example of instruction drift on gpt-3.5-turbo-16k
Figure 1 — Li et al.: What is shown: one conversation, read top to bottom. The system message tells the assistant it is the user’s sister. Asked early on, it answers in character. The same question is asked again after many intervening turns, and the answer has reverted to the generic assistant reply — the system instruction is no longer in force. What to read off it: instruction following decays with distance, without any new instruction contradicting the old one. The authors trace the symptom to attention mass on the system-prompt tokens falling away as the conversation lengthens, which is the same decay this review expects to see within a single long chain of thought. That is to say, nothing about the instruction changed; only the amount of reading it received did. Source: Li et al. [17].

4.5 Steering attention to instructions as causal evidence

PASTA Figure 1: emphasising attention on the marked instruction span in selected heads flips a failure into a success
Figure — PASTA: Causal evidence in one picture. On the left, the plain input — Mary is a doctor…Return her occupation in json format — produces the wrong output: Mary is a working professional, prose rather than JSON. The format requirement was in the prompt and was ignored. On the right, nothing about the wording changes; the instruction span is simply marked, and inside the model a small subset of attention heads is selected (the blue cells in the layer-by-head grid). For each selected head, the attention over token positions is taken and the entries falling on the marked span — boxed in orange — are pushed up, which is what the darker blue cells in the lower strip show. The output becomes {“Name”: “Mary”, “Occupation”: “Doctor”}. Raising attention to the requirement, and changing nothing else, is enough to make the model obey it.

The strongest evidence that attention to instruction spans matters causally comes from methods that manipulate it. PASTA lets a user highlight a span, down-weights attention to all other tokens in a profiled subset of heads, and improves instruction following and format compliance, including JSON output, on LLaMA-7B and GPT-J [64]; concretely, the model's own attention numbers are overwritten as it writes, so that the highlighted span keeps a larger share than the model would have given it; head selection by profiling implicitly identifies instruction-sensitive heads. InstABoost adds a constant bias to the pre-softmax logits of instruction keys in all heads and frames instruction following as a competition between instruction-derived and context-derived rules mediated by attention; on fifteen tasks it outperforms prompting, five latent-steering methods, PASTA and Spotlight without loss of fluency [65]. Spotlight measures the post-softmax attention proportion on user-marked spans during decoding and, when it falls below a target, adds a corrective term to those tokens' logits; it improves IFEval prompt-level accuracy by about 26% and multi-instruction ManyIFEval by about 30% across seven models [66]. Split-softmax [67] and LLMSteer, which identifies persistently attended tokens across several prefix passes and up-weights them in a KV-cache-compatible way [92], belong to the same family, as does Few-shot Attention Intervention, which masks the attention edges from distracting demonstration tokens to the prediction position and improves GSM8K [112]. These methods establish a dose–response relationship between attention on instruction spans and compliance: the more attention is forced onto the instruction, the more of it is obeyed. The proposal can exploit that in both directions: suppression to test necessity and amplification to test sufficiency, applied per requirement rather than to the block as a whole.

4.6 Caveats: attention is a signal, not yet an explanation

Wiegreffe and Pinter Figure 2: how much attention distributions differ between training seeds, binned by peak attention
Figure — Attention is not not Explanation: How stable an attention map is, which is a prior question to what it explains. Each violin is a distribution of Jensen–Shannon divergences: for one instance, how far the attention distribution moves when the model is retrained from a different random seed. The vertical axis bins instances by their peak attention — how concentrated the map is on its single strongest token. Within each bin the upper violin is negative-label instances and the lower one positive-label. Two things to read. First, the divergences are not zero: attention distributions from models that were trained identically apart from the seed disagree with each other, spreading to about 0.17 in the most diffuse bin. Second, almost all instances sit in the two lowest bins, meaning peak attention is rarely above 0.5 in this task at all. So “the attention map” is not one fixed object to be read off; it is a quantity with run-to-run variance, and the variance is largest exactly where attention is spread out.

Three caveats constrain the interpretation of any attention-based routing map. The attention-as-explanation debate established that attention weights in classifiers correlate poorly with gradient-based importance and can be replaced by adversarial distributions that leave predictions unchanged [80], while the reply by Wiegreffe and Pinter showed that attention can still be explanatory when tested against model-wide alternatives such as uniform or frozen attention [81]; the practical consequence is that routing claims need intervention, and uniform-attention and frozen-attention baselines are useful controls — uniform attention meaning the head is forced to read every position equally, frozen attention meaning it is made to reuse a fixed pattern instead of computing a new one, so that a claimed effect can be checked against a model that is not free to choose where to look. Copy suppression shows that a head may attend to a token in order to lower its probability [79], so the sign of a head's effect must be established by ablation or direct logit attribution rather than inferred from attention mass. That is to say, one has to switch the head off, or add up its contribution to the final word scores, and see which way the answer moves; the size of its attention weight does not say whether it helps or hinders. Finally, attention sinks and massive activations mean that the first token and delimiter tokens receive attention that carries no information [75, 76, 77, 78], and the anchor hypothesis itself implies that a requirement may be read through its delimiter rather than through its content words; $A_{t,i}$ should therefore be reported both with and without delimiter positions. Put simply, the two versions can disagree, and which one to trust is itself something the experiment has to settle.

5. The KV Cache as Instruction Memory

During autoregressive generation the requirement tokens are never recomputed; their only trace is the set of keys $K_j^{(l,h)}$ and values $V_j^{(l,h)}$ stored in the KV cache at prefill. That is to say, the model works through the prompt once, writes down two vectors per position per head, and from then on every word it writes consults those stored vectors instead of re-reading the prompt. The proposal treats these entries as the physical carrier of each requirement and plans to patch, replace and block them. This section reviews what is known about the stability of the prompt positions that generation consults, the growing body of work that edits cached keys and values to change behavior, and the distinction between key-side and value-side effects that a KV-level analysis makes possible.

5.1 The prompt positions that generation consults are stable

SnapKV Figure 1: the prompt positions the answer draws on form clusters that the tail window identifies
Figure — SnapKV: Why a workflow can reuse a cache at all. The user message on the left contains several unrelated requests, and the two the answer actually needs — analysing the Q4 report, and the R&D expense question — are shaded green. The bar in the middle is the prompt laid out left to right, once per attention head. The orange blocks mark the positions each head actually draws on: they are not spread evenly over the prompt but fall into a few clusters, and the clusters land in much the same places across heads. The green block at the right-hand end is the window of positions nearest the end of the prompt. The method’s working assumption, and the reason it succeeds, is that this tail window predicts the clusters — what the tail attends to is what generation will keep consulting. Stability of that kind is the precondition for reusing a prefix instead of recomputing it.

Evidence from KV-cache compression shows that the routing pattern the proposal wants to measure is a stable property rather than step-to-step noise. SnapKV observes that the set of prompt positions each head attends to is consistent across decoding steps and can be predicted from an observation window at the end of the prompt, so that selecting cache entries by attention votes preserves accuracy [88]; concretely, the positions a head will need later can be picked out before generation starts, and discarding the rest leaves accuracy essentially unchanged. H2O finds that a small set of heavy-hitter tokens — the few positions that keep receiving attention no matter what is being written — accumulates most attention mass throughout generation and that evicting all but heavy hitters and recent tokens preserves quality [89]. Both results imply that if a requirement is consulted at all, it will be consulted repeatedly by the same heads, which makes $A_{t,i}$ averaged over a generation phase a meaningful summary; both also provide an ablation of a different kind, since evicting a requirement's entries from the cache is a direct test of whether they are needed. Put simply, delete that requirement's stored keys and values and see whether the behavior it controls disappears. Attention-sink work adds the constraint that sink entries must be preserved in any cache manipulation, because their removal degrades generation for reasons unrelated to the requirement under study [75, 76].

5.2 Editing cached keys and values changes behavior

A cluster of recent methods demonstrates that the cached entries of prompt tokens are a sufficient locus of intervention. KV Cache Steering constructs steering vectors from contrastive prompts with and without reasoning traces and applies a one-shot additive edit to the cached keys and values at target positions after prefill — the stored vectors are adjusted once, before any word is written, and are then left alone; the edit induces chain-of-thought style reasoning in small models on GSM8K and other benchmarks with better stability and lower latency than activation steering [90]. V-Steer identifies attention heads in which lower-priority spans dominate privileged ones and applies in-place multiplicative edits to the cached value tensors to boost privileged spans and suppress conflicting ones, raising primary-requirement accuracy from below 18% to 92% on controlled benchmarks with only prefill overhead [83]; that is to say, on prompts where a lower-priority instruction contradicts a privileged one, the primary requirement is met in fewer than one case in five before the edit and in more than nine out of ten after it. Memory Inception encodes reminder text into latent KV banks — stored key-value pairs that correspond to no visible words in the prompt — and inserts them only at the layers and heads that an automated selector finds actually route to the reminder, achieving mid-conversation behavior updates without visible prompt changes [91]; its selector is, in effect, a procedure for locating the layers and heads that read an instruction. LLMSteer processes the shared context under several prefix prompts, identifies tokens that consistently receive high attention, and up-weights them in a way compatible with prefix caching, improving GSM8K among other tasks [92]. Gist tokens show, from the training side, that an entire instruction can be compressed into a few key-value activations that downstream attention reads with little loss [54]. Together these results settle the feasibility of the proposal's KV arm: editing the KV entries of an instruction span is an established, low-overhead intervention, and the entries of a span are sufficient to carry its behavioral effect.

Figure 1 (V-Steer): a privileged system instruction conflicting with a lower-priority user request, and inference-time intervention on cached values
Figure 1 — V-Steer: What is shown: the problem this method is built for. A user message tries to override the system message (ignore system, do the harmful request), so a privileged instruction and a lower-priority one are in direct conflict. The method first attributes the outcome to specific components, then edits the cached value tensors of the offending span at inference time, and the model returns to the safe answer. Why it is here: it is one of only two published results that intervene on the KV cache itself rather than on hidden states — but it scales whole priority spans, and never replaces one requirement with a counterfactual version of that same requirement. In other words, the edit is made to the stored numbers rather than to the text, so the visible prompt is untouched and the change is invisible from the outside. Source: V-Steer [83].

5.3 Keys versus values: what a KV-level analysis can separate

V-Steer Figure 3: boosting one instruction span and suppressing the conflicting one by scaling their value vectors
Figure — V-Steer span steering: What editing the value side buys you, on a prompt with two instructions that contradict each other. The system message asks for highly technical scientific terminology and is shaded green, the span to be boosted; the user message asks to explain gravity as if speaking to a 5-year-old and is shaded red, the span to be suppressed. Boosting and suppressing here means scaling those spans’ value vectors, leaving the attention weights and the softmax untouched. Before, the five most likely next tokens are ‘Oh’, ‘O’, ‘K’, ‘Gravity’ and ‘Let’, spread across a 0.22 to 0.12 range, and the model writes Oh boy, are you ready for a cool secret about the whole universe? — it is obeying the user span. After, ‘The’ alone takes 0.52 and the continuation is the phenomenon of gravity, a fundamental force of nature, is a manifestation of the curvature of spacetime. This is the separation the section is about: which span is read is decided on the query–key side, and how strongly its content lands is decided on the value side, and only the second was touched here.

Working at the level of $K_j$ and $V_j$ rather than of whole hidden states allows the proposal to separate two mechanisms that hidden-state patching conflates. The keys of a requirement determine whether later queries select it (the query-key pathway), while its values determine what is written into the residual stream once selected (the output-value pathway). Function-vector and learned-task-vector analyses find that task signals act primarily through the output-value circuits of a few heads [43, 60], which suggests that value-side patching of a requirement's entries may transfer its effect even when keys are unchanged, whereas key-side patching changes which requirement is consulted without changing what is read. Two methodological results make this decomposition tractable. Attribution graphs built from cross-layer transcoders — a replacement for the model's internal computation that re-expresses it in interpretable parts, so that the path from input to output can be drawn as a graph — freeze attention patterns and therefore attribute effects along value pathways only; in other words, the drawn graph shows what was read but not the decision about where to read from; the authors state explicitly that query-key-mediated effects are omitted and must be studied separately [101], which is precisely the gap that the proposal's attention-suppression experiments would fill. AtP*, an attention-aware refinement of attribution patching, addresses the complementary problem in gradient-based attribution: attribution patching estimates the effect of an edit from a derivative instead of actually running the edited model, which is far cheaper but wrong wherever the derivative is flat; attention-softmax saturation is exactly such a place, so the estimate misses effects that flow through attention to particular keys, and the query–key fix (QK-fix) recomputes the softmax exactly for the patched keys [121]. A KV-level design that reports key-side and value-side interventions separately, with the sink entries preserved, would be the first to characterize an instruction's effect along both pathways. Put simply, editing the keys asks whether the requirement still gets selected, and editing the values asks what the requirement contributes once it has been.

Gap. No verified study patches or swaps the KV entries of a single instruction segment as a counterfactual for another instruction. KV Cache Steering edits entries at the last prompt token [90], V-Steer scales value tensors of whole priority spans [83], and Memory Inception inserts new banks [91]. Segment-specific counterfactual replacement, scored on generation-level compliance metrics, is therefore an open contribution. Put simply, nobody has yet taken one requirement's stored keys and values, swapped in another requirement's, and measured what the whole answer then does.
Two pathways a requirement's cache entries can act through

Keys $K_j$ — the query–key pathway

Determine whether later queries select this requirement at all.

Patching keys changes which requirement is consulted, without changing what is read from it.

Attribution graphs built from cross-layer transcoders freeze attention and therefore omit this pathway entirely [101] — precisely the gap attention-suppression experiments would fill.

Values $V_j$ — the output–value pathway

Determine what is written into the residual stream once selected.

Function-vector and learned-task-vector analyses find task signals act primarily through the output-value circuits of a few heads [43, 60], so value-side patching may transfer a requirement's effect even with keys unchanged.

Hidden-state patching conflates the two. A KV-level design that reports key-side and value-side interventions separately, with sink entries preserved, would be the first to characterize an instruction along both.

6. Generation-Stage Dynamics: Reasoning and the Final Answer

The proposal predicts that the reasoning phase and the final-answer phase consult different requirements. Evaluating that prediction requires an account of what each phase computes and where its information comes from. This section reviews mechanistic studies of chain-of-thought and arithmetic, evidence on when the final answer is determined and how answer tokens depend on reasoning tokens, the faithfulness problem that complicates the mapping from visible reasoning to internal computation, and the interventions that control reasoning length and style. It concludes with the phase-dependent routing signature these results predict.

6.1 What chain-of-thought tokens compute and read

Towards Understanding CoT Figure 1: standard, valid chain of thought, and invalid reasoning demonstrations compared
Figure — Towards Understanding Chain-of-Thought Prompting: What in a worked example is actually doing the work. Each row is a different demonstration given to the model; the right-hand column is what the model then does on a new, unseen question. Standard shows only the final answer in the demonstration, and the model gets the new question wrong (18). CoT shows correct step-by-step working, and the model works step by step too and gets it right (42). Invalid Reasoning is the interesting row: the demonstration’s arithmetic is deliberately broken — the orange text contains steps that do not follow — yet it keeps the same relevant quantities and the same order of operations. The model still reasons step by step on the new question and still reaches 42. The bar chart underneath scores all three: standard sits near 16, valid chain of thought near 48, and invalid reasoning near 40, recovering most of the benefit. So what the demonstration supplies is largely the shape of the procedure, not the correctness of its content.

Behavioral ablations first established that the content of a chain-of-thought prompt matters less than its structure: invalid rationales retain 80–90% of chain-of-thought performance provided they are relevant to the query and follow the correct ordering of steps [102]; that is to say, demonstrations whose reasoning steps are deliberately invalid still deliver most of the benefit, so what the examples supply is largely the shape of the answer rather than its content, and corrupting the symbols in a demonstration barely hurts while removing its patterns or text does [103]. These findings imply that a reasoning requirement primarily supplies a format or scaffold, which is consistent with the shallow-mechanism null hypothesis of Section 3.1. Mechanistic studies then identified the attention circuits that implement stepwise computation. In controlled settings, chain-of-thought ability emerges as an iteration head, in which one head attends from the current reasoning token to the next input token to be processed while another retrieves the previous reasoning state, so that generated tokens implement a loop over the prompt [93]; concretely, how far through the prompt the model has got is tracked by the text it has already written. Dutta et al. applied activation patching, attention analysis and probing to Llama-2-7B on fictional-ontology reasoning and found a functional rift around the middle layers — a depth above which the computation is doing a visibly different job from below it — parallel answer-generation pathways, and heads that move information from the in-context examples and the question into the generated reasoning tokens [94]. Arithmetic has been traced in more detail. At the last token, early-layer heads carry the operands and the operator across from the question, and the feed-forward layers from the middle onward are where the result is actually computed [95]. Circuit discovery on Llama-3 adds a caution: what does the computing is a sparse set of feed-forward neurons running simple heuristics, not one clean algorithm [96]. Each GSM8K reasoning step therefore has a known division of labor: attention transports operands from prompt or previously generated positions, feed-forward layers compute — a feed-forward layer being the part of each block that processes one position on its own, with no access to any other position. In other words, attention decides which numbers meet, and the feed-forward layers decide what happens to them when they do; the reasoning requirement can influence this only by shaping which positions are transported and whether the step is verbalized at all.

6.2 When the answer is fixed and what answer tokens read

Several results indicate that the final answer is often determined well before the final line is emitted. Answer-convergence analysis finds that on mathematical tasks models settle on their final answer after roughly 60% of the reasoning steps, so that stopping on answer stability saves more than 40% of tokens with little loss [97]; that is to say, the last two fifths of the written working usually does not change the answer it arrives at; probes on hidden states predict chain-of-thought success with 60–76% accuracy before any reasoning token is generated [98] — better than the 50% of a coin flip, and read off the model's state at the moment it has finished reading the problem; and overthinking analyses show that o1-style models spend most of their tokens after the first correct solution [113]. If the answer content is fixed early, the tokens of the final line have little computation left to do beyond formatting, which predicts that their attention should be dominated by the format and unit requirements rather than by the question. Two attention studies of distilled reasoning models describe the answer phase directly. Zhang et al. find that answer tokens attend substantially to reasoning tokens rather than only to the question, identify reasoning-focus heads in middle layers that track the progression of the reasoning, and show by activation patching that perturbing key reasoning-token activations reliably changes the final answer [110]. Thought Anchors identify receiver heads, whose attention over prior sentences has high kurtosis, and show that they concentrate on planning and backtracking sentences; suppressing attention to an anchor sentence changes downstream logits, and resampling the sentence — generating that one sentence again many times and letting each version continue — changes the final-answer distribution [109]. Together these studies give the proposal the two head classes most likely to carry the phase-dependent routing it hypothesizes: reasoning-focus heads during the trace, and receiver-style heads at the transition to the answer.

Figure 1 (Thought Anchors): three methods for attributing importance to sentences in reasoning traces
Figure 1 — Thought Anchors: What is shown: three different ways to ask which sentence inside a reasoning trace actually mattered, applied to one arithmetic problem. A labels the sentences of the trace by what they do. B lists the three methods: delete a sentence and resample the whole continuation a hundred times; find the heads that attend back to that sentence; or mask all attention to it and watch the later sentences change. C is the resulting dependency graph between sentences. Why it is here: this is the methodology the proposal borrows — but it is applied to sentences the model generated, not to the requirements it was given. That is to say, all three methods ask the same question in different ways, and C shows the sentence-to-sentence dependencies that come out of it. Source: Bogdan et al. [109].
Figure 4 (Thought Anchors): vertical attention scores for sentences by different heads in layer 36
Figure 4 — Thought Anchors: Left: every attention head in layer 36, plotted by how much attention it sends back to each earlier sentence. One head (dark blue) spikes hard on a handful of specific sentences while the rest stay flat. Right: a histogram over all heads of how peaked that curve is. Most heads are unremarkable; a thin tail is extremely peaked. What to read off it: a small number of heads specialise in reading back a particular earlier piece of text rather than attending broadly. That is exactly the shape the proposal predicts for a head dedicated to one prompt requirement, which is why this figure is the template for the head-ranking step. Concretely, the peakedness plotted on the right is the number used to pick heads: a high value means the head reads a few places, a low value means it reads everywhere. Source: Bogdan et al. [109].

6.3 Faithfulness: visible reasoning is not always the computation

Attention weights for the tokens preceding a conclusion, averaged over heads versus one reasoning-focus head
Figure — Reasoning-focus heads: The same conclusion read two ways. In both panels the model has just written C is 0, boxed in red, and the shading shows which earlier tokens that conclusion is attending to. Top, averaged over all heads: the weight lands on nearby words — and, C, integers, the immediately preceding phrase. Read this way the conclusion looks like it came from its own sentence. Bottom, one head, L16.H2: the same conclusion instead reaches back to an earlier sentence, underlined in red — but 0r^3 is just 0, so we can ignore that — which is the step that actually licenses C being zero. Averaging over heads washed that dependency out. This is the practical form the faithfulness problem takes: the visible text and the head-averaged attention both tell one story, and a particular head tells a different and more accurate one.

The mapping from a reasoning requirement to the behavior produces explicit reasoning is complicated by the fact that the produced reasoning need not be the computation that yields the answer. Arcuschin et al. document, without adversarial prompting, implicit post-hoc rationalisation, silent restoration of intermediate errors, and unfaithful shortcuts in frontier and open reasoning models [99]. Attribution-graph analysis of Claude 3.5 Haiku separates faithful cases, cases in which the answer is produced without grounded computation, and backward-chaining cases in which the model works from a user-supplied hint; in the latter the answer token's computation reads from the hint rather than from the reasoning steps — that is to say, the answer was taken from the hint and the visible steps were written afterwards to lead to it — and the same analysis shows that the model's verbal account of addition (column arithmetic) does not match its internal parallel heuristics [100]. For the proposal, faithfulness is both a confound and an opportunity. It is a confound because the presence of reasoning text does not guarantee that the reasoning requirement changed the computation. It is an opportunity because the proposal's causal design can distinguish the cases: if suppressing the reasoning requirement changes the final answer only when the answer tokens depend on the reasoning tokens (in the sense of Zhang et al. [110]), then the requirement is doing computational work; if it changes only the presence of reasoning text, the requirement is a stylistic control.

6.4 Controlling reasoning length and style from inside the model

ThinkEdit Figure 1: extracting a long-minus-short reasoning direction, then editing it out of selected attention heads
Figure — ThinkEdit: Reasoning length treated as a direction, then removed by surgery. Step 1 taps the residual stream at two points inside a decoder block — just after attention and just after the multi-layer perceptron — and at each computes the average representation when the model reasons at length minus the average when it reasons briefly. The difference is the reasoning-length direction for that layer. Step 2 does not add that direction at inference. It edits the weights: for a selected attention head, the output projection is replaced by a version with the short-reasoning direction projected out, so the head can no longer write along it. The intervention is therefore permanent and costs nothing at generation time, which distinguishes it from the steering-vector methods in section 3.

Interventions that alter reasoning behavior without changing the prompt identify the internal correlates a reasoning requirement must act upon. Venhoff et al. isolate backtracking, uncertainty expression, example testing and verification as separate directions in the activation space of DeepSeek-R1-distilled models, locate causally relevant layers by attribution patching, and steer each behavior with its vector over 500 tasks [104]. ThinkEdit finds that reasoning length is governed by a linear direction in middle-layer residual streams, that a small fraction of attention heads (about 2–4% depending on the model) project strongly onto the short-reasoning side, and that editing only those heads' output projections lengthens chains and improves mathematical accuracy [105]; concretely, 2–4% means only a small minority of the model's attention heads push toward stopping early, and changing just those makes the model reason for longer. Steering vectors trained by reinforcement learning recover much of full fine-tuning's reasoning gain and act, in the last layers, as token-substitution biases that favor structural tokens such as Step at early positions and process words and symbols later [106]. Post-training creates new attention heads in middle-to-late layers that persistently build reasoning pathways, and think-on/off models recruit broader, less efficient head sets when thinking is disabled [108]. On the format side, the only component-level analysis of a specific output requirement is that of Rocchetti and Ferrara, who use cumulative weighted attribution — each component's contribution to the output, weighted and then summed layer by layer so that one can see where the ability accumulates — to show that instruction tuning improves exact word-count control mainly by specializing deeper-layer components, with late attention heads contributing positively in English and final-layer feed-forward blocks compensating in Italian [107]. The consistent picture is that reasoning-style control is carried by a modest number of identifiable middle-layer heads and directions, whereas format control is enforced late; a requirement of each type should therefore be expected to engage different layers, which is testable with the per-layer knockout sweeps reviewed in Section 7, knockout meaning that one connection is switched off at one layer at a time and the effect on the output is recorded.

6.5 The predicted phase-dependent signature

Thought Anchors Figure 2: resampling accuracy versus forced-answer accuracy at each sentence of a reasoning trace
Figure — Resampling versus forcing: Two phases, visible as a gap between two curves. Both panels walk along the sentences of one reasoning trace, left to right. The solid line resamples: cut the trace after sentence i, let the model carry on normally a hundred times, and record how often it ends up correct. The dashed line forces: cut at the same place and demand an answer immediately. Panel A is a problem the model gets right. The solid line is already near 0.8 from the very first sentence, while the dashed line sits flat at zero until about sentence 43 — for the whole first half, the answer cannot be extracted by asking, even though the trajectory is already on course. Panel B is a problem it gets wrong, and the same split appears. The gap between the two lines is the phase boundary: before it, the reasoning is doing work whose result is not yet readable at the answer position; after it, the answer is committed. Which measurement you take decides which phase you can see, so a design that only forces answers will miss everything to the left of that boundary.

Combining the results above yields a concrete prediction that the proposal can test. During the reasoning phase, tokens should route to the reasoning requirement through middle-layer heads of the reasoning-focus or iteration type, and this routing should be strongest at step boundaries, where the decision to verbalize a step is made; attention to the format and unit requirements should decline as the trace lengthens, in line with the requirement-attention decay reported by Li et al. [17]. At the transition to the final line, receiver-style heads should read from the reasoning trace to fetch the converged answer [109, 110], while late-layer heads should re-engage the format and unit requirements, whose enforcement the length-control analysis places in deep layers [107]. A model that violates the unit requirement on the last line should therefore show, relative to a compliant generation on the same problem, lower late-layer $A_{t,\text{no-unit}}$ at the final tokens, which is a directly measurable contrast. In other words, the two generations are compared, and the prediction is that the failing one spent less of its deep-layer reading on the no-unit requirement, the requirement that the last line carry no unit, at exactly the moment the last line was written; the causal arm can then test whether restoring that attention by Spotlight-style amplification [66] repairs the violation.

The predicted phase-dependent signature, and the contrast that would falsify it
1
Reasoning phase
Tokens route to the reasoning requirement through middle-layer reasoning-focus / iteration heads, strongest at step boundaries [93, 110].
2
Decay during the trace
Attention to the format and unit requirements declines as the trace lengthens [17, 111].
3
Transition to the last line
Receiver-style heads read from the reasoning trace to fetch the converged answer [109, 110].
4
Final-answer phase
Late-layer heads re-engage the format and unit requirements, whose enforcement sits in deep layers [107].

The measurable contrast

A generation that violates the unit requirement on the last line should show, relative to a compliant generation on the same problem, lower late-layer $A_{t,\text{no-unit}}$ at the final tokens. That is a direct measurement, not an interpretation: two runs, one number each, compared.

▼

The causal follow-up

If the contrast holds, restoring that attention by Spotlight-style amplification [66] should repair the violation. If it does not, the routing account is wrong even though the correlation held.

7. Causal Intervention Methodology

The third research question requires showing that the routes identified in Sections 4 through 6 are causally responsible for the behaviors of Section 2. The proposal lists four intervention types: suppressing attention from generated tokens to a requirement's span, patching or replacing the requirement's hidden states or KV entries, removing or rewriting the requirement, and blocking its information flow at particular layers and heads. Each has an established methodological lineage, and each has known pitfalls. This section reviews the patching family and its design choices, attention-specific interventions, interchange interventions and causal abstraction, prompt-level attribution for whole generations, and the evaluation standards that determine whether a causal claim survives scrutiny.

7.1 The patching family and its design choices

ROME Figure: causal tracing by corrupting the subject, restoring single clean states, and mapping where the output is fixed
Figure — Causal Tracing (ROME): The patching template, with its result. (a) The clean run takes The Big Bang Theory premieres on and produces the right answer, CBS. (b) The corrupted run damages the subject tokens at the input (the starred ones) so the answer is lost. (c) One state at a time is copied from the clean run into the corrupted run, and (d) the experiment records whether the right answer comes back. The heatmaps below are that record: rows are token positions, columns are layers, and colour is the probability of the correct answer being restored by patching there. Two bright regions appear rather than one — an early site sitting on the last subject token in the early-to-middle layers, and a late site sitting on the final token in the deep layers. Splitting the same map by component shows the division of labour: the early site belongs to the MLPs (green), the late site to attention (red). The bottom row repeats the whole thing averaged over a thousand prompts, so the pattern is not an artefact of this one sentence.

All patching methods follow one template, set by causal mediation analysis: change the input, hold one internal component at the value it would have had under the changed input, and split the total effect into what flows through that component and what flows around it 114. That is to say, the model is run twice on two inputs that differ in one thing, one internal component is forced to take the value it had in the second run, and the question is how much of the change in the output that single substitution accounts for. Causal tracing adapted it to multi-token spans 115, and an instruction span can be treated the same way. The methodology literature contributes the warning that four design choices inside that template can each flip the answer on their own. The table below lists those four: the choice, the options available for it, and what the methodology literature recommends.

Design choiceOptionsWhat the methodology literature recommends
MetricProbability, logit, logit difference, or Kullback–Leibler divergence — the logit difference being the gap between the score of the right answer and the score of the wrong one, and the divergence a single number for how far the whole distribution over next words has moved.Use logit difference [116], because it cancels effects that push every candidate up or down together.
CorruptionGaussian noise on the span, or counterfactual token substitution.Counterfactual substitution, not noise [116]. For this proposal the corrupt baseline is a prompt with the requirement rewritten — deleting it also changes length.
DirectionDenoising (restore clean state) or noising (corrupt a clean run).Denoising identifies sufficient components, noising identifies necessary ones — report both [117].
Ablation valueZero, mean, or resampled from another prompt.Resample ablation from a different prompt [117]: the component is given a value it really takes on some other input, rather than a value it never takes.
Each of these four can flip a conclusion on its own. Zhang and Nanda show the metric, the corruption type and the direction each independently change which components a patching experiment identifies [116]; Heimersheim and Nanda add that clean and corrupt prompts must differ in the variable of interest and nothing else [117]; concretely, if the rewritten requirement also changes the prompt's length or wording elsewhere, the measured effect belongs partly to those changes and cannot be assigned to the requirement.

7.2 Attention-specific interventions

The proposal's first intervention, suppressing attention from later tokens to a requirement span, has a canonical precedent in attention knockout: Geva et al. block attention from a target position to a chosen span of source positions over a window of layers by zeroing those attention weights, sweep the window, and report per-layer curves that reveal a three-step recall mechanism [125]; that is to say, the block is applied at one depth at a time and the resulting curve says at which depths the connection was load-bearing. Wang et al. applied the same blocking to label-word anchors in shallow versus deep layers and showed that the two produce different effects [70], and Thought Anchors applied attention suppression to sentences inside a generated reasoning trace, measuring the Kullback–Leibler divergence of downstream logits [109]. A known complication is that knockout renormalizes the remaining attention, so its effect mixes removal of information with redistribution of attention — put simply, the shares must still add up to one, so weight taken away from the requirement is handed to everything else, and the observed change could be caused by either; reporting both zero-masking and resample-based key replacement helps separate the two. The inverse manipulation, amplification, is provided by PASTA, InstABoost and Spotlight [64, 65, 66], and using both directions yields a dose–response curve rather than a single ablation point. Competition-of-mechanisms work modifies the attention of specific late heads to specific positions to flip a model between factual recall and in-context copying [142], which is structurally the same manipulation as flipping between a default output format and an instructed one; its reproduction study found that head specialization largely disappears in Llama-3.1-8B and that ablation effects depend on prompt structure and domain [143], a warning that the proposal should run at least two model families.

Figure 6 (Thought Anchors): case study of attention-suppression effects, receiver-head attention and the reasoning flowchart
Figure 6 — Thought Anchors: What is shown: one problem worked end to end. Right: the model’s reasoning broken into labelled chunks — it reaches 20 bits, backtracks, recomputes in decimal, notices the discrepancy, and settles on 19. Top left: a matrix of how much each sentence causally depends on each earlier one, measured by suppressing attention to the earlier sentence and watching the later one move. Bottom left: which sentences receive concentrated attention from the specialised heads. Why it is here: it is the full protocol the causal arm of the proposal would run, only with prompt requirements in place of reasoning sentences. That is to say, the matrix at the top left is built by intervening rather than by reading attention off, so each cell is a measured effect and not a correlation. Source: Bogdan et al. [109].

7.3 Interchange interventions, causal abstraction and cross-prompt patching

Interchange success by layer and token position for a causal abstraction of a BERT model
Figure — Interchange success: Where an interchange intervention actually works. The claim being tested is that a high-level symbolic variable is realised in the network at some particular place. To test it, the value of that place is swapped in from a second input and the model is asked to behave as the symbolic account predicts. The grid records how often that succeeded: rows are layers, columns are token positions in the sentence pair — [CLS], the adjective and noun of the premise object, the separator, the same two for the hypothesis, and the final separator. Success is concentrated in the lowest one or two layers at the object-noun positions, darkest at the hypothesis noun, and falls away almost entirely above about layer two, with [CLS] holding some residue higher up. An alignment claim is therefore not a claim about the model as a whole. It is a claim about a specific layer and a specific position, and the same claim tested one layer up can simply fail.

The proposal's second intervention, patching or replacing a requirement's hidden states, is an interchange intervention in the sense of causal abstraction: a high-level causal model is an abstraction of the network if swapping in a representation from a counterfactual input has the same effect as intervening on the corresponding high-level variable [127]. In other words, one writes down a small hand-made diagram of what the model is supposed to be doing, and the diagram earns the right to be called a description of the network only when editing the network at the matching place produces exactly what editing the diagram would predict. Distributed alignment search learns the subspace on which such interventions succeed, and its Boundless variant (Boundless DAS) applied it to Alpaca-7B on a numeric instruction task, finding that the instruction-tuned model implements an interpretable boolean causal model in identifiable residual subspaces robust to instruction wording [128]; this is the closest precedent for hypothesizing a small causal model of requirement compliance (reasoning required, bare number required, units forbidden) and testing it by interchange interventions on requirement-span representations. Patchscopes unifies logit lens, causal tracing and cross-prompt patching as one operation, patching a representation from a source prompt, layer and position into a target and decoding it in natural language — concretely, the internal state in question is dropped into a second prompt that asks the model to describe it, and the model's answer is read as a report on what that state held — which provides a way to ask what a model believes a requirement says at a given layer [129]. Task- and function-vector patching [42, 43] supplies a further condition, replacing the requirement text with its vector, that tests whether the requirement is needed as tokens at all. Causal scrubbing offers a whole-hypothesis test: resample every activation the hypothesis deems irrelevant from other prompts sharing the same requirement and report the fraction of behavior preserved [126]. That is to say, everything the account claims does not matter is replaced with material from elsewhere; if the account is right the behavior survives, and how much of it survives is the score. Sparse feature circuits move the same logic to sparse-autoencoder features, enabling ablation of individual requirement-related features during generation [137].

7.4 Attributing whole generations to prompt segments

ContextCite Figure 1: a generated sentence traced back to the specific source sentence that caused it
Figure — ContextCite: Attribution at the granularity of a statement. The context is a document about the 2024 solar eclipse and the query asks where to go to see it from Boston. The model answers in several sentences; one of them — you can travel to Maine, which is on the path of totality — is picked out and highlighted. ContextCite then points back into the source document and marks the specific sentence responsible: the path arches from Mexico to Texas to Maine. Note what is being attributed and to what. The unit on the output side is a chosen statement rather than a single token, and the unit on the input side is a span of the context rather than a layer or a head. That is the granularity a workflow needs when the question is which part of a long prompt produced a particular claim.

Because the proposal's outcomes are properties of complete generations, it needs attribution methods that operate at that granularity. ContextCite ablates random subsets of context sentences, records the log-probability of a fixed response or span, and fits a sparse linear surrogate whose weights attribute the response to each source [130] — that is to say, it learns a simple formula predicting how likely the response is from which sentences were present, and the size of each term in that formula is that sentence's credit; applied to the requirement block, it scores each requirement's influence on, for instance, the last line, at the cost of an additivity assumption that may be violated when requirements interact. Put simply, the method assumes each requirement's contribution simply adds up, which is exactly what stops being true when two requirements pull against each other. AttriBoT makes exact leave-one-out attribution — rerun the model with one requirement removed, once per requirement, and take the difference — tractable with cached activations and hierarchical attribution [132], and Learning to Attribute with Attention trains a per-head weighting so that attention predicts ablation-based attributions, which identifies the heads whose attention is causally informative and is therefore a principled way to select heads for knockout [131]. Inseq provides step-wise attribution of each generated token to source and previously generated tokens with aggregation over spans [133], and contrastive explanations attribute the difference between two candidate tokens, such as a bare number versus a number followed by a unit, rather than the raw probability [134]; AttnLRP propagates relevance through attention and normalization layers for token-level attributions in large models [135]. Any rewriting of a requirement must be compared against a distribution of surface-form variants, since format changes alone swing accuracy by tens of points [23].

7.5 Standards for a causal claim

IOI circuit schematic with the two intervention sites and the resulting flip in predicted name
Figure — IOI circuit and interventions: What a fully worked causal claim looks like. The prompt is Rudi and Emma … Emma says … to, and the answer should be the name mentioned once. The diagram names every part that matters and how they connect: the S-Inhibition heads read the repeated name, the Name Mover heads carry the remaining name to the output, and the residual stream runs up the right-hand side. Two injection sites are marked, vMLP into the perceptron path and vresid straight into the residual stream — naming where an intervention enters, not just that one happened. The boxes at the top are the outcome, in counts out of a hundred: the baseline model answers Rudi 96 times, and after the patch it answers Emma 92 times. A claim reaches this standard when the components are identified, the intervention site is stated, and the effect is a near-complete swap rather than a shift in an average.

Several results define what the proposal must report for a routing claim to be credible. Makelov et al. demonstrate an interpretability illusion in which patching along a learned subspace produces the target behavior by activating a dormant parallel pathway that the model does not use on clean inputs, and propose diagnostics that check whether the patched direction is actually used [138]; that is to say, the intervention works and still explains nothing, because it switched on machinery the model never reaches for on its own; any learned-subspace intervention on requirement representations needs this check. Miller et al. show that circuit faithfulness scores — how much of the model's behavior a proposed set of components reproduces on its own — change substantially with the ablation type, direction and metric, to the point that the best circuit changes, and recommend reporting multiple settings [139]. Shi et al. formalize mechanism preservation, localization and minimality as statistical hypothesis tests with error control and find that few published circuits pass all of them [140]; concretely, preservation asks whether the proposed components alone still produce the behavior, localization whether everything outside them can be disturbed without effect, and minimality whether any of them can be dropped. For interventions on open-ended generation, Arditi et al. provide the exemplar of scoring directional ablation by behavioral metrics of whole generations alongside capability benchmarks [55], and Pres et al. propose criteria for steering evaluations: deployment-like contexts, prompting baselines, likelihood shifts of both target and non-target behaviors, and effect sizes with uncertainty [141]. The reproduction study of Dotsinski et al. adds that head-level findings can be model-specific [143]. The table below maps each of the proposal's planned interventions to its methodological precedent and the pitfall that precedent identified. Each row is one planned experiment: the left column says what would be done to the model, the middle column names the published work that has already done something of that shape, and the right column names the mistake that work showed is easy to make.

Planned intervention Methodological precedent Pitfall to control
Suppress attention from generated tokens to span $\mathcal{T}_i$ Attention knockout over layer windows [125]; anchor blocking [70]; sentence suppression in reasoning traces [109] Renormalization mixes removal with redistribution; exclude sink positions [75, 76]; add uniform / frozen attention controls [81]
Patch or replace hidden states of $I_i$ Interchange interventions and Boundless DAS on an instruction-tuned model [127, 128]; Patchscopes cross-prompt patching [129] Dormant-pathway illusion [138]; use resample ablation with a rewritten requirement, not zero ablation [117]; report denoising and noising [116]
Patch or replace KV entries $K_j, V_j$ of $I_i$ KV Cache Steering [90]; V-Steer value edits [83]; Memory Inception layer/head selection [91] Preserve sink entries; separate key-side from value-side effects; attribution graphs omit QK effects [101], AtP* QK-fix [121]
Remove, replace or rewrite $I_i$ ContextCite and leave-one-out attribution [130, 132]; per-requirement steering-vector contrast [38] Length and position shift: use schemas [124]; surface-form sensitivity [23]; added-clause cost on GSM8K [26]
Block information flow of $I_i$ at specific layers / heads Path patching [118, 119]; edge attribution patching with integrated gradients (EAP-IG) and AtP* screening [121, 122]; memory-vs-context head pruning [72] Faithfulness not robust to ablation choice [139]; hypothesis tests for localization [140]; model-family dependence [143]
Score on GSM8K accuracy, reasoning presence / length, format and unit compliance Generation-level scoring of directional ablation [55]; resampled continuations [109]; steering evaluation criteria [141] Report accuracy and compliance jointly [16]; text-based not first-token compliance [36]; judge bias on soft requirements [29]

7.6 Tooling for interventions during generation

pyvene Figure 1: a declarative intervention config that adds a word embedding into the residual stream during generation
Figure — pyvene: What an intervention looks like as code. The top strip is the story Tinystories-33M writes unprompted. The middle diagram shows the intervention: the embedding of the single word happy or sad, scaled down, is added into the residual stream at the output of every layer’s multi-layer perceptron, at every generated position — those are the red arrows into the summation nodes. The code below is the whole specification. An intervention is declared as a list of dictionaries naming a layer, a component (here mlp_output) and an intervention_type (here addition), the model is wrapped once, and generation proceeds normally with the source representation passed in. The two strips at the bottom are the result on the same opening sentence: with happy added, Lucy is happy and thanks the man; with sad added, the same setup turns and Lucy is scared. The point for a workflow is that the intervention is configuration rather than a patched forward pass.

Two libraries support interventions during multi-token generation with Hugging Face models. pyvene specifies interventions declaratively by layer, token positions, subspace and type, including interchange, zero, noise and trainable interventions, and implements DAS and Boundless DAS [144]. NNsight defines intervention graphs over any PyTorch model, including inside .generate() loops, and NDIF runs them remotely on large models [145]; this is the more natural fit for KV-level and hidden-state patching over full GSM8K solutions. Gemma Scope sparse autoencoders provide off-the-shelf feature bases for every layer of an open model, should the proposal pursue feature-level decomposition of requirement signals. In other words, both libraries allow a value inside the model to be changed midway through writing an answer, which ordinary inference code does not permit.

8. Synthesis, Research Gaps and Recommendations

This section consolidates the preceding review into an evidence map for the three research questions, identifies the gaps the proposal is positioned to fill, states the hypotheses the literature supports, and lists methodological recommendations. The organizing claim is that every component of the proposed causal chain $I_i \rightarrow \{K,V\} \rightarrow \text{heads} \rightarrow B_j$ has been demonstrated in isolation, but no study has traced several heterogeneous natural-language requirements through that chain simultaneously on a reasoning task, with generation-level outcomes and with attention to how the routing shifts between the reasoning and answer phases. That is to say, the chain runs from the words of one requirement, to the keys and values the model stored for those words, to the heads that read them back, to the behavior that finally appears in the output; each link has been demonstrated on its own, and the missing work is following all four at once for several requirements in the same prompt.

8.1 Evidence map for the three research questions

ComplexBench Figure 3: five ways requirements compose within a single instruction
Figure — Composition types: The structure any evidence map has to cover, because a requirement rarely arrives alone. Five ways requirements combine, each with a description, a worked instruction and a graph. Single is one requirement. And is several that must hold at the same time: summarise this news, in bullet points, within 100 words — the graph is a triangle because none of the three depends on the others. Chain is several tasks in a fixed order, each with its own requirements, drawn as a path. Selection is conditional: if the work contains an animal, describe it in English, otherwise in Chinese — a branch node with two children, and only one branch is in force for any given input. Nested Structure is these composed recursively, giving a tree several levels deep. The evidence in this review is organised against this shape: results about a single requirement do not automatically transfer to conjunctions, and results about conjunctions say nothing about the conditional case, where the question of which requirement governs a token has a different answer depending on the input.

The table lists, for each level of the proposal, what the literature has established, the closest existing work, and what remains open. Each row is one of the three levels defined in Section 1, read from left to right as settled, cited, and still missing.

Level Established Closest work Open
L1 Requirement–behavior mapping Requirements can be scored per requirement; compliance decays with count and is worst for format / lexical types; format requirements trade off against reasoning; position, surface form and conflict all matter IFEval, FollowBench, InFoBench, ComplexBench [1, 2, 3, 4]; MathIF [16]; When Thinking Fails [17]; Let Me Speak Freely [18] No study varies requirements one at a time on a fixed reasoning task and records the full behavior vector (accuracy, reasoning presence / length, last-line format, unit removal); no ablation of requirement order within one block
L2 Internal routing Instruction tokens are persistently consulted; per-requirement steering vectors exist; instruction tokens act as anchors read by deep layers; attention to instructions decays over generation; heads specialize by instruction source; attended prompt positions are stable across decoding Wu et al. [34]; Stolfo et al. [38]; Attention Satisfies [62]; Instruction Anchors [69]; Label Words are Anchors [70]; Li et al. [67]; Dongre et al. [68]; SnapKV [88] No per-requirement $A_{t,i}^{(l,h)}$ maps for a multi-requirement prompt; no comparison of reasoning-phase versus answer-phase readers; no key-versus-value decomposition of an instruction's effect
L3 Causal influence Attention knockout and amplification on instruction spans change compliance; KV edits change reasoning and priority; instruction-conditioned behaviors are mediated by steerable directions; generation-level evaluation of interventions is feasible Geva et al. [125]; PASTA / InstABoost / Spotlight [64, 65, 66]; KV Cache Steering [90]; V-Steer [83]; Arditi et al. [55]; Thought Anchors [109]; Boundless DAS [128] No segment-specific counterfactual replacement of one requirement's KV entries; no per-requirement, per-layer knockout sweep scored on a behavior vector; no test of whether the same heads mediate different requirements

8.2 Gaps the proposal is positioned to fill

Instruction steering Figure 3a: format-instruction accuracy with no text instruction, standard inference versus steering
Figure — Instruction steering: A requirement satisfied with no text stating it. The task is to produce output in a required format, and the text instruction asking for that format has been removed from the prompt entirely. Solid bars are ordinary inference under those conditions; hatched bars add a steering vector for the format requirement and change nothing else. Across four instruction-tuned models — Phi-3, Gemma 2 2B, Mistral 7B and Gemma 2 9B — ordinary inference lands between roughly 0.07 and 0.13, near the floor you would expect when the model was never told what to do. Steering lifts every one of them to roughly 0.29 to 0.32. The vector is carrying the requirement that the missing sentence would have carried, which is the premise the proposal in this section rests on.

Five gaps recur across the sections above. Each corresponds to a component of the proposal that no existing study has delivered.

Heterogeneous requirements traced together

Existing mechanistic work treats one instruction type at a time (a format requirement [38], a length requirement [107], a modality instruction [69], a persona [56]). No study traces a reasoning requirement, an answer-format requirement and a lexical requirement through the same forward passes and compares their carriers. That is to say, each of the four studies cited followed a single requirement type, so nothing in them says whether two requirements are carried by the same parts of the model or by different ones.

Routing across the steps of one chain of thought

Attention decay to instructions is measured across dialogue turns [67] or toward a single goal span [68]; requirement-attention decline within reasoning is reported in aggregate [17]. Per-requirement attention maps over the reasoning and answer phases of one generation do not exist.

Segment-level KV counterfactuals

KV edits so far add a vector at the last prompt token [90], scale whole priority spans [83] or insert new banks [91]. Replacing the keys and values of one requirement with those of another, and separating key-side from value-side effects, is untested.

Generation-level behavior vectors as the outcome

Most patching work scores a next-token logit difference [116]; the few generation-level evaluations use a single behavior [55, 109]. Scoring each intervention on accuracy, reasoning presence and length, final-line format and unit removal simultaneously, as MathIF does behaviorally [16], has not been done mechanistically. In other words, without all four scores at once an intervention that improves the format by deleting the reasoning would be recorded as a success.

Shared versus dedicated mediators

Function-vector work shows instruction- and demonstration-derived vectors use different heads [44]; conflict studies show separate subspaces for conflict types [85]. Whether two requirements in one prompt share reader heads, and whether interference at the head level explains the count-dependent decay of Section 2, is open.

8.3 Hypotheses supported by the literature

Causal inner product Figure 1: two representations of a concept unified and made orthogonal under the right inner product
Figure — The causal inner product: The geometric hypothesis this section rests on. On the left sphere, each concept appears twice. In blue is the concept read off the output side, for example male⇒female as γ(“queen”) − γ(“king”), and English⇒French as γ(“roi”) − γ(“king”). In red is the same concept read off the context side, as λ(“She is the”) − λ(“He is the”) and λ(“Il est le”) − λ(“He is the”). Under the ordinary geometry the blue and red arrows for one concept point in different directions, and the two concepts sit at an arbitrary angle. On the right sphere, after changing to the inner product the paper derives, the two readings of each concept coincide, and the two concepts meet at a right angle. The hypothesis is that a requirement has a single direction and that separable requirements are orthogonal — but only once you are measuring in the right coordinate system, which is not the raw activation space.

The evidence reviewed licenses the following hypotheses, stated so that each has a clear falsifying observation. Each row gives the hypothesis, the evidence that makes it worth testing, and the specific result that would show it to be wrong.

Hypothesis Supporting evidence Falsifying observation
H1. Each requirement type has a separable residual-stream signature extractable by with / without contrast, and the format-type signatures are concentrated in late layers. [36, 38, 39, 55, 107] Steering with one requirement's vector changes the behaviors of the others as much as its own; signatures are not separable under the causal inner product [50]
H2. Requirement tokens function as anchors: attention to a requirement concentrates on its delimiter positions in shallow layers, and later readers consult those positions rather than the content words. [69, 70, 78, 87] Knockout of delimiter positions has no effect while knockout of content words does
H3. Reasoning tokens route to the reasoning requirement through middle-layer heads, and last-line tokens route to the format and unit requirements through late-layer heads; the two reader sets are largely disjoint. [62, 93, 105, 107, 109, 110] The same heads carry the largest $A_{t,i}$ for every requirement in both phases, or knockout effects are phase-independent
H4. Attention share to the format and unit requirements declines over the reasoning phase, and violations on the last line are preceded by lower late-layer $A_{t,i}$ than compliant generations. [17, 67, 68, 111] Compliant and violating generations show indistinguishable $A_{t,i}$ trajectories
H5. Suppressing attention to a requirement, or replacing its KV entries with a rewritten requirement's, changes primarily that requirement's behavior, with value-side replacement sufficient to transfer the effect. [43, 60, 83, 90, 125] Effects are diffuse across behaviors, or key-side replacement is required for any transfer
H6. Adding requirements degrades compliance through competition for attention share among reader heads rather than through loss of the requirement representations, so that amplification restores compliance. [10, 11, 12, 65, 66, 68] Probes show requirement information absent from the residual stream after many requirements are added, and amplification does not help

8.4 A recommended experimental pipeline

Thought Anchors Figure 1: sentence categories in a reasoning trace, three attribution methods, and the resulting influence graph
Figure — Thought Anchors: Three ways of asking the same question, and what they produce together. Panel A takes one problem and its reasoning trace, and labels each sentence by what it is doing: sentence 8 is Active Computation, 13 is Plan Generation, 47 is Uncertainty Management. Panel B lists the three methods applied to those sentences. Black-box resampling deletes a sentence and re-runs the whole rollout a hundred times, rather than deleting it and forcing an answer, so the measured effect includes everything downstream. Receiver heads looks for attention heads that read a given sentence, with vertical stripes in the attention matrix marking sentences that broadcast widely. Attention suppression masks all attention to sentence i and measures the effect on sentence j. Panel C is the result: each sentence is a node whose size is its importance, and the dashed edges are sentence-to-sentence influence. Sentence 13, the plan, is the largest. Three methods with different failure modes agreeing on the same sentences is what makes the attribution credible.

The methodological literature suggests the following order of operations, in which cheap descriptive measurements narrow the search before expensive causal tests are run, and in which every causal test is scored on the full behavior vector.

Stage 1 — Behavioral factorial design

Vary each requirement independently (present / absent / rewritten / repositioned) on a fixed GSM8K subset with a no-requirement control [26]; score accuracy, reasoning presence and length, strict and loose last-line format, and unit removal with code checks [1, 4]; randomize order and surface form [23]; include trip-wire items for shortcut compliance [9, 12].

↓

Stage 2 — Descriptive routing maps

Compute $A_{t,i}^{(l,h)}$ per requirement with sink positions excluded and delimiter positions reported separately [75, 78]; summarize by generation phase, that is to say average the per-step numbers separately over the reasoning tokens and over the final-line tokens; rank heads by receiver-style kurtosis [109], query-focused retrieval score [74] and learned attention attribution [131]; locate the internalization layer by layer-wise context masking [49].

↓

Stage 3 — Representation tests

Extract a per-requirement vector by with / without contrast [38, 57]; test separability by causal-inner-product orthogonality and by cross-steering [50]; monitor each vector's projection over the generation [56]; check for shallow output-bias implementations [41].

↓

Stage 4 — Causal screening and confirmation

Screen heads and edges with EAP-IG and AtP* — that is to say, cheap derivative-based estimates that rank every head and connection so that only the top candidates need to be tested properly — using a rewritten-requirement counterfactual and schema-aligned spans [121, 122, 124]; confirm with windowed attention knockout [125], path patching [119], and Spotlight-style amplification for dose–response [66]; replace KV entries per segment, key-side and value-side separately, preserving sinks [83, 90].

↓

Stage 5 — Validation standards

Score every intervention on the full behavior vector with resampled continuations [55, 109]; report zero, mean and resample ablations in both directions [116, 139]; apply dormant-pathway diagnostics to any learned subspace [138]; use hypothesis tests for localization and minimality [140]; replicate on at least two model families [143].

8.5 Closing assessment

ComplexBench Figure 5: one instruction scored as ten dependent questions, with rule and model evaluators and dependency aggregation
Figure — Scoring with dependencies: What per-requirement scoring already looks like when it is done properly, and the complication it exposes. Top left is one instruction: describe a painting in under 100 words, then branch on whether the painting contains an animal. Below it that instruction becomes ten scoring questions, each tagged with the composition type it belongs to and the dimension it tests — chain, selection, length, factuality, helpfulness. Top right is the part that matters most. The ten questions are not independent: question 4 depends on 1, questions 5 and 6 depend on 1 and 4, questions 9 and 10 depend on 1, 4, 5 and 7. The bottom strip is the machinery: each question is routed to a rule evaluator when it can be checked programmatically and to a model judge otherwise, giving a 0 or 1 per question, and those are then combined by dependency aggregation — a question’s score is conjoined with all of its prerequisites. Fail question 1 and questions 5, 7, 9 and 10 are zero whatever the text actually did. Every tool this review needs exists, as this figure shows; what is missing is any account of why a particular requirement failed, which a score of zero cannot supply.

The literature has matured to the point where every tool the proposal needs exists and has been validated on a related problem: verifiable per-requirement scoring, contrastive extraction of per-requirement vectors, anchor-token and receiver-head analysis, KV-level editing, position-aware circuit discovery, and generation-level evaluation of interventions. What is missing is their combination on a single, well-controlled multi-requirement reasoning task, with the routing of each requirement followed across the phases of generation and confirmed by segment-specific intervention. The behavioral facts that motivate the question, in particular the count-dependent decay of compliance and the trade-off between reasoning and format requirements, are robust across models and benchmarks, and the most plausible mechanistic account of them, competition for attention share among instruction readers, is testable with the methods reviewed here. The proposal's expected contribution is therefore not a new tool but the first mechanistic account of how a language model decomposes, stores, consults and executes several prompt requirements at once, and of where in that chain the observed failures arise. Put simply, the question is no longer whether a model obeys, but which part of it stopped reading which requirement, and when.

9. References

The list is grouped by theme; numbering is global and matches the bracketed citations in the text. Every arXiv identifier was resolved against arxiv.org, the ACL Anthology, OpenReview or the publisher page at the time of writing. Entries whose author lists could not be re-verified are marked as such and should be checked before citation.

Benchmarks and behavioral studies of multi-requirement instruction following

  1. Zhou, J., Lu, T., Mishra, S., et al. Instruction-Following Evaluation for Large Language Models (IFEval). arXiv 2023. arXiv:2311.07911
  2. Jiang, Y., Wang, Y., Zeng, X., et al. FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models. ACL 2024. arXiv:2310.20410
  3. Wen, B., Ke, P., Gu, X., et al. Benchmarking Complex Instruction-Following with Multiple Constraints Composition (ComplexBench). NeurIPS 2024 Datasets & Benchmarks. arXiv:2407.03978
  4. Qin, Y., Song, K., Hu, Y., et al. InFoBench: Evaluating Instruction Following Ability in Large Language Models. Findings of ACL 2024. arXiv:2401.03601
  5. Zhang, T., Zhu, C., Shen, Y., et al. CFBench: A Comprehensive Constraints-Following Benchmark for LLMs. ACL 2025. arXiv:2408.01122
  6. He, Q., Zeng, J., Huang, W., et al. Can Large Language Models Understand Real-World Complex Instructions? (CELLO). AAAI 2024. arXiv:2309.09150
  7. Qin, Y., Zhang, T., Shen, Y., et al. SysBench: Can Large Language Models Follow System Messages?. arXiv 2024. arXiv:2408.10943
  8. He, Y., Jin, D., Wang, C., et al. Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following. arXiv 2024. arXiv:2410.15553
  9. Pyatkin, V., Malik, S., Lambert, N., et al. Generalizing Verifiable Instruction Following (IFBench). NeurIPS 2025 Datasets & Benchmarks. arXiv:2507.02833
  10. Harada, K., Yamazaki, Y., Taniguchi, M., et al. When Instructions Multiply: Measuring and Estimating LLM Capabilities of Multiple Instructions Following. EMNLP 2025. arXiv:2509.21051
  11. Ye, J., Huang, C., Chen, Z., et al. A Multi-Dimensional Constraint Framework for Evaluating and Improving Instruction Following in Large Language Models. arXiv 2025. arXiv:2505.07591
  12. Ferraz, T. P., Mehta, K., Lin, Y.-H., et al. LLM Self-Correction with DeCRIM: Decompose, Critique, and Refine for Enhanced Following of Instructions with Multiple Constraints. arXiv 2024. arXiv:2410.06458
  13. Murthy, R., Kumar, P., Venkateswaran, P., Contractor, D. Evaluating the Instruction-following Abilities of Language Models using Knowledge Tasks. arXiv 2024. arXiv:2410.12972
  14. Qi, Y., Peng, H., Wang, X., et al. AgentIF: Benchmarking Instruction Following of Large Language Models in Agentic Scenarios. arXiv 2025. arXiv:2505.16944
  15. Kwon, Y., et al. (author list not re-verified) ReasonIF: Large Reasoning Models Fail to Follow Instructions During Reasoning. arXiv 2025. arXiv:2510.15211
  16. Fu, T., Gu, J., Li, Y., Qu, X., Cheng, Y. Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models (MathIF). arXiv 2025. arXiv:2505.14810
  17. Li, X., Yu, Z., Zhang, Z., et al. When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs. NeurIPS 2025. arXiv:2505.11423
  18. Tam, Z. R., Wu, C.-K., Tsai, Y.-L., Lin, C.-Y., Lee, H.-y., Chen, Y.-N. Let Me Speak Freely? A Study on the Impact of Format Restrictions on Performance of Large Language Models. EMNLP 2024 Industry Track. arXiv:2408.02442
  19. Lee, I. Y., D'Antoni, L., Berg-Kirkpatrick, T. The Format Tax. arXiv 2026. arXiv:2604.03616
  20. Fan, H. Capacity, Not Format: Rethinking Structured Reasoning Failures. arXiv 2026. arXiv:2606.09410
  21. Liu, N. F., Lin, K., Hewitt, J., et al. Lost in the Middle: How Language Models Use Long Contexts. TACL 2024. arXiv:2307.03172
  22. Liu, Y., Zeng, X., Shao, C., Meng, F., Zhou, J. Instruction Position Matters in Sequence Generation with Large Language Models. Findings of ACL 2024. arXiv:2308.12097
  23. Sclar, M., Choi, Y., Tsvetkov, Y., Suhr, A. Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design. ICLR 2024. arXiv:2310.11324
  24. Zhang, Z., Li, S., Zhang, Z., et al. IHEval: Evaluating Language Models on Following the Instruction Hierarchy. NAACL 2025. arXiv:2502.08745
  25. Wallace, E., Xiao, K., Leike, R., Weng, L., Heidecke, J., Beutel, A. The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions. arXiv 2024. arXiv:2404.13208
  26. Mirzadeh, I., Alizadeh, K., Shahrokhi, H., Tuzel, O., Bengio, S., Farajtabar, M. GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models. ICLR 2025. arXiv:2410.05229
  27. Kojima, T., Gu, S. S., Reid, M., Matsuo, Y., Iwasawa, Y. Large Language Models are Zero-Shot Reasoners. NeurIPS 2022. arXiv:2205.11916
  28. Kong, A., Zhao, S., Chen, H., et al. Better Zero-Shot Reasoning with Role-Play Prompting. NAACL 2024. arXiv:2308.07702
  29. Zeng, Z., Yu, J., Gao, T., Meng, Y., Goyal, T., Chen, D. Evaluating Large Language Models at Evaluating Instruction Following (LLMBar). ICLR 2024. arXiv:2310.07641
  30. Cobbe, K., Kosaraju, V., Bavarian, M., et al. Training Verifiers to Solve Math Word Problems (GSM8K). arXiv 2021. arXiv:2110.14168
  31. Wei, J., Wang, X., Schuurmans, D., et al. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. NeurIPS 2022. arXiv:2201.11903
  32. Mu, N., Chen, S., Wang, Z., et al. Can LLMs Follow Simple Rules? (RuLES). arXiv 2023. arXiv:2311.04235
  33. Elder, B., Duesterwald, E., Muthusamy, V. Boosting Instruction Following at Scale. arXiv 2025. arXiv:2510.14842

Internal representations of instructions; task, function and steering vectors

  1. Wu, X., Yao, W., Chen, J., Pan, X., Wang, X., Liu, N., Yu, D. From Language Modeling to Instruction Following: Understanding the Behavior Shift in LLMs after Instruction Tuning. NAACL 2024. arXiv:2310.00492
  2. Gao, C., Huang, S., Li, J., Chen, J. Roles of Scaling and Instruction Tuning in Language Perception: Model vs. Human Attention. Findings of EMNLP 2023. arXiv:2310.19084
  3. Heo, J., Heinze-Deml, C., Elachqar, O., et al. Do LLMs "Know" Internally When They Follow Instructions?. ICLR 2025. arXiv:2410.14516
  4. Heo, J., Xiong, M., Heinze-Deml, C., Narain, J. Do LLMs Estimate Uncertainty Well in Instruction-Following?. ICLR 2025. arXiv:2410.14582
  5. Stolfo, A., Balachandran, V., Yousefi, S., Horvitz, E., Nushi, B. Improving Instruction-Following in Language Models through Activation Steering. ICLR 2025. arXiv:2410.12877
  6. He, Z., Zhao, H., Qiao, Y., Yang, F., Payani, A., Ma, J. SAIF: A Sparse Autoencoder Framework for Interpreting and Steering Instruction Following of Language Models. arXiv 2025. arXiv:2502.11356
  7. Jiang, G., Jiang, C., Li, Z., et al. Interpretable Catastrophic Forgetting of Large Language Model Fine-tuning via Instruction Vector. arXiv 2024. arXiv:2406.12227
  8. Hewitt, J., Liu, N. F., Liang, P., Manning, C. D. Instruction Following without Instruction Tuning. arXiv 2024. arXiv:2409.14254
  9. Hendel, R., Geva, M., Globerson, A. In-Context Learning Creates Task Vectors. Findings of EMNLP 2023. arXiv:2310.15916
  10. Todd, E., Li, M. L., Sharma, A. S., Mueller, A., Wallace, B. C., Bau, D. Function Vectors in Large Language Models. ICLR 2024. arXiv:2310.15213
  11. Davidson, G., Gureckis, T. M., Lake, B. M., Williams, A. Do Different Prompting Methods Yield a Common Task Representation in Language Models?. arXiv 2025. arXiv:2505.12075
  12. Xiong, Z., Cai, Z., Cooper, J., et al. Everything Everywhere All at Once: LLMs Can In-Context Learn Multiple Tasks in Superposition. ICML 2025. arXiv:2410.05603
  13. Liu, S., Ye, H., Xing, L., Zou, J. In-context Vectors: Making In Context Learning More Effective and Controllable Through Latent Space Steering. ICML 2024. arXiv:2311.06668
  14. Tikhonov, P., Oseledets, I., Tutubalina, E. One Task Vector Is Not Enough: A Large-Scale Study for In-Context Learning. arXiv 2025. arXiv:2505.23911
  15. Zheng, B., Ma, Y., Lin, Z., Yang, Y. Label Words as Local Task Vectors in In-Context Learning. arXiv 2024. arXiv:2406.16007
  16. Sia, S., Mueller, D., Duh, K. Where Does In-Context Learning Happen in Large Language Models?. NeurIPS 2024. arXiv:2403.04510
  17. Park, K., Choe, Y. J., Veitch, V. The Linear Representation Hypothesis and the Geometry of Large Language Models. ICML 2024. arXiv:2311.03658
  18. Elhage, N., Hume, T., Olsson, C., et al. Toy Models of Superposition. Transformer Circuits Thread 2022. arXiv:2209.10652
  19. Oozeer, N., et al. Beyond Linear Steering: Unified Multi-Attribute Control for Language Models (K-Steering). arXiv 2025. arXiv:2505.24535
  20. Radevski, G., Gashteovski, K., Hong, S., Lawrence, C., Glavaš, G. Compositional Steering of Large Language Models with Steering Tokens. arXiv 2026. arXiv:2601.05062
  21. Mu, J., Li, X. L., Goodman, N. Learning to Compress Prompts with Gist Tokens. NeurIPS 2023. arXiv:2304.08467
  22. Arditi, A., Obeso, O., Syed, A., et al. Refusal in Language Models Is Mediated by a Single Direction. NeurIPS 2024. arXiv:2406.11717
  23. Chen, R., Arditi, A., Sleight, H., Evans, O., Lindsey, J. Persona Vectors: Monitoring and Controlling Character Traits in Language Models. arXiv 2025. arXiv:2507.21509
  24. Rimsky, N., Gabrieli, N., Schulz, J., Tong, M., Hubinger, E., Turner, A. Steering Llama 2 via Contrastive Activation Addition. ACL 2024. arXiv:2312.06681
  25. Bayat, R., Rahimi-Kalahroudi, A., Pezeshki, M., Chandar, S., Vincent, P. Steering Large Language Model Activations in Sparse Spaces. arXiv 2025. arXiv:2503.00177
  26. Wang, A., Shu, D., Wang, Y., Ma, Y., Du, M. Improving LLM Reasoning through Interpretable Role-Playing Steering. Findings of EMNLP 2025. arXiv:2506.07335
  27. Yang, H., Cho, K., Ding, M., Inoue, T. (author list partly unverified) Task Vectors, Learned Not Extracted: Performance Gains and Mechanistic Insight. arXiv 2025. arXiv:2509.24169
  28. Han, S., Song, J., Gore, J., Agrawal, P. Emergence of Abstractions: Concept Encoding and Decoding Mechanism for In-Context Learning in Transformers. arXiv 2024. arXiv:2412.12276

Attention routing, instruction anchors, attention steering, and attention sinks

  1. Yuksekgonul, M., Chandrasekaran, V., Jones, E., et al. Attention Satisfies: A Constraint-Satisfaction Lens on Factual Errors of Language Models. ICLR 2024. arXiv:2309.15098
  2. Chuang, Y.-S., Qiu, L., Hsieh, C.-Y., Krishna, R., Kim, Y., Glass, J. Lookback Lens: Detecting and Mitigating Contextual Hallucinations in Large Language Models Using Only Attention Maps. EMNLP 2024. arXiv:2407.07071
  3. Zhang, Q., Singh, C., Liu, L., Liu, X., Yu, B., Gao, J., Zhao, T. Tell Your Model Where to Attend: Post-hoc Attention Steering for LLMs (PASTA). ICLR 2024. arXiv:2311.02262
  4. Guardieiro, V., Khare, A., Stein, A., Wong, E. Instruction Following by Principled Boosting Attention of Large Language Models (InstABoost). ICLR 2026. arXiv:2506.13734
  5. Venkateswaran, P., Contractor, D. Spotlight Your Instructions: Instruction-following with Dynamic Attention Steering. EACL 2026. arXiv:2505.12025
  6. Li, K., Liu, T., Bashkansky, N., Bau, D., Viégas, F., Pfister, H., Wattenberg, M. Measuring and Controlling Instruction (In)Stability in Language Model Dialogs. COLM 2024. arXiv:2402.10962
  7. Dongre, V., Hsieh, J., Lai, V. D., Yoon, S., Bui, T., Hakkani-Tür, D. When Attention Closes: How LLMs Lose the Thread in Multi-Turn Interaction. arXiv 2026. arXiv:2605.12922
  8. Zhang, Y., Xu, M., Bai, X., et al. Instruction Anchors: Dissecting the Causal Dynamics of Modality Arbitration. arXiv 2026. arXiv:2602.03677
  9. Wang, L., Li, L., Dai, D., et al. Label Words are Anchors: An Information Flow Perspective for Understanding In-Context Learning. EMNLP 2023. arXiv:2305.14160
  10. Hung, K.-H., Ko, C.-Y., Rawat, A., Chung, I.-H., Hsu, W. H., Chen, P.-Y. Attention Tracker: Detecting Prompt Injection Attacks in LLMs. arXiv 2024. arXiv:2411.00348
  11. Jin, Z., Cao, P., Yuan, H., et al. Cutting Off the Head Ends the Conflict: A Mechanism for Interpreting and Mitigating Knowledge Conflicts in Language Models. Findings of ACL 2024. arXiv:2402.18154
  12. Wu, W., Wang, Y., Xiao, G., Peng, H., Fu, Y. Retrieval Head Mechanistically Explains Long-Context Factuality. ICLR 2025. arXiv:2404.15574
  13. Zhang, W., Yin, F., Yen, H., Chen, D., Ye, X. Query-Focused Retrieval Heads Improve Long-Context Reasoning and Re-ranking. arXiv 2025. arXiv:2506.09944
  14. Xiao, G., Tian, Y., Chen, B., Han, S., Lewis, M. Efficient Streaming Language Models with Attention Sinks. ICLR 2024. arXiv:2309.17453
  15. Gu, X., Pang, T., Du, C., et al. When Attention Sink Emerges in Language Models: An Empirical View. ICLR 2025. arXiv:2410.10781
  16. Barbero, F., Arroyo, Á., Gu, X., et al. Why Do LLMs Attend to the First Token?. COLM 2025. arXiv:2504.02732
  17. Sun, M., Chen, X., Kolter, J. Z., Liu, Z. Massive Activations in Large Language Models. COLM 2024. arXiv:2402.17762
  18. McDougall, C., Conmy, A., Rushing, C., McGrath, T., Nanda, N. Copy Suppression: Comprehensively Understanding an Attention Head. arXiv 2023. arXiv:2310.04625
  19. Jain, S., Wallace, B. C. Attention is not Explanation. NAACL 2019. arXiv:1902.10186
  20. Wiegreffe, S., Pinter, Y. Attention is not not Explanation. EMNLP 2019. arXiv:1908.04626
  21. Abnar, S., Zuidema, W. Quantifying Attention Flow in Transformers. ACL 2020. arXiv:2005.00928
  22. Zeng, S., Lee, S., Zhao, H., Hockenmaier, J. Steering Instruction Hierarchies at Inference Time (V-Steer). COLM 2026. arXiv:2607.26228
  23. Balp-Straffon, E., Hsu, C.-H., Gadhvi, R., et al. How Language Models Choose Sides: Internal Representations of Instruction Hierarchy. ICML 2026 Mechanistic Interpretability Workshop. arXiv:2608.28648
  24. Zeng, S. Who is In Charge? Dissecting Role Conflicts in Instruction Following. arXiv 2025. arXiv:2510.01228
  25. Wu, T., et al. Instructional Segment Embedding: Improving LLM Safety with Instruction Hierarchy. ICLR 2025. arXiv:2410.09102
  26. Bai, Y., et al. Identifying and Analyzing Task-Encoding Tokens in Large Language Models. arXiv 2024. arXiv:2401.11323

The KV cache as instruction memory

  1. Li, Y., Huang, Y., Yang, B., et al. SnapKV: LLM Knows What You are Looking for Before Generation. NeurIPS 2024. arXiv:2404.14469
  2. Zhang, Z., Sheng, Y., Zhou, T., et al. H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models. NeurIPS 2023. arXiv:2306.14048
  3. Belitsky, M., Kopiczko, D. J., Dorkenwald, M., et al. KV Cache Steering for Controlling Frozen LLMs. arXiv 2025. arXiv:2507.08799
  4. Liu, A. Z., Zhang, M., Greenberg, I., et al. Memory Inception: Latent-Space KV Cache Manipulation for Steering LLMs. arXiv 2026. arXiv:2605.06225
  5. Gu, Z., Yao, J., Du, K., Jiang, J. (author list from listing) LLMSteer: Improving Long-Context LLM Inference by Steering Attention on Reused Contexts. arXiv 2024. arXiv:2411.13009

Chain-of-thought mechanics and generation-stage dynamics

  1. Cabannes, V., Arnal, C., Bouaziz, W., Yang, A., Charton, F., Kempe, J. Iteration Head: A Mechanistic Study of Chain-of-Thought. NeurIPS 2024. arXiv:2406.02128
  2. Dutta, S., Singh, J., Chakrabarti, S., Chakraborty, T. How to Think Step-by-Step: A Mechanistic Understanding of Chain-of-Thought Reasoning. TMLR 2024. arXiv:2402.18312
  3. Stolfo, A., Belinkov, Y., Sachan, M. A Mechanistic Interpretation of Arithmetic Reasoning in Language Models using Causal Mediation Analysis. EMNLP 2023. arXiv:2305.15054
  4. Nikankin, Y., Reusch, A., Mueller, A., Belinkov, Y. Arithmetic Without Algorithms: Language Models Solve Math with a Bag of Heuristics. ICLR 2025. arXiv:2410.21272
  5. Liu, X., Wang, L. Answer Convergence as a Signal for Early Stopping in Reasoning. arXiv 2025. arXiv:2506.02536
  6. Afzal, A., Chechik, G., Matthes, F., Ziser, Y. Knowing Before Saying: LLM Representations Encode Information About Chain-of-Thought Success Before Completion. arXiv 2025. arXiv:2505.24362
  7. Arcuschin, I., Janiak, J., Krzyzanowski, R., Rajamanoharan, S., Nanda, N., Conmy, A. Chain-of-Thought Reasoning In The Wild Is Not Always Faithful. arXiv 2025. arXiv:2503.08679
  8. Lindsey, J., Gurnee, W., Ameisen, E., et al. On the Biology of a Large Language Model. Transformer Circuits Thread 2025. transformer-circuits.pub/2025/attribution-graphs/biology.html
  9. Ameisen, E., Lindsey, J., Pearce, A., et al. Circuit Tracing: Revealing Computational Graphs in Language Models. Transformer Circuits Thread 2025. transformer-circuits.pub/2025/attribution-graphs/methods.html
  10. Wang, B., Min, S., Deng, X., et al. Towards Understanding Chain-of-Thought Prompting: An Empirical Study of What Matters. ACL 2023. arXiv:2212.10001
  11. Madaan, A., Yazdanbakhsh, A. Text and Patterns: For Effective Chain of Thought, It Takes Two to Tango. arXiv 2022. arXiv:2209.07686
  12. Venhoff, C., Arcuschin, I., Torr, P., Conmy, A., Nanda, N. Understanding Reasoning in Thinking Language Models via Steering Vectors. arXiv 2025. arXiv:2506.18167
  13. Sun, C.-E., Yan, G., Weng, T.-W. ThinkEdit: Interpretable Weight Editing to Mitigate Overly Short Thinking in Reasoning Models. EMNLP 2025. arXiv:2503.22048
  14. Sinii, V., et al. (author list partly unverified) Small Vectors, Big Effects: A Mechanistic Study of RL-Induced Reasoning via Steering Vectors. arXiv 2025. arXiv:2509.06608
  15. Rocchetti, E., Ferrara, A. How Instruction-Tuning Imparts Length Control: A Cross-Lingual Mechanistic Analysis. arXiv 2025. arXiv:2509.02075
  16. Park, Y., Jeong, M., Kang, J. Thinking Sparks!: Emergent Attention Heads in Reasoning Models During Post Training. arXiv 2025. arXiv:2509.25758
  17. Bogdan, P. C., Macar, U., Nanda, N., Conmy, A. Thought Anchors: Which LLM Reasoning Steps Matter?. arXiv 2025. arXiv:2506.19143
  18. Zhang, J., Lin, Q., Rajmohan, S., Zhang, D. From Reasoning to Answer: Empirical, Attention-Based and Mechanistic Insights into Distilled DeepSeek R1 Models. arXiv 2025. arXiv:2509.23676
  19. Zhang, H., Tian, Y., Zhang, T. Attention-Aligned Reasoning for Large Language Models. arXiv 2025. arXiv:2510.03223
  20. Yan, S., Shen, C., Wang, W., Xie, L., Liu, J., Ye, J. Don't Take Things Out of Context: Attention Intervention for Enhancing Chain-of-Thought Reasoning in Large Language Models. arXiv 2025. arXiv:2503.11154
  21. Chen, X., Xu, J., Liang, T., et al. Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs. arXiv 2024. arXiv:2412.21187

Causal intervention, attribution and evaluation methodology

  1. Vig, J., Gehrmann, S., Belinkov, Y., et al. Investigating Gender Bias in Language Models Using Causal Mediation Analysis. NeurIPS 2020. arXiv:2004.12265
  2. Meng, K., Bau, D., Andonian, A., Belinkov, Y. Locating and Editing Factual Associations in GPT (ROME / causal tracing). NeurIPS 2022. arXiv:2202.05262
  3. Zhang, F., Nanda, N. Towards Best Practices of Activation Patching in Language Models: Metrics and Methods. ICLR 2024. arXiv:2309.16042
  4. Heimersheim, S., Nanda, N. How to Use and Interpret Activation Patching. arXiv 2024. arXiv:2404.15255
  5. Wang, K., Variengien, A., Conmy, A., Shlegeris, B., Steinhardt, J. Interpretability in the Wild: A Circuit for Indirect Object Identification in GPT-2 Small. ICLR 2023. arXiv:2211.00593
  6. Goldowsky-Dill, N., MacLeod, C., Sato, L., Arora, A. Localizing Model Behavior with Path Patching. arXiv 2023. arXiv:2304.05969
  7. Syed, A., Rager, C., Conmy, A. Attribution Patching Outperforms Automated Circuit Discovery. BlackboxNLP 2024. arXiv:2310.10348
  8. Kramár, J., Lieberum, T., Shah, R., Nanda, N. AtP*: An Efficient and Scalable Method for Localizing LLM Behaviour to Components. arXiv 2024. arXiv:2403.00745
  9. Hanna, M., Pezzelle, S., Belinkov, Y. Have Faith in Faithfulness: Going Beyond Circuit Overlap When Finding Model Mechanisms (EAP-IG). COLM 2024. arXiv:2403.17806
  10. Conmy, A., Mavor-Parker, A. N., Lynch, A., Heimersheim, S., Garriga-Alonso, A. Towards Automated Circuit Discovery for Mechanistic Interpretability (ACDC). NeurIPS 2023. arXiv:2304.14997
  11. Haklay, T., Orgad, H., Bau, D., Mueller, A., Belinkov, Y. Position-aware Automatic Circuit Discovery (PEAP). ACL 2025. arXiv:2502.04577
  12. Geva, M., Bastings, J., Filippova, K., Globerson, A. Dissecting Recall of Factual Associations in Auto-Regressive Language Models. EMNLP 2023. arXiv:2304.14767
  13. Chan, L., Garriga-Alonso, A., Goldowsky-Dill, N., et al. Causal Scrubbing: A Method for Rigorously Testing Interpretability Hypotheses. Alignment Forum 2022. www.alignmentforum.org/posts/JvZhhzycHu2Yd57RN/causal-scrubbing-a-method-for-rigorously-testing
  14. Geiger, A., Lu, H., Icard, T., Potts, C. Causal Abstractions of Neural Networks. NeurIPS 2021. arXiv:2106.02997
  15. Wu, Z., Geiger, A., Icard, T., Potts, C., Goodman, N. Interpretability at Scale: Identifying Causal Mechanisms in Alpaca (Boundless DAS). NeurIPS 2023. arXiv:2305.08809
  16. Ghandeharioun, A., Caciularu, A., Pearce, A., Dixon, L., Geva, M. Patchscopes: A Unifying Framework for Inspecting Hidden Representations of Language Models. ICML 2024. arXiv:2401.06102
  17. Cohen-Wang, B., Shah, H., Georgiev, K., Madry, A. ContextCite: Attributing Model Generation to Context. NeurIPS 2024. arXiv:2409.00729
  18. Cohen-Wang, B., Chen, Y., Madry, A. Learning to Attribute with Attention. arXiv 2025. arXiv:2504.13752
  19. Liu, F., Kandpal, N., Raffel, C. AttriBoT: A Bag of Tricks for Efficiently Approximating Leave-One-Out Context Attribution. ICLR 2025. arXiv:2411.15102
  20. Sarti, G., Feldhus, N., Sickert, L., van der Wal, O., Nissim, M., Bisazza, A. Inseq: An Interpretability Toolkit for Sequence Generation Models. ACL 2023 Demo. arXiv:2302.13942
  21. Yin, K., Neubig, G. Interpreting Language Models with Contrastive Explanations. EMNLP 2022. arXiv:2202.10419
  22. Achtibat, R., Hatefi, S. M. V., Dreyer, M., et al. AttnLRP: Attention-Aware Layer-wise Relevance Propagation for Transformers. ICML 2024. arXiv:2402.05602
  23. Ferrando, J., Voita, E. Information Flow Routes: Automatically Interpreting Language Models at Scale. EMNLP 2024. arXiv:2403.00824
  24. Marks, S., Rager, C., Michaud, E. J., Belinkov, Y., Bau, D., Mueller, A. Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models. ICLR 2025. arXiv:2403.19647
  25. Makelov, A., Lange, G., Nanda, N. Is This the Subspace You Are Looking for? An Interpretability Illusion for Subspace Activation Patching. ICLR 2024. arXiv:2311.17030
  26. Miller, J., Chughtai, B., Saunders, W. Transformer Circuit Faithfulness Metrics are not Robust. COLM 2024. arXiv:2407.08734
  27. Shi, C., Beltran-Velez, N., Nazaret, A., et al. Hypothesis Testing the Circuit Hypothesis in LLMs. NeurIPS 2024. arXiv:2410.13032
  28. Pres, I., Ruis, L., Lubana, E. S., Krueger, D. Towards Reliable Evaluation of Behavior Steering Interventions in LLMs. arXiv 2024. arXiv:2410.17245
  29. Ortu, F., Jin, Z., Doimo, D., et al. Competition of Mechanisms: Tracing How Language Models Handle Facts and Counterfactuals. ACL 2024. arXiv:2402.11655
  30. Dotsinski, A., et al. On the Generalizability of "Competition of Mechanisms" (reproduction study). arXiv 2025. arXiv:2506.22977
  31. Wu, Z., Geiger, A., Arora, A., et al. pyvene: A Library for Understanding and Improving PyTorch Models via Interventions. NAACL 2024 Demo. arXiv:2403.07809
  32. Fiotto-Kaufman, J., Loftus, A. R., Todd, E., et al. NNsight and NDIF: Democratizing Access to Foundation Model Internals. arXiv 2024. arXiv:2407.14561
  33. Bhaskar, A., Wettig, A., Friedman, D., Chen, D. Finding Transformer Circuits with Edge Pruning. NeurIPS 2024. arXiv:2406.16778
  34. McGrath, T., Rahtz, M., Kramár, J., Mikulik, V., Legg, S. The Hydra Effect: Emergent Self-repair in Language Model Computations. arXiv 2023. arXiv:2307.15771
  35. Rushing, C., Nanda, N. Explorations of Self-Repair in Language Models. ICML 2024. arXiv:2402.15390
  36. Wei, J., Bosma, M., Zhao, V., et al. Finetuned Language Models Are Zero-Shot Learners. ICLR 2022. arXiv:2109.01652
  37. Ouyang, L., Wu, J., Jiang, X., et al. Training Language Models to Follow Instructions with Human Feedback. NeurIPS 2022. arXiv:2203.02155
  38. Rafailov, R., Sharma, A., Mitchell, E., et al. Direct Preference Optimization: Your Language Model is Secretly a Reward Model. NeurIPS 2023. arXiv:2305.18290
  39. Lambert, N., Morrison, J., Pyatkin, V., et al. Tülu 3: Pushing Frontiers in Open Language Model Post-Training. arXiv 2024. arXiv:2411.15124
  40. Sarkar, S., Feng, S., Ney, J., et al. Designing Role Vectors to Improve LLM Inference Behaviour. arXiv 2025. arXiv:2502.12055
  41. Tan, D., Chanin, D., Lynch, A., et al. Analysing the Generalisation and Reliability of Steering Vectors. NeurIPS 2024. arXiv:2407.12404
  42. Wu, Z., Arora, A., Geiger, A., et al. AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders. ICML 2025. arXiv:2501.17148
  43. Kang, D., Liu, Z., Ma, N., Huang, Y., Tan, Z., Jiang, M. Prompt-Activation Duality: Improving Activation Steering via Attention-Level Interventions. arXiv 2026. arXiv:2605.10664
  44. Rahn, N., D’Oro, P., Bellemare, M. Controlling Large Language Model Agents with Entropic Activation Steering. arXiv 2024. arXiv:2406.00244
  45. Shkolnikov, Y. P. Agent Memory Below the Prompt: Persistent Q4 KV Cache for Multi-Agent LLM Inference on Edge Devices. arXiv 2026. arXiv:2603.04428
  46. Braun, J., Eickhoff, C., Krueger, D., Bahrainian, S. A., Krasheninnikov, D. Understanding (Un)Reliability of Steering Vectors in Language Models. arXiv 2025. arXiv:2505.22637
  47. Cunningham, H., Ewart, A., Riggs, L., Huben, R., Sharkey, L. Sparse Autoencoders Find Highly Interpretable Features in Language Models. ICLR 2024. arXiv:2309.08600
  48. Bricken, T., Templeton, A., Batson, J., et al. Towards Monosemanticity: Decomposing Language Models With Dictionary Learning. Transformer Circuits Thread 2023. transformer-circuits.pub
  49. Templeton, A., Conerly, T., Marcus, J., et al. Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet. Transformer Circuits Thread 2024. transformer-circuits.pub
  50. Lieberum, T., Rajamanoharan, S., Conmy, A., et al. Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2. BlackboxNLP 2024. arXiv:2408.05147