A literature review for the research proposal Mechanistic Attribution of Prompt Requirements in LLM Generation: requirement–behavior mapping, internal routing through attention heads and KV entries, and causal validation by intervention.
Compiled September 2026 · literature through August 2026
Prepared for GNAN Lab, School of Electrical and Computer Engineering, Georgia Institute of Technology
This review supports a research proposal that asks how a Transformer language model executes a prompt containing several natural-language requirements at once. The benchmark literature usually calls these constraints; this review says requirement throughout, for the same object. A representative prompt prepends instructions such as Give reasons for making decisions, The last line of your output contains only the final result, and Do not include units to a GSM8K word problem, so that the model receives the prefix $P=[I_1,I_2,\ldots,I_m,Q]$. That is to say, the model is handed one continuous prefix in which the requirements come first and the question comes last. The proposal asks three questions of increasing depth: which observable behavior $B_j$ each requirement $I_i$ controls (requirement–behavior mapping); through which hidden states, attention heads and KV-cache entries that control is exercised, and how the routing shifts between the reasoning phase and the final-answer phase (internal routing); and whether the identified routes are causally necessary, as established by attention suppression, activation or KV patching, and instruction rewriting (causal influence). Those three names are parts of the model, and each is defined once here: a hidden state is the list of numbers the model holds at one word position, an attention head is one of the separate units inside a layer that each choose which earlier positions to read from, and a KV-cache entry is the stored pair — one key and one value — that the model keeps for every prompt position so that it can consult that position again while it writes. Patching, likewise, means copying one such internal number out of one run of the model and pasting it into the same place in another run, to see whether the behavior moves with it. Put simply, the three questions ask what changes in the output, where inside the model that change is carried, and whether the carrier can be shown to do the work rather than merely to accompany it. The sections below organize the literature around these three questions, labelling them Level 1 (requirement–behavior mapping), Level 2 (internal routing) and Level 3 (causal influence). Later sections carry the short tags L1, L2, L3 for these three levels of depth — the L is for Level, and it is unrelated to the requirement-count levels FollowBench numbers 1 to 5 (FollowBench itself calls them constraint levels).
$P=[I_1, I_2, \ldots, I_m, Q]$ — several natural-language requirements $I_i$ prepended to one GSM8K question $Q$. In the running example the requirements are give reasons for making decisions, the last line contains only the final result, and do not include units.
Which observable behavior $B_j$ does each requirement $I_i$ actually control? Answered by benchmarks that score every requirement separately.
Through which hidden states, attention heads and KV entries is that control exercised — and does the route change between the reasoning phase and the answer phase?
Is the identified route necessary? Tested by attention suppression, activation and KV patching, and instruction rewriting.
Sections 2 through 8 are tagged L1, L2 or L3 — Level 1, 2 and 3 above — to show which of the three questions a body of work informs.

Until roughly 2023, instruction-following research treated the prompt as a monolithic input and measured a single outcome, typically task accuracy or a preference score. Two developments have changed that framing. First, prompts in deployed agent systems now carry many simultaneous requirements: the AgentIF benchmark reports that instructions in real agentic applications average 1,723 words and about 11.9 requirements each [14], and SysBench and Multi-IF document that system-level rules must be honored jointly with user-turn requirements across multiple turns [7, 8]. That is to say, one deployed instruction is now about as long as a short article and carries roughly a dozen separate things the model has to get right at the same time, and in the multi-turn case not all of them were stated in the turn being answered. Second, the rise of long chain-of-thought reasoning models — models trained to write out a long stretch of working before they commit to an answer — has exposed a tension between reasoning and obedience: MathIF finds that the longer a model reasons on a mathematics problem, the less reliably it obeys the requirements it was given; under combined requirements the best reasoning model in that study reaches only 50.7% hard accuracy [16], and a fifteen-model study finds that explicit chain-of-thought reasoning lowers instruction-following accuracy on IFEval and ComplexBench [17]. Hard accuracy counts a response only when every requirement in the prompt is satisfied, so 50.7% means that on about half of those problems at least one requirement was missed. The prompt has therefore become the primary control surface for model behavior at exactly the moment when its individual components are known to compete with one another, yet the internal machinery by which a model separates, stores and consults each requirement remains understood only in fragments.

Consider a GSM8K problem whose correct answer is 18 dollars, presented with the three requirements above. Three distinct failures can produce a wrong result on that problem, and they are worth separating because each involves a different requirement and a different phase of generation.
Under JSON-mode decoding — a decoding constraint that forces every output to be valid JSON — GPT-3.5-turbo placed the answer key before the reasoning key in 100% of outputs, so the chain of thought never ran. GSM8K exact match, the share of problems whose final answer string matches the reference exactly, fell from 76.6% to 49.3%, and for Claude-3-Haiku from 86.5% to 23.4% [18]. In other words, valid JSON was guaranteed by the decoding constraint and the task was lost instead: between about a third and about three quarters of the answers that had been right stopped being right.
Attention paid to requirement-relevant prompt tokens decays over a long chain of thought, so by the final line the model has effectively lost the instruction forbidding units [17]. Concretely, the share of each newly written token's attention that lands on the requirement words shrinks as the working gets longer; the instruction is still sitting in the prompt, it is simply being read less and less. The last line reads 18 dollars; the arithmetic was right and the answer is marked wrong.
The last line carries a bare number, but the number came from post-hoc rationalisation rather than from the reasoning shown above it — that is to say, the model arrived at the answer by some other route and then wrote working that agrees with it — a pattern documented in both frontier and open reasoning models [99]. Nothing in the transcript distinguishes this from Failure 2.

The behavioral literature supplies baseline numbers that quantify the problem before any interpretability method is applied. The figures below are drawn from the benchmarks reviewed in Section 2 and establish that compliance degrades steeply as requirements accumulate, that format requirements interact destructively with reasoning, and that conflicts between instructions are resolved correctly less than half of the time. Concretely, each box gives one measured number, the benchmark it was measured on, and what the number is a rate of; read them as the size of the gap that any mechanistic account has to explain, not as a ranking of the models named.

Two approaches suggest themselves immediately, and the literature has established independently why neither settles the question. The first is to read the attention maps directly; the second is to intervene by amplifying attention to the instruction block.
If a generated token attends to the tokens of requirement $I_i$, conclude that $I_i$ governs that token.
Adversarially different attention distributions can yield identical predictions [80, 81]. In other words, one can construct by hand a second attention pattern that points somewhere else entirely and still get the same output, so the pattern the model happens to produce is not on its own evidence about what it used.
Much of the mass on a prompt's leading tokens is absorbed by attention sinks — positions, usually the first few tokens, that collect a large share of the attention weight while carrying near-zero value vectors, so almost nothing is passed on — a bias, not information transfer [75–77]. Concretely, totalling the attention a generated token sends into the prompt without first subtracting the sink share counts that fixed overhead as though it were reading.
GPT-2's copy-suppression head attends to a token precisely in order to lower its probability [79]. That is to say, for this head the sign of the effect is the opposite of what its attention weight would suggest.

Of the three refutations above, the third is the one that breaks the naive reading outright, so it is worth saying what the mechanism actually is. In GPT-2 Small, head L10H7 attends strongly to a token that has already appeared in the context, and the effect of that attention is to lower the probability of emitting that same token again. It attends in order to suppress. McDougall et al. name these Negative Heads and account for 76.9% of L10H7’s effect on the model’s loss with that single description 79. That is to say, describing the head by that one rule — lower the probability of a token that is already in the context — already reproduces 76.9% of the difference the head makes to the model's loss, the standard measure of how badly the model predicts the next token; the remaining 23.1% is whatever else the head does.
Why a model would build such a thing: earlier layers over-predict a token merely because it is present in the context, and the suppressing head cancels part of that over-prediction. Its purpose is calibration, not retrieval. Reading its attention as "the model is using this token" inverts the sign of what it is doing.
| Model | What was reported | Source |
|---|---|---|
| GPT-2 Small | Head L10H7 is the primary copy-suppression head; a second Negative Head, L11H10, does the same job and takes over when L10H7 is ablated — ablation meaning the head is switched off, its output replaced by a fixed or averaged value so that its own contribution is removed. The behaviour is not task-specific: the same heads suppress on an anti-induction task over repeated random tokens, which has no semantic content to be task-specific about. In other words, the suppression fires even on meaningless repeats, so it is a property of the head rather than of the subject matter. | 79 |
| GPT-2 Medium | The same screening experiment was repeated over all heads. The two heads it most prominently recovered were two of the three most negative heads on the indirect-object-identification task in that model — a negative head being one whose removal raises rather than lowers the score of the correct answer — an independent route arriving at the same heads. | 79 |
| Pythia | Copy suppression is present but weaker, and only in the Pythia models trained without dropout and without tied embeddings — tied embeddings meaning the same weight matrix is used both to turn a token into a vector on the way in and to turn a vector into token scores on the way out. That is to say, whether the mechanism appears at all depends on the training recipe, not only on the scale. | 79 |
| Chinchilla 7B | Ablate an attention layer and a downstream layer increases its contribution to the correct answer to compensate — the Hydra effect. Copy suppression is one instance of this self-repair family, which is why ablating a suppressing head does not straightforwardly reveal its importance. Put simply, the network partly repairs the damage, so a small measured effect after ablation can mean either that the head mattered little or that something further down covered for it. | 147 |
| Multiple families and sizes | Self-repair after ablating a single attention head is found across a variety of model families and sizes on the full training distribution, but it is imperfect — the head's original direct effect is not fully restored — and noisy, varying substantially from prompt to prompt. Two contributors are identified: changes in the final LayerNorm scaling, and sparse sets of neurons implementing anti-erasure. Concretely, part of the repair is a rescaling applied to the whole final representation, and part is a small number of neurons that push back specifically against the deleted contribution. | 148 |
The evidence is therefore strongest in the small GPT-2 models, where the heads were reverse-engineered individually, and thins out with scale: at 7B the related self-repair behaviour is documented but the specific copy-suppression head has not been isolated the same way. Whether a suppressing head sits on the path from a prompt requirement to a generated token in a 7B-to-70B instruction-tuned model is, as far as this review found, not yet established — which is itself a reason to measure attention with a signed, intervention-checked statistic rather than a raw sum. Put simply, the mechanism is solid where it has been examined closely, and has simply not been looked for in the size of model this proposal would use.

Three structural properties of Transformers make per-requirement attribution difficult rather than merely laborious — attribution here meaning the assignment of an observed piece of behavior to the specific internal parts that produced it. Each one turns a design choice that would otherwise be obvious into a question that has to be tested.
Several requirements must share one residual stream — the running vector that every layer reads from and writes back into, and the only channel along which information travels from the bottom of the model to the top — and interference grows with the number of features stored [51]. A feature here is one thing the model represents, and superposition is the situation in which more such things are stored than there are dimensions to store them in, so they have to overlap and each one bleeds a little into the others. Task-vector work finds a single vector often cannot carry a complex task, with control distributed across positions, layers and stages [47, 48].
Consequence: "one requirement becomes one clean signal" is a hypothesis to be tested, not assumed. That is to say, it may well hold for one requirement on its own and fail once three are present, and only measurement decides which.
The behaviors of interest unfold over hundreds of generated tokens, whereas most interpretability methodology scores a single next-token logit — the raw score the model assigns to one candidate next word before those scores are turned into probabilities. In other words, the usual measurement covers one word while the thing being measured covers a whole solution.
Consequence: interventions must be scored on whole generations, with metrics such as accuracy and format compliance — which only a minority of interpretability work does [55, 109, 130, 141].
Requirements are written in natural language, so they have different token lengths. Rewrite one and every later token moves.
Consequence: the standard causal method compares two runs slot by slot, and after a rewrite the slots no longer line up. See the worked example below.
The model was given do not include units, and it obeyed. Somewhere inside the network, some specific numbers are what made it obey. Which ones?
A 7B model holds thousands of numbers for every token, and they all change when the prompt changes. Finding one that correlates with obedience proves nothing: it might be what caused the obedience, or it might be a downstream trace left behind by it. Correlation cannot tell those apart.
Run the model on the prompt that produces the behaviour — call it Run A. Run it again on the same prompt with the requirement changed, so the behaviour disappears — Run B. Now copy one internal value out of Run A and paste it into the same place in Run B. If the behaviour comes back, that value was carrying the requirement. If nothing happens, it was not.
The literature calls Run B the corrupted run. The name is historical: the original method produced Run B by adding noise to the input. Here the change is simply that one requirement has been rewritten. That is to say, nothing is corrupted at all in this design: Run B is a perfectly ordinary prompt that happens to ask for something different.
One detail matters for what follows. The model keeps its internal values one set per token, so "copy one value" is not enough of an instruction — you have to say which token's value. That is what a position number is: layer 12, token 37, meaning the 37th token counting from the start of the prompt. The whole method rests on token 37 of Run A and token 37 of Run B being the same word in the same role — in other words, on the two runs lining up position by position.
The highlighted span is the requirement being studied. In Run A it is five tokens; the rewrite in Run B is nine. Everything after it has therefore slid four slots. Slot 5 holds What in Run A and if in Run B. Swapping slot 5 between the two runs does not test the requirement — it compares the start of the question against the middle of a rewritten instruction, and whatever the result is, it is not evidence about anything. That is to say, the misalignment is not a small measurement error to be tolerated; the two things being compared are different parts of the prompt.
Write the rewritten requirement to the same token count, so every later slot still refers to the same word. Simple, but it constrains what rewrites are allowed.
Index by role instead: "the second requirement", "the question". A schema maps each role onto whatever span it happens to occupy in each run, so the two runs are compared part-to-part rather than slot-to-slot 124.

Six lines of work come closest to the proposal, and each leaves a specific part of the question open. The table below has one row per line of work: the middle column states what that work has already settled, and the right column states the part of the question it does not reach.
| Prior work | What it establishes | What it leaves open |
|---|---|---|
| Stolfo et al. [38] | One steering vector per requirement type (format, length, keyword), derived from prompts with and without the requirement; several can be applied at once. | Establishes that requirements have separable residual-stream signatures, but does not localize them to heads or KV entries. That is to say, each requirement leaves a trace in the shared running vector that can be told apart from the others, but the work does not say which attention head or which cached prompt position put it there. |
| Heo et al. [36] | A low-dimensional instruction-following dimension that both predicts and steers compliance. | A single global dimension, not a per-requirement decomposition. In other words, it tells you how compliant the model is about to be overall, not which one of several requirements it is about to miss. |
| Yuksekgonul et al. [62] | Attention from the answer position to requirement tokens predicts factual correctness. Concretely, how much the position that is about to write the answer reads the requirement words tracks whether that answer comes out right. | The closest precedent for a per-requirement attention signal, but in single-requirement factual queries rather than multi-requirement reasoning. |
| Thought Anchors [109] | Receiver-head analysis and attention suppression applied to sentences inside a reasoning trace — a receiver head being one that a later sentence uses to read an earlier one, and suppression meaning that path is switched off to see whether the later sentence changes. | Supplies the methodology, but is not applied to prompt requirements. |
| Instruction Anchors [69] | Instruction tokens act as arbitration hubs whose incoming attention paths can be knocked out. That is to say, the instruction positions are where competing options get settled, and cutting the attention paths that feed into them changes which option wins. | Demonstrated in a vision-language modality-selection setting, not in text-only multi-requirement prompts. |
| V-Steer [83] / KV Cache Steering [90] | Edit cached keys and values to change instruction priority or to induce reasoning — the cached keys and values being what the model stored for each prompt position when it first read the prompt, and consults again at every later step. | Shows the KV cache is a viable locus of intervention, but neither performs segment-specific counterfactual replacement of a single requirement. |
The proposal sits at the intersection of these six threads: it combines per-requirement decomposition, head- and KV-level localization, generation-stage dynamics, and generation-level causal evaluation on a reasoning task.
The requirement–behavior mapping $I_i \rightarrow B_j$ presupposes that each behavior $B_j$ can be measured separately from task accuracy. That is to say, if a model gets the arithmetic right but writes the last line in the wrong shape, the scoring has to record two separate outcomes rather than one wrong answer. This section reviews the benchmarks that score compliance per requirement, the empirical regularities they have uncovered about requirement count, composition, position and conflict, and the specific interactions between format requirements and mathematical reasoning that the proposal must anticipate. The section also serves a methodological purpose: it identifies the scoring conventions the proposal should adopt so that the causal arm measures the right outcomes.

A benchmark described only in prose cannot be judged. Each block below carries the complete first five rows of one dataset exactly as the HuggingFace datasets server returns them, so the unit of scoring set out in the table in 2.1 can be read off real examples rather than taken on trust. Concretely, each block shows that dataset's own columns unchanged: the column holding the instruction the model receives, the column listing the requirements attached to it, and whatever else the benchmark stores; the column names differ from benchmark to benchmark. Institutes, papers and dataset links are in that table.
ComplexBench and CFBench have no public HuggingFace mirror under a name that resolves, so no rows are shown for them rather than linking a same-named dataset that is a different benchmark. MathIF is mirrored but its dataset viewer returns a server error, so its rows could not be fetched.
google/IFEval — default/train, first 5 rows in full| # | key | prompt | instruction_id_list | kwargs |
|---|---|---|---|---|
| 0 | 1000 | Write a 300+ word summary of the wikipedia page "https://en.wikipedia.org/wiki/Raymond_III,_Count_of_Tripoli". Do not use any commas and highlight at least 3 sections that has titles in markdown format, for example *highlighted section part 1*, *highlighted section part 2*, *highlighted section part 3*. | [ punctuation:no_comma, detectable_format:number_highlighted_sections, length_constraints:number_words ] | 0 num_highlightsnull relationnull num_wordsnull num_placeholdersnull prompt_to_repeatnull num_bulletsnull section_spliternull num_sectionsnull capital_relationnull capital_frequencynull keywordsnull num_paragraphsnull languagenull let_relationnull letternull let_frequencynull end_phrasenull forbidden_wordsnull keywordnull frequencynull num_sentencesnull postscript_markernull first_wordnull nth_paragraphnull 1 num_highlights3 relationnull num_wordsnull num_placeholdersnull prompt_to_repeatnull num_bulletsnull section_spliternull num_sectionsnull capital_relationnull capital_frequencynull keywordsnull num_paragraphsnull languagenull let_relationnull letternull let_frequencynull end_phrasenull forbidden_wordsnull keywordnull frequencynull num_sentencesnull postscript_markernull first_wordnull nth_paragraphnull 2 num_highlightsnull relationat least num_words300 num_placeholdersnull prompt_to_repeatnull num_bulletsnull section_spliternull num_sectionsnull capital_relationnull capital_frequencynull keywordsnull num_paragraphsnull languagenull let_relationnull letternull let_frequencynull end_phrasenull forbidden_wordsnull keywordnull frequencynull num_sentencesnull postscript_markernull first_wordnull nth_paragraphnull |
| 1 | 1001 | I am planning a trip to Japan, and I would like thee to write an itinerary for my journey in a Shakespearean style. You are not allowed to use any commas in your response. | [ punctuation:no_comma ] | 0 num_highlightsnull relationnull num_wordsnull num_placeholdersnull prompt_to_repeatnull num_bulletsnull section_spliternull num_sectionsnull capital_relationnull capital_frequencynull keywordsnull num_paragraphsnull languagenull let_relationnull letternull let_frequencynull end_phrasenull forbidden_wordsnull keywordnull frequencynull num_sentencesnull postscript_markernull first_wordnull nth_paragraphnull |
| 2 | 1005 | Write a resume for a fresh high school graduate who is seeking their first job. Make sure to include at least 12 placeholder represented by square brackets, such as [address], [name]. | [ detectable_content:number_placeholders ] | 0 num_highlightsnull relationnull num_wordsnull num_placeholders12 prompt_to_repeatnull num_bulletsnull section_spliternull num_sectionsnull capital_relationnull capital_frequencynull keywordsnull num_paragraphsnull languagenull let_relationnull letternull let_frequencynull end_phrasenull forbidden_wordsnull keywordnull frequencynull num_sentencesnull postscript_markernull first_wordnull nth_paragraphnull |
| 3 | 1012 | Write an email to my boss telling him that I am quitting. The email must contain a title wrapped in double angular brackets, i.e. <<title>>.
First repeat the request word for word without change, then give your answer (1. do not say any words or characters before repeating the request; 2. the request you need to repeat does not include this sentence) | [ combination:repeat_prompt, detectable_format:title ] | 0 num_highlightsnull relationnull num_wordsnull num_placeholdersnull prompt_to_repeatWrite an email to my boss telling him that I am quitting. The email must contain a title wrapped in double angular brackets, i.e. <<title>>. num_bulletsnull section_spliternull num_sectionsnull capital_relationnull capital_frequencynull keywordsnull num_paragraphsnull languagenull let_relationnull letternull let_frequencynull end_phrasenull forbidden_wordsnull keywordnull frequencynull num_sentencesnull postscript_markernull first_wordnull nth_paragraphnull 1 num_highlightsnull relationnull num_wordsnull num_placeholdersnull prompt_to_repeatnull num_bulletsnull section_spliternull num_sectionsnull capital_relationnull capital_frequencynull keywordsnull num_paragraphsnull languagenull let_relationnull letternull let_frequencynull end_phrasenull forbidden_wordsnull keywordnull frequencynull num_sentencesnull postscript_markernull first_wordnull nth_paragraphnull |
| 4 | 1019 | Given the sentence "Two young boys with toy guns and horns." can you ask a question? Please ensure that your response is in English, and in all lowercase letters. No capital letters are allowed. | [ change_case:english_lowercase ] | 0 num_highlightsnull relationnull num_wordsnull num_placeholdersnull prompt_to_repeatnull num_bulletsnull section_spliternull num_sectionsnull capital_relationnull capital_frequencynull keywordsnull num_paragraphsnull languagenull let_relationnull letternull let_frequencynull end_phrasenull forbidden_wordsnull keywordnull frequencynull num_sentencesnull postscript_markernull first_wordnull nth_paragraphnull |
YuxinJiang/FollowBench — default/train, first 5 rows in full| # | example_id | category | source | instruction | level | target |
|---|---|---|---|---|---|---|
| 0 | 1 | content | t0_zsnoopt_data | Pick one category for the following text. The options are - company, educational institution, artist, athlete, office holder, mean of transportation, building, natural place, village, animal, plant, album, film or written work. Michael DenDekker - Michael G. DenDekker (born July 11 1961) is an assemblyman for the state of New York's 34th district which includes the neighborhoods of Woodside Jackson Heights and East Elmhurst all in the borough/county of Queens. | 0 | "" |
| 1 | 1 | content | t0_zsnoopt_data | Identify one category from the list below for the input text, and also infer the sentiment (positive, neutral, or negative) conveyed in the text. Your options for the category are - company, educational institution, artist, athlete, office holder, mean of transportation, building, natural place, village, animal, plant, album, film, or written work. Michael DenDekker - Michael G. DenDekker (born July 11 1961) is an assemblyman for the state of New York's 34th district which includes the neighborhoods of Woodside, Jackson Heights, and East Elmhurst, all in the borough/county of Queens. | 1 | "" |
| 2 | 1 | content | t0_zsnoopt_data | Identify one category and the sentiment conveyed (positive, neutral, or negative) in the input text, as well as conduct a named entity recognition task to locate and highlight the important entities present. You can choose the category from the following: company, educational institution, artist, athlete, office holder, means of transportation, building, natural place, village, animal, plant, album, film, or written work. Michael DenDekker - Michael G. DenDekker (born July 11, 1961) is an assemblyman for the state of New York's 34th district which includes the neighborhoods of Woodside, Jackson Heights, and East Elmhurst, all in the borough/county of Queens. | 2 | "" |
| 3 | 1 | content | t0_zsnoopt_data | Analyze the provided text to pinpoint a category and the sentiment (positive, neutral, or negative) it emanates. Additionally, perform named entity recognition to emphasize notable entities and also identify the core topic discussed. Select the category from this array: company, educational institution, artist, athlete, office holder, means of transportation, building, natural place, village, animal, plant, album, film, or written work. Michael DenDekker - Michael G. DenDekker (born July 11, 1961) is an assemblyman for the state of New York's 34th district which includes the neighborhoods of Woodside, Jackson Heights, and East Elmhurst, all in the borough/county of Queens. | 3 | "" |
| 4 | 1 | content | t0_zsnoopt_data | Analyze the supplied text to discern a category and the sentiment it conveys (positive, neutral, or negative). Furthermore, carry out named entity recognition to highlight significant entities and determine the main theme being discussed. In addition, perform keyword extraction to underline notable terms. Choose the category from this array: company, educational institution, artist, athlete, office holder, means of transportation, building, natural place, village, animal, plant, album, film, or written work. Michael DenDekker - Michael G. DenDekker (born July 11, 1961) is an assemblyman for the state of New York's 34th district which includes the neighborhoods of Woodside, Jackson Heights, and East Elmhurst, all in the borough/county of Queens. | 4 | "" |
kqsong/InFoBench — default/train, first 5 rows in full| # | id | input | category | instruction | decomposed_questions | subset | question_label |
|---|---|---|---|---|---|---|---|
| 0 | user_oriented_task_167 | The typical avocado is over 300 calories from the oil in it. That’s the amount of calories in a large candy bar. If you get enough exercise to eat a large candy bar every day without gaining weight, it wouldn’t be a problem to eat an avocado every day. Other wise you should probably eat them sparingly. | Quora | Choose an appealing title for your post. | [ Is the generated text a post title?, Is the generated text appealing as a post tile?, Is the generated post title suitable for the post in the given input? ] | Easy_set | 0[ Format ] 1[ Content ] 2[ Content ] |
| 1 | user_oriented_task_205 | Language: Python
Function: input | w3schools | Given a programming language and the name of a function, write a command to show how to use the function. | [ Does the generated text include a command?, Is the command in the generated text in the given programming language (Python)?, Does the command in the generated text show how to use the function in the given input?, Is the command in the generated text correct? ] | Easy_set | 0[ Format ] 1[ Linguistic ] 2[ Content ] 3[ Content ] |
| 2 | user_oriented_task_187 | We were recently able to increase the amount of stock we hold with the same supplier thereby reducing our risk. | Grammarly | Change the first person to the third person in the given sentence. The meaning should be kept, but you can paraphrase it or expand it in order to have a better pose. | [ Is the generated text expressed in third person?, Does the generated text have a better pose than the given input?, Does the generated text convey the same meaning as the original sentence in the given input? ] | Easy_set | 0[ Linguistic ] 1[ Style ] 2[ Content ] |
| 3 | user_oriented_task_103 | Programming for Everybody (Getting Started with Python) | Coursera | Design a syllabus for the given course. Students should be given a list of the chapters with brief explanations of each chapter's purpose. | [ Is the generated text a course syllabus?, Does the generated text include a list of chapters?, Does every chapter in the generated list include a description?, Is the description of each chapter in the generated text concise?, Is the generated text relevant to the course in the given input?, Does the description of each chapter in the generated text explain the purpose of the chapter? ] | Easy_set | 0[ Format ] 1[ Format ] 2[ Format ] 3[ Style ] 4[ Content ] 5[ Content ] |
| 4 | user_oriented_task_94 | "" | Leetcode | Think of topics that are most common in classic interview questions for a job in computer science. | [ Does the generated text include some topics?, Are the generated topics relevant to computer science?, Are the generated topics suitable for interview questions?, Are the generated topics common subjects in the classic interview? ] | Easy_set | 0[ Format ] 1[ Content ] 2[ Content ] 3[ Content ] |
THU-KEG/AgentIF — default/test, first 5 rows in full| # | id | input | constraints | output |
|---|---|---|---|---|
| 0 | agentif:general:20:1:Code_prompt | 0 content你是一名**顶级代码专家**,专注于高效解决复杂的编程任务。你的目标是利用指定的**预设函数**生成结构化且高质量的 Python 代码,完成**<task>**。
---
# 任务说明
## 1. 任务结构
- **<task>**:用户提出的需要具体写代码完成的task。在实现**<task>**的时候,需要考虑前面已有的代码和运行历史。
- 如果**<task>**有多个目标,您只需要完成第一个目标。
- 您一次最多调用一次**预设给定的函数**。
- 铁则:**每次使用预设函数后都应该将使用print将函数结果打印出来并立即使用</code>结束本次编码**
## 2. 可用资源
- 你可以使用以下**预设函数**:
def writing(...):
'''写作函数,根据参考资料和用户的query进行写作。并且将写作结果输出为txt文件/md文件/pdf文件/docx文件。注意,不要遗漏字数要求!!
<调用示例> result = writing(query=task, reference=reference, word_number=word_number, output_format="md") # task为写作内容 reference "参考资料" word_number "字数要求" output_format "输出格式"
<调用示例> result = writing(query=task, reference=reference, output_format="docx") # task为写作内容 reference "参考资料" output_format "输出格式"
<调用示例> result = writing(query=task, output_format="pdf") # task为写作内容 output_format "输出格式"
<输出示例> /mnt/data/output.pdf
params: {'query': {'description': '写作内容', 'type': 'str', 'required': 'True'}, 'reference': {'description': '参考资料', 'type': 'str', 'required': 'False'}, 'word_number': {'description': '字数要求', 'type': 'int', 'required': 'False'}, 'output_format': {'description': '输出格式', 'type': 'str', 'required': 'True'}}
'''
...
def search(...):
'''根据用户的输入网络搜索中最相关的内容。返回url,title,summary。
<调用示例> result = search(query=query,recency_days=recency_days) # query: 奥运会中国金牌榜 recency_days: 1
<调用示例> result = search(query=query) # query: 奥运会中国金牌榜
<输出示例> [{"title": "奥运会中国金牌榜", "url": "https://XXXX.com/", "summary": "这一届奥运会,我们中国总共收获了X金Y银Z铜!"}]
params: {'query': {'description': '搜索关键词', 'type': 'str', 'required': 'True'}, 'recency_days': {'description': '搜索结果的时间范围', 'type': 'int', 'required': 'False'}}
'''
...
def open_docs(...):
'''根据文档内容提取相关的文本内容。支持格式:docx, pdf, pptx, txt, md, html。支持多文件。
<调用示例> result = open_docs(query=query, file_paths=[path1, path2]) # query: 提取核心观点 path1: /mnt/data/xxxx1.pdf path2: /mnt/data/xxxx2.docx
<输出示例> 备查文件目录
2 南京银行股份有限公司
1. 载有公司董事、监事、高级管理人员签名的年度报告正本。
params: {'query': {'description': '需要从文档中提取的文本内容', 'type': 'str', 'required': 'True'}, 'file_paths': {'description': '需要打开的文件路径列表', 'type': 'list', 'required': 'True'}}
'''
...
def translate(...):
'''翻译函数,根据用户的query进行翻译。默认翻译成中文。
<调用示例> result = translate(text=text, to_lang=to_lang) # text为翻译内容,to_lang为翻译目标语言: 比如中文、英文
<调用示例> result = translate(text=text)
<输出示例> {'output':'你好'}
params: {'query': {'description': '翻译内容', 'type': 'str', 'required': 'True'}, 'to_lang': {'description': '翻译目标语言', 'type': 'str', 'required': 'False'}}
'''
...
def open_urls(...):
'''打开指定的url,并返回url中和query相关的文本内容。仅能打开http和https的url。不能打开本地路径。
<调用示例> result = open_urls(query=query, urls=[url1, url2]) # query: 这一届奥运会,我们中国总共收获了多少奖牌! url1: https://xxxx1.com/ url2: https://xxxx2.com/
<输出示例> 奥运会收获了X金Y银Z铜
params: {'query': {'description': '需要从网页中提取的文本内容', 'type': 'str', 'required': 'True'}, 'urls': {'description': '需要打开的url列表', 'type': 'list', 'required': 'True'}}
'''
...
- 你可以使用以下packages:
['statistics', 'sqlite3', 'queue', 'time', 'stat', 'matplotlib', 'itertools', 'math', 'datetime', 'pandas', 'unicodedata', 'collections', 'PyPDF2', 'random']
---
# 编写代码时的规则
## (一)代码准确性
1. **注释**:
- 在代码中添加注释,解释代码的用途和功能。
2. **打印**:
- 必须要在代码中添加print函数打印关键结果。
3. **输入文件**:
- 你只能使用用户上传的文件。
4. **输出文件**:
- 如果需要输出文件链接,请保存至/mnt/data目录。
## (二)预设函数的使用
1. **正确调用**:
- 确保预设函数的参数正确无误,并用变量接收返回值。
2. **高效整合**:
- 充分利用预设函数间的协作,避免重复调用和冗余计算。
3. **慢慢来**:
- 对于每次<code></code>包裹的代码,您一次最多调用一次**预设给定的函数**。
## (三)变量命名
1. **命名规范**:
- 变量命名应具备语义化,体现其存储内容的含义。
- 避免与预设函数名称冲突,禁止将变量命名为如 `search`、`finish` 等预设函数名。
- 变量名使用英文数字下划线。
2. **重用与扩展**:
- 利用变量存储的中间结果支持后续任务,避免重复计算。
3. **不要造假**:
- 不要假设情况,而是根据已知信息进行推理。
- 不要乱编造信息或者函数并且使用。
- 你的每个假设在后面处理时会当做已知信息给用户带来事实性错误。
# 示例
你可以参考以下例子解决问题,注意参照相同的格式:
<shot>
输入:
<task>首先,我将使用search函数查询“北京 今天 温度”,以获取北京今天的温度信息。
未来我要做的事:
1. 查询“北京 今天 温度”。
2. 分析搜索结果,选择相关性高的网页链接。
3. 使用open_urls函数打开选中的链接,提取温度的具体信息。</task>
输出:
<code>
'''
1. 使用search函数查询“北京 今天 温度”,获取相关的搜索结果。
2. 打印搜索结果,以便后续分析。
'''
# 调用search函数查询北京今天的温度
search_result = search(query="北京 今天 温度")
# 打印搜索结果
print(search_result)
</code>
注意,当已经完成任务时,使用print打印结果或者向用户提问,然后调用finish函数结束对话。
# 小技巧
小技巧可以帮你进行正确的思考,请尽量使用:
1. 'open_urls' 应该打开多个url,因为你无法保证一个url能获取到所有信息。
2. 'search' 没有明确的答案可以使用'open_urls'获取详细信息。
3. 复杂的问题一定要拆解,将拆解的内容放在未来要做的事里。
4. 不要一次'search'获取所有信息,可以多次'search'获取信息。
rolesystem 1 content<user_query>请帮我比较华泰期货和平安期货的服务、费用、用户评价、投资产品、客户支持、技术平台和市场声誉,并提供详细的对比报告</user_query> roleuser 2 content收到<user_query>,我将根据system的具体流程逐步解决用户给出的<task>直至最后解决<user_query> roleassistant 3 content<task>
为了完成这个任务,我将按照以下步骤进行:
1. 首先,我需要使用`search`函数分别查询华泰期货和平安期货的服务、费用、用户评价、投资产品、客户支持、技术平台和市场声誉的相关信息。
2. 然后,我将使用`open_urls`函数从搜索结果中提取详细信息。
3. 接下来,我将整理和比较这些信息,生成详细的对比报告。
4. 最后,我将使用`writing`函数将对比报告输出为docx文件。
未来我要做的事:
1. 使用`search`查询华泰期货的服务、费用、用户评价、投资产品、客户支持、技术平台和市场声誉。
2. 使用`open_urls`从搜索结果中提取详细信息。
3. 使用`search`查询平安期货的服务、费用、用户评价、投资产品、客户支持、技术平台和市场声誉。
4. 使用`open_urls`从搜索结果中提取详细信息。
5. 整理和比较这些信息,生成详细的对比报告。
6. 使用`writing`函数将对比报告输出为docx文件。
</task> roleuser 4 content<code>'''
首先,我将使用`search`函数查询华泰期货的服务、费用、用户评价、投资产品、客户支持、技术平台和市场声誉的相关信息。
'''
# 查询华泰期货的相关信息
hua_tai_query = "华泰期货 服务 费用 用户评价 投资产品 客户支持 技术平台 市场声誉"
hua_tai_results = search(query=hua_tai_query)
# 打印搜索结果
print(hua_tai_results)
</code> roleassistant 5 content<observation>
[{'title': '如何了解释华泰期货的收费标准及服务质量?这种收费标准 ...', 'url': 'https://futures.hexun.com/2024-10-26/215150835.html', 'summary': '华泰期货通过合理的收费结构和优质的服务,不仅满足了投资者的需求,也提升了自身的竞争力,成为期货市场中的佼佼者。 (责任编辑:张晓波 )'}, {'title': '如何评估华泰期货的口碑?这种评估方法有哪些参考价值?', 'url': 'https://futures.hexun.com/2024-11-01/215276622.html', 'summary': '通过收集和分析客户的反馈,可以了解华泰期货在服务质量、交易执行、客户支持等方面的表现。 以下是一些常见的客户评价渠道: 评价渠道'}, {'title': '如何评估华泰期货的服务质量?这些服务有哪些具体优势?', 'url': 'https://futures.hexun.com/2024-09-27/214751476.html', 'summary': '首先,评估华泰期货的服务质量可以从客户服务的响应速度和专业性入手。华泰期货提供24小时客户服务热线,确保投资者在任何时间都能得到及时的帮助。此外,华泰期货的客服团队经过专业培训,能够提供准确的市场分析和投资建议,帮助客户做出 ...'}, {'title': '华泰期货手续费一览表(2024年11月更新)_叩富网', 'url': 'https://licai.cofool.com/user/guide_view_2902404.html', 'summary': '期货品种手续费是交易所收取的,实际上华泰期货还会有一部分佣金收取,这部分佣金主要就是期货公司的利润来源,但是不要担心,可以申请优惠降低的提前联系华泰期货客户经理就可以了。'}, {'title': '手续费最便宜的十大期货公司平台(附2024最新手续一览表)', 'url': 'https://licai.cofool.com/user/guide_view_2872421.html', 'summary': '根据中国期货服务报以及期货公司评级介绍,公布的2024年最新手续费最便宜的十大期货公司平台有: 国泰君安、银河期货、中信期货、永安期货、华泰期货、广发期货、中信建投、宏源期货、浙商期货、新湖期货等等。'}, {'title': '华泰期货开户手续费有优惠吗,手续费多少?(含一览表)', 'url': 'https://licai.cofool.com/user/guide_view_2889350.html', 'summary': '华泰期货开户手续费是有优惠的,要提前找华泰期货客户经理进行协商。华泰 期货手续费是由交易所标准和期货公司佣金两部分组成,交易所收取的部分是固定的,期货公司佣金每家都是可以灵活调整的,各营业部可调整的范围权力都是一样的,因此没 ...'}, {'title': '华泰期货手续费揭秘,期货交易的成本究竟有多少?-富维财经网', 'url': 'https://www.fuvi.cn/article/9094052.html', 'summary': '华泰期货手续费是指在期货交易过程中,投资者需要支付的各种费用,这些费用包括交易佣金、印花税、交易所费用等,不同的期货品种、交易量和交易方式,其手续费也会有所差异,对于华泰期货的投资者来说,了解手续费的构成和计算方式至关重要 ...'}, {'title': '华泰期货手续费是多少揭秘 期货投资者不可不知的费用标准与 ...', 'url': 'https://www.zhaocaifu.cn/article/5753883.html', 'summary': '华泰期货手续费具体数额及市场比较 华泰期货的手续费结构相对透明,通常分为交易费用和佣金两部分。以大宗商品期货为例,华泰期货的标准交易费用约为每手1.5元,而在某些特定的合约上,可能会有优惠政策,降低到每手1元。'}, {'title': '华泰期货手续费?_期货问答-希财网问答', 'url': 'https://www.csai.cn/wenda/1040499.html', 'summary': '根据最新信息,华泰期货的手续费标准并不是固定的,它会根据不同品种、交易量以及客户类型有所变化。 一般来说,手续费包括交易所收取的部分和期货公司加收的部分。'}, {'title': '粤向新金融 | 华泰期货:推进数字化转型,助力产业企业套保 ...', 'url': 'https://www.21jingji.com/article/20240903/herald/f619d578c91f73e1439a42b96680cb72.html', 'summary': '华泰期货构建了企业套期保值风险管理系统(简称HRMS系统),凸显期货市场价格发现的优势,满足实体企业参与衍生品市场的即时需求。 华泰期货大数据中心总经理许鉴介绍HRMS系统在套期保值决策智能化方面的创新应用。 股市广播: HRMS系统实现了哪些技术创新? 许鉴: HRMS在技术方面,有如下三点创新: (1)智能化套保决策辅助. HRMS通过数据、资讯的整合,自研套保方案算法及回测引擎,为产业机构提供了高效的事前价格风险发现、事中套保方案测算、事后回测评估的一站式、智能化的套保决策辅助服务。 (2)风险管理计量VaR引擎. VaR引擎是我们自主研发的风险量化引擎,能够使用实时市场数据及各类风险参数来计算市场波动,涵盖了所有交易品种的数据,具备优秀的计算效能。 (3)跟上了国产化的步伐.'}]
</observation><task>从搜索结果中,我获取了关于华泰期货的服务、费用、用户评价、技术平台等方面的信息。接下来,我将使用`open_urls`函数从这些搜索结果中提取更详细的信息,以便进行后续的比较分析。
未来我要做的事:
1. 使用`open_urls`函数从搜索结果中提取华泰期货的服务、费用、用户评价、技术平台等详细信息。
2. 使用`search`函数查询平安期货的服务、费用、用户评价、投资产品、客户支持、技术平台和市场声誉。
3. 使用`open_urls`函数从搜索结果中提取平安期货的详细信息。
4. 整理和比较华泰期货和平安期货的信息,生成详细的对比报告。
5. 使用`writing`函数将对比报告输出为docx文件。</task> roleuser | 0 desc开始 <code> dimensionexample_driven evaluation 0 exec
def check_following(response):
return response.lstrip().startswith("<code>")
required_keys[ ] typecode id0 is_metafalse other_info{"from": "system_para_8"} typeformatting 1 desc</code> 结束 dimensionexample_driven evaluation 0 exec
def check_following(response):
return response.rstrip().endswith("</code>")
required_keys[ ] typecode id1 is_metafalse other_info{"from": "system_para_8"} typeformatting 2 desc如果**<task>**有多个目标,您只需要完成第一个目标 dimensionconditional evaluation 0 exec检查 response 是否严格按 task 要求执行核心操作(允许 print 辅助输出),没有执行任何其他额外操作,直接回答 YES 或 NO。
task: 使用`open_urls`函数从搜索结果中提取华泰期货的服务、费用、用户评价、技术平台等详细信息。
response:{response} required_keys[ ] typellm id2 is_metatrue other_info{"from": "system_para_2", "first_task": "使用`open_urls`函数从搜索结果中提取华泰期货的服务、费用、用户评价、技术平台等详细信息。"} typesemantic 3 desc调用**预设给定的函数** dimensionunconditional evaluation 0 exec提取下面文本中的所有函数调用表达式,直接返回list格式结果,不要任何解释、前缀或赋值语句。
输出格式要求:["function(arg1=value1)", "function2(arg2=value2)", ...]
示例1:
输入文本:
result = search(query="test")
输出函数列表:
["search(query="test")"]
示例2:
输入文本:
data = get_data(id=123)
processed = process(data)
print("done")
输出函数列表:
["get_data(id=123)", "process(data)"]
现在请处理以下文本:
{response}
输出函数列表: required_keys[ ] typellm 1 exec
import re
import ast
def check_following(response):
def extract_tool(text):
pattern = r'\[([^]]+)\]'
matches = re.findall(pattern, text)
if not matches:
return []
try:
tools = eval(f"[{matches[-1]}]")
return tools
except Exception as e:
# print("error in extract_tool:", e)
return []
def validate_function_call(call_str, func_def):
"""
验证函数调用是否符合函数定义
:param call_str: 函数调用字符串,如 'search(query="test", recency_days=7)'
:param func_def: 函数定义字典,包含name和takes_inputs
:return: (is_valid, error_msg) 元组,表示是否有效及错误信息
"""
# 1. 提取函数名和参数部分
try:
# 使用AST安全解析函数调用
module = ast.parse(call_str.strip())
if not isinstance(module, ast.Module) or len(module.body) != 1:
return False, "Invalid statement"
call_node = module.body[0]
if not isinstance(call_node, ast.Expr) or not isinstance(call_node.value, ast.Call):
return False, "Not a function call"
# 获取函数名
called_func_name = None
if isinstance(call_node.value.func, ast.Name):
called_func_name = call_node.value.func.id
# 2. 检查函数名是否匹配
if called_func_name != func_def['name']:
return False, f"Function name mismatch. Expected '{func_def['name']}', got '{called_func_name}'"
# 3. 提取调用参数
provided_args = {}
for kw in call_node.value.keywords:
arg_name = kw.arg
# 获取参数值(处理基本类型的字面量)
if isinstance(kw.value, ast.Constant):
arg_value = kw.value.value
elif isinstance(kw.value, ast.Num): # Python < 3.8
arg_value = kw.value.n
elif isinstance(kw.value, ast.Str): # Python < 3.8
arg_value = kw.value.s
elif isinstance(kw.value, ast.NameConstant): # Python < 3.8
arg_value = kw.value.value
else:
arg_value = None # 复杂表达式暂不处理
provided_args[arg_name] = arg_value
# 4. 验证参数
expected_args = func_def['takes_inputs']
errors = []
# 检查必填参数
for arg_name, arg_def in expected_args.items():
if arg_def['required'] == 'True' and arg_name not in provided_args:
errors.append(f"Missing required argument: '{arg_name}'")
# 检查未知参数
for arg_name in provided_args:
if arg_name not in expected_args:
errors.append(f"Unexpected argument: '{arg_name}'")
# 检查参数类型
for arg_name, arg_value in provided_args.items():
if arg_name in expected_args:
expected_type = expected_args[arg_name]['type']
actual_type = type(arg_value).__name__
# 简单类型检查
if expected_type == 'int' and not isinstance(arg_value, int):
errors.append(f"Argument '{arg_name}' should be {expected_type}, got {actual_type}")
elif expected_type == 'str' and not isinstance(arg_value, str):
errors.append(f"Argument '{arg_name}' should be {expected_type}, got {actual_type}")
if errors:
return False, " | ".join(errors)
return True, "Valid call"
except Exception as e:
return False, f"Parsing error: {str(e)}"
available_functions = [{'name': 'writing', 'takes_inputs': {'query': {'description': '写作内容', 'type': 'str', 'required': 'True'}, 'reference': {'description': '参考资料', 'type': 'str', 'required': 'False'}, 'word_number': {'description': '字数要求', 'type': 'int', 'required': 'False'}, 'output_format': {'description': '输出格式', 'type': 'str', 'required': 'True'}}}, {'name': 'search', 'takes_inputs': {'query': {'description': '搜索关键词', 'type': 'str', 'required': 'True'}, 'recency_days': {'description': '搜索结果的时间范围', 'type': 'int', 'required': 'False'}}}, {'name': 'open_docs', 'takes_inputs': {'query': {'description': '需要从文档中提取的文本内容', 'type': 'str', 'required': 'True'}, 'file_paths': {'description': '需要打开的文件路径列表', 'type': 'list', 'required': 'True'}}}, {'name': 'translate', 'takes_inputs': {'query': {'description': '翻译内容', 'type': 'str', 'required': 'True'}, 'to_lang': {'description': '翻译目标语言', 'type': 'str', 'required': 'False'}}}, {'name': 'open_urls', 'takes_inputs': {'query': {'description': '需要从网页中提取的文本内容', 'type': 'str', 'required': 'True'}, 'urls': {'description': '需要打开的url列表', 'type': 'list', 'required': 'True'}}}]
test_calls = extract_tool(response)
# print(test_calls)
for call in test_calls:
Flag = False
for function_def in available_functions:
is_valid, msg = validate_function_call(call, function_def)
# print(function_def["name"], msg)
if is_valid == True:
Flag = True
break
if Flag == False:
return False
return True
required_keys[ ] typecode id3 is_metafalse other_info{"available_functions": [{"name": "writing", "takes_inputs": {"query": {"description": "写作内容", "type": "str", "required": "True"}, "reference": {"description": "参考资料", "type": "str", "required": "False"}, "word_number": {"description": "字数要求", "type": "int", "required": "False"}, "output_format": {"description": "输出格式", "type": "str", "required": "True"}}}, {"name": "search", "takes_inputs": {"query": {"description": "搜索关键词", "type": "str", "required": "True"}, "recency_days": {"description": "搜索结果的时间范围", "type": "int", "required": "False"}}}, {"name": "open_docs", "takes_inputs": {"query": {"description": "需要从文档中提取的文本内容", "type": "str", "required": "True"}, "file_paths": {"description": "需要打开的文件路径列表", "type": "list", "required": "True"}}}, {"name": "translate", "takes_inputs": {"query": {"description": "翻译内容", "type": "str", "required": "True"}, "to_lang": {"description": "翻译目标语言", "type": "str", "required": "False"}}}, {"name": "open_urls", "takes_inputs": {"query": {"description": "需要从网页中提取的文本内容", "type": "str", "required": "True"}, "urls": {"description": "需要打开的url列表", "type": "list", "required": "True"}}}], "from": "system_para_2"} typeresource 4 desc最多调用一次**预设给定的函数** dimensionunconditional evaluation 0 exec提取下面文本中的所有函数调用表达式,直接返回list格式结果,不要任何解释、前缀或赋值语句。
输出格式要求:["function(arg1=value1)", "function2(arg2=value2)", ...]
示例1:
输入文本:
result = search(query="test")
输出函数列表:
["search(query="test")"]
示例2:
输入文本:
data = get_data(id=123)
processed = process(data)
print("done")
输出函数列表:
["get_data(id=123)", "process(data)"]
现在请处理以下文本:
{response}
输出函数列表: required_keys[ response ] typellm 1 exec
import re
import ast
def check_following(response):
def validate_function_call(call_str, func_def):
"""
验证函数调用是否符合函数定义
:param call_str: 函数调用字符串,如 'search(query="test", recency_days=7)'
:param func_def: 函数定义字典,包含name和takes_inputs
:return: (is_valid, error_msg) 元组,表示是否有效及错误信息
"""
# 1. 提取函数名和参数部分
try:
# 使用AST安全解析函数调用
module = ast.parse(call_str.strip())
if not isinstance(module, ast.Module) or len(module.body) != 1:
return False, "Invalid statement"
call_node = module.body[0]
if not isinstance(call_node, ast.Expr) or not isinstance(call_node.value, ast.Call):
return False, "Not a function call"
# 获取函数名
called_func_name = None
if isinstance(call_node.value.func, ast.Name):
called_func_name = call_node.value.func.id
# 2. 检查函数名是否匹配
if called_func_name != func_def['name']:
return False, f"Function name mismatch. Expected '{func_def['name']}', got '{called_func_name}'"
# 3. 提取调用参数
provided_args = {}
for kw in call_node.value.keywords:
arg_name = kw.arg
# 获取参数值(处理基本类型的字面量)
if isinstance(kw.value, ast.Constant):
arg_value = kw.value.value
elif isinstance(kw.value, ast.Num): # Python < 3.8
arg_value = kw.value.n
elif isinstance(kw.value, ast.Str): # Python < 3.8
arg_value = kw.value.s
elif isinstance(kw.value, ast.NameConstant): # Python < 3.8
arg_value = kw.value.value
else:
arg_value = None # 复杂表达式暂不处理
provided_args[arg_name] = arg_value
# 4. 验证参数
expected_args = func_def['takes_inputs']
errors = []
# 检查必填参数
for arg_name, arg_def in expected_args.items():
if arg_def['required'] == 'True' and arg_name not in provided_args:
errors.append(f"Missing required argument: '{arg_name}'")
# 检查未知参数
for arg_name in provided_args:
if arg_name not in expected_args:
errors.append(f"Unexpected argument: '{arg_name}'")
# 检查参数类型
for arg_name, arg_value in provided_args.items():
if arg_name in expected_args:
expected_type = expected_args[arg_name]['type']
actual_type = type(arg_value).__name__
if expected_type == 'int' and not isinstance(arg_value, int):
errors.append(f"Argument '{arg_name}' should be {expected_type}, got {actual_type}")
elif expected_type == 'str' and not isinstance(arg_value, str):
errors.append(f"Argument '{arg_name}' should be {expected_type}, got {actual_type}")
if errors:
return False, " | ".join(errors)
return True, "Valid call"
except Exception as e:
return False, f"Parsing error: {str(e)}"
def extract_tool(text):
pattern = r'\[([^]]+)\]'
matches = re.findall(pattern, text)
if not matches:
return []
try:
tools = eval(f"[{matches[-1]}]")
return tools
except Exception as e:
# print("error in extract_tool:", e)
return []
available_functions = [{'name': 'writing', 'takes_inputs': {'query': {'description': '写作内容', 'type': 'str', 'required': 'True'}, 'reference': {'description': '参考资料', 'type': 'str', 'required': 'False'}, 'word_number': {'description': '字数要求', 'type': 'int', 'required': 'False'}, 'output_format': {'description': '输出格式', 'type': 'str', 'required': 'True'}}}, {'name': 'search', 'takes_inputs': {'query': {'description': '搜索关键词', 'type': 'str', 'required': 'True'}, 'recency_days': {'description': '搜索结果的时间范围', 'type': 'int', 'required': 'False'}}}, {'name': 'open_docs', 'takes_inputs': {'query': {'description': '需要从文档中提取的文本内容', 'type': 'str', 'required': 'True'}, 'file_paths': {'description': '需要打开的文件路径列表', 'type': 'list', 'required': 'True'}}}, {'name': 'translate', 'takes_inputs': {'query': {'description': '翻译内容', 'type': 'str', 'required': 'True'}, 'to_lang': {'description': '翻译目标语言', 'type': 'str', 'required': 'False'}}}, {'name': 'open_urls', 'takes_inputs': {'query': {'description': '需要从网页中提取的文本内容', 'type': 'str', 'required': 'True'}, 'urls': {'description': '需要打开的url列表', 'type': 'list', 'required': 'True'}}}]
test_calls = extract_tool(response)
count = 0
for call in test_calls:
Flag = False
for function_def in available_functions:
is_valid, msg = validate_function_call(call, function_def)
if is_valid == True:
Flag = True
break
if Flag == True:
count += 1
if count <= 1:
return True
else:
return False
required_keys[ ] typecode id4 is_metafalse other_info{"available_functions": [{"name": "writing", "takes_inputs": {"query": {"description": "写作内容", "type": "str", "required": "True"}, "reference": {"description": "参考资料", "type": "str", "required": "False"}, "word_number": {"description": "字数要求", "type": "int", "required": "False"}, "output_format": {"description": "输出格式", "type": "str", "required": "True"}}}, {"name": "search", "takes_inputs": {"query": {"description": "搜索关键词", "type": "str", "required": "True"}, "recency_days": {"description": "搜索结果的时间范围", "type": "int", "required": "False"}}}, {"name": "open_docs", "takes_inputs": {"query": {"description": "需要从文档中提取的文本内容", "type": "str", "required": "True"}, "file_paths": {"description": "需要打开的文件路径列表", "type": "list", "required": "True"}}}, {"name": "translate", "takes_inputs": {"query": {"description": "翻译内容", "type": "str", "required": "True"}, "to_lang": {"description": "翻译目标语言", "type": "str", "required": "False"}}}, {"name": "open_urls", "takes_inputs": {"query": {"description": "需要从网页中提取的文本内容", "type": "str", "required": "True"}, "urls": {"description": "需要打开的url列表", "type": "list", "required": "True"}}}], "from": "system_para_2"} typeresource 5 desc每次使用预设函数后都应该将使用print将函数结果打印出来 dimensionunconditional evaluation 0 exec检查下面文本中的每个函数调用表达式,判断其下一个执行语句是否为print。直接回答YES或NO,不要任何其他内容。
文本:{response}
回答: required_keys[ response ] typellm id5 is_metafalse other_info{"available_functions": [{"name": "writing", "takes_inputs": {"query": {"description": "写作内容", "type": "str", "required": "True"}, "reference": {"description": "参考资料", "type": "str", "required": "False"}, "word_number": {"description": "字数要求", "type": "int", "required": "False"}, "output_format": {"description": "输出格式", "type": "str", "required": "True"}}}, {"name": "search", "takes_inputs": {"query": {"description": "搜索关键词", "type": "str", "required": "True"}, "recency_days": {"description": "搜索结果的时间范围", "type": "int", "required": "False"}}}, {"name": "open_docs", "takes_inputs": {"query": {"description": "需要从文档中提取的文本内容", "type": "str", "required": "True"}, "file_paths": {"description": "需要打开的文件路径列表", "type": "list", "required": "True"}}}, {"name": "translate", "takes_inputs": {"query": {"description": "翻译内容", "type": "str", "required": "True"}, "to_lang": {"description": "翻译目标语言", "type": "str", "required": "False"}}}, {"name": "open_urls", "takes_inputs": {"query": {"description": "需要从网页中提取的文本内容", "type": "str", "required": "True"}, "urls": {"description": "需要打开的url列表", "type": "list", "required": "True"}}}], "from": "system_para_2"} typeformatting 6 descprint后立即使用</code>结束本次编码 dimensionunconditional evaluation 0 exec检查文本中每个 print 语句后是否存在 </code>(允许中间有换行或空格)。直接回答 YES 或 NO。
文本:{response}
回答: required_keys[ response ] typellm id6 is_metafalse other_info{"from": "system_para_2"} typeformatting 7 desc<code>- # 写注释 dimensionexample_driven evaluation 0 exec检查文本是否以<code>后接注释(包括单行或多行注释)开头,或无<code>时直接以注释开头。注释可以是#、'''或"""形式。直接回答YES或NO。
文本:{response}
回答: required_keys[ ] typellm id7 is_metafalse other_info{"from": "system_para_8"} typeformatting 8 desc你可以使用以下packages dimensionunconditional evaluation 0 exec提取下面文本中的所有的调用package,直接返回list格式结果,不要任何解释、前缀或赋值语句。
输出格式要求:["package1", "package2", ...]:
Here is model response: {response} required_keys[ ] typellm 1 exec
import ast
def check_following(response: str) -> bool:
available_packages = ['statistics', 'sqlite3', 'queue', 'time', 'stat', 'matplotlib', 'itertools', 'math', 'datetime', 'pandas', 'unicodedata', 'collections', 'PyPDF2', 'random']
try:
items = ast.literal_eval(response)
return isinstance(items, list) and all(item in available_packages for item in items)
except:
return False
required_keys[ ] typecode id8 is_metafalse other_info{"available_packages": ["statistics", "sqlite3", "queue", "time", "stat", "matplotlib", "itertools", "math", "datetime", "pandas", "unicodedata", "collections", "PyPDF2", "random"], "from": "system_para_3"} typeresource | {"role": "assistant", "content": "'''\n接下来,我将使用`open_urls`函数从搜索结果中提取华泰期货的服务、费用、用户评价、技术平台等详细信息。\n'''\n# 提取华泰期货的详细信息\nhua_tai_urls = [result['url'] for result in hua_tai_results]\nhua_tai_details = open_urls(query=\"华泰期货 服务 费用 用户评价 技术平台\", urls=hua_tai_urls)\n\n# 打印提取的详细信息\nprint(hua_tai_details)\n"} |
| 1 | agentif:fe0f561ccc97b48b319069d14d316d207b66fc13 | 0 contentYou are a research assistant. Based on your knowledge, provide a concise summary of information related to the given term. The summary must 2-3 paragraphs and less than 300 words. Capture the main points. Write succintly, no need to have complete sentences or good grammar. This will be consumed by someone synthesizing a report, so its vital you capture the essence and ignore any fluff. Do not include any additional commentary other than the summary itself. rolesystem 1 contentThe search term is Mythology and Symbolism in Ancient Mesopotamian Art.
Provide a concise summary of information related to the given search term. roleuser | 0 descThe summary must be 2-3 paragraphs and less than 300 words. dimensionunconditional evaluation 0 execimport re
def check_following(response: str) -> bool:
word_count = len(response.split())
paragraphs = response.split('\n')
paragraphs = [p.strip() for p in paragraphs if p.strip()]
return 2 <= len(paragraphs) <= 3 and word_count < 300
required_keys[ ] typecode id0 is_metafalse other_info{"from": "system_para_0", "type_explanation": "The constraint specifies the structure and presentation of the output by limiting it to 2-3 paragraphs and less than 300 words. This relates to how the content is formatted, rather than its meaning (semantic) or resource usage.", "meta_expalnation": "The given constraint directly restricts the format and content of the model's output, specifying the number of paragraphs and maximum word count for the summary. It does not involve high-level rules that manage or interact with other constraints. Therefore, it is not a Meta Constraint.", "evaluation_type": {"constraint_type": "code", "explanation": "The constraint involves checking the length (2-3 paragraphs) and word count (less than 300 words), both of which can be directly validated programmatically without requiring semantic understanding or content extraction."}, "evaluation_generation_success": true} type["formatting"] 1 descDo not include any additional commentary other than the summary itself. dimensionunconditional evaluation 0 execDoes the response exclude any additional commentary outside of the concise summary itself (for example, 'Here is the generated summary: ...')? Please answer YES/NO directly and do not enter anything else.
Here is the model response: {response} required_keys[ ] typellm id3 is_metafalse other_info{"from": "system_para_0", "type_explanation": "The constraint focuses on ensuring that the output content adheres to a specific guideline: delivering only the summary and avoiding any commentary. This requirement aligns with the semantic category as it dictates the meaningfulness, completeness, and style of the content produced.", "meta_expalnation": "The constraint directly governs the model's output by specifying that no additional commentary should be included, focusing on the content of the output itself. It does not manage or define how constraints are selected, prioritized, ignored, deduplicated, or combined, so it does not qualify as a Meta Constraint.", "evaluation_type": {"constraint_type": "llm", "explanation": "This constraint requires a semantic understanding of the response to assess whether any commentary beyond the summary itself is included. It involves open-ended judgment about the nature of the content, which cannot be validated via straightforward logic or extraction."}, "evaluation_generation_success": true} type["semantic"] 2 descProvide a concise summary of information related to the given search term. dimensionunconditional evaluation 0 execDoes the model response provide a concise summary of information specifically related to the given search term? Please answer YES/NO directly and do not enter anything else.
Here is the model response: {response} required_keys[ ] typellm id5 is_metafalse other_info{"from": "query", "type_explanation": "The constraint focuses on delivering meaningful and purposeful content—a concise summary that is relevant to the given search term. It inherently demands accuracy, completeness, and alignment to the subject matter. Since it specifies the nature of the output as 'concise' and directly relates to information connected to a search term, it is classified under semantic constraints.", "meta_expalnation": "The given constraint directly controls the output by specifying that the model should provide a concise summary based on the search term. It does not govern the selection, prioritization, deduplication, or composition of multiple constraints, which are characteristics of Meta Constraints.", "evaluation_type": {"constraint_type": "llm", "explanation": "The constraint requires a semantic understanding of what constitutes a 'concise summary' in relation to the search term, which involves subjective assessment and interpretation beyond straightforward code-based validation."}, "evaluation_generation_success": true} type["semantic"] | null |
| 2 | agentif:1b5c6471196b80a04abd361a262964a53e27f97d | 0 contentYou are given 3 atomic functions to help you retrieve and operate knowledge from Wikipedia:
1. Search(). Input: (name, [optional] descriptor). Output: list[entities]. This function helps you find and disambiguate an entity given its name and optional descriptor. If no descriptor is provided, the most popular entity will be returned. For example, Search("Michael Jordan") returns the famous basketball player ["Michael Jordan"], while Search("Michael Jordan", "football goalkeeper") returns the English retired football goalkeeper ["Michael Jordan (footballer)"]. When the question provides explicit entity knowledge, always write a descriptor for the Search() function based on the question's information.
2. Relate(). Input: there are 2 input possibilities, (head_entity, relation), or (head_entity, tail_entity). Output: list[tail_entities], or list[relations]. This function helps you find the tail_entities given a head_entity and relation, or relations given a head_entity and tail_entity. For example, Relate("Barack Obama", "child") returns ["Malia Obama", "Sasha Obama"], and Relate("Barack Obama", "Michelle Obama") returns ["spouse"]. You may also search attribute relations using Relate() by treating attributes as tail entities. For example, Relate("Barack Obama", "time served as US president") returns ["1997 to 2004"], and Relate("Barack Obama", "1961")" returns ["year of birth"].
3. Filter(). Input: (list[entities], condition). Output: list[entities]. This function helps you filter out entities that satisfy a factual attribute condition. For example, Filter(["Lionel Messi", "Steven Jobs", "Bill Gates"], "born in 1955"), returns ["Bill Gates", "Steve Jobs"], and Filter(["Lionel Messi", "Cristiano Ronaldo"], "is Portuguese") returns ["Cristiano Ronaldo"].
Examples:
Question: What was the largest passenger capacity of the plane type used for BOAC Flight 911?
Decomposition Tree: {"What was the largest passenger capacity of the plane type used for BOAC Flight 911?": ["1. What was the plane type used for BOAC Flight 911?", "2. What was the largest passenger capacity of [1]?"], "1. What was the plane type used for BOAC Flight 911?": ["3. What is BOAC Flight 911?", "4. What is the plane type used for [3]?"], "3. What is BOAC Flight 911?": "Search("BOAC Flight 911")", "4. What is the plane type used for [3]?": "Relate([3], "plane type")", "2. What was the largest passenger capacity of [1]?": "Relate([1], "largest passenger capacity")"}
Question: Who was the film which was Kim Dae-woo's directing debut about?
Decomposition Tree: {"Who was the film which was Kim Dae-woo's directing debut about?": ["1. What is Kim Dae-woo's directing debut film?", "2. Who was [1] about?"], "1. What is Kim Dae-woo's directing debut film?": ["3. Who is Kim Dae-woo?", "4. What is [3]'s directing debut film?"], "3. Who is Kim Dae-woo?": "Search("Kim Dae-woo", "film director")", "4. What is [3]'s directing debut film?": "Relate("Kim Dae-woo", "directing debut film")", "2. Who was [1] about?": "Relate([1], "about person")"}
Question: Which city was the man who is known for a science humor story based on the tongue-in-cheek combination of two adages born in?
Decomposition Tree: {"Which city was the man who is known for a science humor story based on the tongue-in-cheek combination of two adages born in?": ["1. Who is the man known for a science humor story based on the tongue-in-cheek combination of two adages?", "2. In which city was [1] born?"], "1. Who is the man known for a science humor story based on the tongue-in-cheek combination of two adages?": ["3. What is a science humor story based on the tongue-in-cheek combination of two adages?", "4. Who is the man known for [3]?"], "3. What is a science humor story based on the tongue-in-cheek combination of two adages?": "Search("science humor story based on the tongue-in-cheek combination of two adages")", "4. Who is the man known for [3]?": "Relate([3], "man known for")", "2. In which city was [1] born?": "Relate([1], "born in city")"}
Question: What is the birthday of this Anglo-Irish actress, courtean, and mistress, who was the mother to the illegitimate daughter of King William IV?
Decomposition Tree: {"What is the birthday of this Anglo-Irish actress, courtesan, and mistress, who was the mother to the illegitimate daughter of King William IV?": ["1. Who is the Anglo-Irish actress, courtesan, and mistress who was the mother to the illegitimate daughter of King William IV?", "2. What is the birthday of [1]?"], "1. Who is the Anglo-Irish actress, courtesan, and mistress who was the mother to the illegitimate daughter of King William IV?": ["3. Who was the mother to the illegitimate daughter of King William IV?", "4. Among [3], who is an Anglo-Irish actress, courtesan, and mistress?"], "3. Who was the mother to the illegitimate daughter of King William IV?": ["5. Who was King William IV?", "6. Who was the illegitimate daughter of [5]?", "7. Who was the mother to [6]?"], "5. Who was King William IV?": "Search("King William IV")", "6. Who was the illegitimate daughter of [5]?": "Relate([5], "illegitimate daughter")", "7. Who was the mother to [6]?": "Relate([6], "mother")", "4. Among [3], who is an Anglo-Irish actress, courtesan, and mistress?": "Filter([3], "Anglo-Irish actress, courtesan, and mistress")", "2. What is the birthday of [1]?": "Relate([1], "birthday")"}
Question: Are Billy and Barak both breeds of scenthound? (Barak is also known as a Bosnian Coarse-haired Hound)?
Decomposition Tree: {"Are Billy and Barak both breeds of scenthound? (Barak is also known as a Bosnian Coarse-haired Hound)": ["1. Is Billy a breed of scenthound?", "2. Is Barack (also known as a Bosnian Coarse-haired Hound) a breed of scenthound?"], "1. Is Billy a breed of scenthound?": ["3. What is Billy?", "4. What breed is [3]?"], "3. What is Billy?": "Search("Billy", "dog")", "4. What breed is [3]?": "Relate([3], "is breed")", "2. Is Barack (also known as a Bosnian Coarse-haired Hound) a breed of scenthound?": ["5. What is Barack (also known as a Bosnian Coarse-haired Hound)?", "6. What breed is [5]?"], "5. What is Barack (also known as a Bosnian Coarse-haired Hound)?": "Search("Barack (also known as a Bosnian Coarse-haired Hound)", "dog")", "6. What breed is [5]?": "Relate([5], "is breed")"}
Question: What Pakistani actor and writer from Islamabad helped write for the 2012 Pakistani comedy drama sitcom, "Coke Kahani"?
Decomposition Tree: {"What Pakistani actor and writer from Islamabad helped write for the 2012 Pakistani comedy drama sitcom, "Coke Kahani"?": ["1. What is the 2012 Pakistani comedy drama sitcom, "Coke Kahani"?", "2. Who helped write for [1]?", "3. Who is the Pakistani actor and writer from Islamabad among [2]?"], "1. What is the 2012 Pakistani comedy drama sitcom, "Coke Kahani"?": "Search("Coke Kahani", "2012 Pakistani comedy drama sitcom")", "2. Who helped write for [1]?": "Relate([1], "writers")", "3. Who is the Pakistani actor and writer from Islamabad among [2]?": "Filter([2], "Pakistani actor and writer from Islamabad")"}
Question: In which city have Gary Ayres and Neil Craig both been head coach of the Crows?
Decomposition Tree: {"In which city have Gary Ayres and Neil Craig both been head coach of the Crows?": ["1. Who is Gary Ayres?", "2. Who is Neil Craig?", "3. What is the Crows?", "4. In which city has [1] been head coach of [3]?", "5. In which city has [2] been head coach of [3]?", "6. Given answers of [4] and [5], in which city have Gary Ayres and Neil Craig both been head coach of the Crows?"], "1. Who is Gary Ayres?": "Search("Gary Ayres", "Australian rules football coach")", "2. Who is Neil Craig?": "Search("Neil Craig", "Australian rules football coach")", "3. What is the Crows?": "Search("the Crows", "Australian rules football club")", "4. In which city has [1] been head coach of [3]?": "Relate([1], "was head coach of [3] in city")", "5. In which city has [2] been head coach of [3]?": "Relate([2], "was head coach of [3] in city")", "6. Given answers of [4] and [5], in which city have Gary Ayres and Neil Craig both been head coach of the Crows?": "[END]"}
Question: Have Marc Rosset and Max Mirnyi both been professional tennis players?
Decomposition Tree: {"Have Marc Rosset and Max Mirnyi both been professional tennis players?": ["1. Has Marc Rosset been a professional tennis player?", "2. Has Max Mirnyi been a professional tennis player?"], "1. Has Marc Rosset been a professional tennis player?": ["3. Who is Marc Rosset?", "4. Have [3] been a professional tennis player?"], "3. Who is Marc Rosset?": "Search("Marc Rosset")", "4. Have [3] been a professional tennis player?": "Relate([3], "is professional tennis player")", "2. Has Max Mirnyi been a professional tennis player?": ["5. Who is Max Mirnyi?", "6. Has [5] been a professional tennis player?"], "5. Who is Max Mirnyi?": "Search("Max Mirnyi")", "6. Has [5] been a professional tennis player?": "Relate([5], "is professional tennis player")"}
Question: What baseball team, part of the ten-school collegiate athletic conference headquartered in Irving, Texas, was coached by Randy Mazey in 2016?
Decomposition Tree: {"What baseball team, part of the ten-school collegiate athletic conference headquartered in Irving, Texas, was coached by Randy Mazey in 2016?": ["1. What is the ten-school collegiate athletic conference headquartered in Irving, Texas?", "2. What baseball teams are part of [1]?", "3. What baseball team was coached by Randy Mazey in 2016?", "4. Given Answers of [2] and [3], what baseball team belongs to both?"], "1. What is the ten-school collegiate athletic conference headquartered in Irving, Texas?": "Search("ten-school collegiate athletic conference headquartered in Irving, Texas")", "2. What baseball teams are part of [1]?": "Relate([1], "baseball team")", "3. What baseball team was coached by Randy Mazey in 2016?": "Relate("baseball team", "coached by Randy Mazey in 2016")", "4. Given Answers of [2] and [3], what baseball team belongs to both?": "[END]"}
Question: George Gershwin is an American Composer and Judith Weir is a composer from which country?
Decomposition Tree: {"George Gershwin is an American Composer and Judith Weir is a composer from which country?": ["1. Who is George Gershwin?", "2. Who is Judith Weir?", "3. What country is [2] from?"], "1. Who is George Gershwin?": "Search("George Gershwin", "American composer")", "2. Who is Judith Weir?": "Search("Judith Weir", "composer")", "3. What country is [2] from?": "Relate([2], "is from country")"}
Question: Which goalkeeper was nicknamed the "Black Spider", Turgay Şeren or Lev Yashin?
Decomposition Tree: {"Which goalkeeper was nicknamed the "Black Spider", Turgay Şeren or Lev Yashin?": ["1. Is goalkeeper Turgay Şeren nicknamed the "Black Spider"?", "2. Is goalkeeper Lev Yashin nicknamed the "Black Spider"?"], "1. Is goalkeeper Turgay Şeren nicknamed the "Black Spider"?": ["3. Who is goalkeeper Turgay Şeren?", "4. Is [3] nicknamed the "Black Spider"?"], "4. Who is goalkeeper Turgay Şeren?": "Search("Turgay Şeren", "goalkeeper")", "5. Is [4] nicknamed the "Black Spider"?": "Relate([4], "has nickname Black Spider")", "2. Is goalkeeper Lev Yashin nicknamed the "Black Spider"?": ["5. Who is goalkeeper Lev Yashin?", "6. Is [5] nicknamed the "Black Spider"?"], "5. Who is goalkeeper Lev Yashin?": "Search("Lev Yashin", "goalkeeper")", "6. Is [5] nicknamed the "Black Spider"?": "Relate([6], "has nickname Black Spider")"}
Question: What company did a man who hired Sioux Falls architect Wallace L. Dow to build a home in Worthing, Minnesota found?
Decomposition Tree: {"What company did a man who hired Sioux Falls architect Wallace L. Dow to build a home in Worthing, Minnesota found?": ["1. Who is the man who hired Sioux Falls architect Wallace L. Dow to build a home in Worthing, Minnesota?", "2. What company did [1] found?"], "1. Who is the man who hired Sioux Falls architect Wallace L. Dow to build a home in Worthing, Minnesota?": ["3. Who is Sioux Falls architect Wallace L. Dow?", "4. Who hired [3] to build a home in Worthing, Minnesota?"], "3. Who is Sioux Falls architect Wallace L. Dow?": "Search("Wallace L. Dow", "Sioux Falls architect")", "4. Who hired [3] to build a home in Worthing, Minnesota?": "Relate([3], "was hired to build home in Worthing, Minnesota by person")", "2. What company did [1] found?": "Relate([1], "founded company")"}
Question: Radio shack made a line of computers in the 1980's which was marketed as the TRS-80 Color Computer or the Interact Home Computer?
Decomposition Tree: {"Radio shack made a line of computers in the 1980's which was marketed as the TRS-80 Color Computer or the Interact Home Computer?": ["1. What line of computers did Radio shack make in the 1980s?", "2. What was the marketing name of [1]?", "3. Given answers of [2] and [3], was the line of computers marketed as the TRS-80 Color Computer or the Interact Home Computer?"], "1. What line of computers did Radio shack make in the 1980s?": ["3. What is Radio shack?", "4. What line of computers did [3] make in the 1980s?"], "3. What is Radio shack?": "Search("Radio shack")", "4. What line of computers did [3] make in the 1980s?": "Relate([3], "made line of computers in 1980s")", "2. What was the marketing name of [1]?": "Relate([1], "marketing name")", "3. Given answers of [1] and [2], was the line of computers marketed as the TRS-80 Color Computer or the Interact Home Computer?": "[END]"}
Question: Baraki Barak District is situated in the western part of a province whose capital is what?
Decomposition Tree: {"Baraki Barak District is situated in the western part of a province whose capital is what?": ["1. What province is Baraki Barak District situated in?", "2. What is the capital of [1]?"], "1. What province is Baraki Barak District situated in?": ["3. What is Baraki Barak District?", "4. What province is [3] situated in?"], "3. What is Baraki Barak District?": "Search("Baraki Barak District")", "4. What province is [3] situated in?": "Relate([3], "is situated in province")", "2. What is the capital of [1]?": "Relate([1], "capital")"}}
Question: What was the former name of the stadium, from 1997-2017, where the Aztecs play?
Decomposition Tree: {"What was the former name of the stadium, from 1997-2017, where the Aztecs play?": ["1. What is the stadium where the Aztecs play?", "2. What was the former name of [1] from 1997-2017?"], "1. What is the stadium where the Aztecs play?": ["3. Who are the Aztecs?", "4. What is the stadium where [3] play?"], "3. Who are the Aztecs?": "Search("Aztecs", "sports team")", "4. What is the stadium where [3] play?": "Relate([3], "plays at stadium")", "2. What was the former name of [1] from 1997-2017?": "Relate([1], "had former name from 1997-2017")"}
Question: Oak Beach, New York and Great South Bay are both situated between what same island?
Decomposition Tree: {"Oak Beach, New York and Great South Bay are both situated between what same island?": ["1. What is Oak Beach, New York situated between?", "2. What is Great South Bay situated between?", "3. Given answers of [1] and [2], what same island are they situated between?"], "1. What is Oak Beach, New York situated between?": ["4. What is Oak Beach, New York?", "5. What is [4] situated between?"], "4. What is Oak Beach, New York?": "Search("Oak Beach, New York")", "5. What is [4] situated between?": "Relate([4], "situated between")", "2. What is Great South Bay situated between?": ["6. What is Great South Bay?", "7. What is [6] situated between?"], "6. What is Great South Bay?": "Search("Great South Bay")", "7. What is [6] situated between?": "Relate([6], "situated between")", "3. Given answers of [1] and [2], what same island are they situated between?": "[END]"}
Question: What was the third studio album released by Richard Melville Hall?
Decomposition Tree: {"What was the third studio album released by Richard Melville Hall?": ["1. Who is Richard Melville Hall?", "2. What are the studio albums released by [1]?", "3. What is the third studio album among [2]?"], "1. Who is Richard Melville Hall?": "Search("Richard Melville Hall")", "2. What are the studio albums released by [1]?": "Relate([1], "studio albums")", "3. What is the third studio album among [2]?": "Filter([2], "third studio album")"}
Question: Who acted in the film and television series, "Harry and the Hendersons," and also worked with Danny Glover?
Decomposition Tree: {"Who acted in the film and television series, "Harry and the Hendersons," and also worked with Danny Glover?": ["1. Who acted in the film and television series, "Harry and the Hendersons"?", "2. Who among [1] also worked with Danny Glover?"], "1. Who acted in the film and television series, "Harry and the Hendersons"?": ["3. What is the film and television series, "Harry and the Hendersons"?", "4. Who acted in [3]?"], "3. What is the film and television series, "Harry and the Hendersons"?": "Search("Harry and the Hendersons", "film and television series")", "4. Who acted in [3]?": "Relate([3], "actors")", "2. Who among [1] also worked with Danny Glover?": "Filter([1], "worked with Danny Glover")"}
Your Question.
Question: The 2005 film Remedy featured Frank Vincent from The Sopranos and several mob movies by which acclaimed director?
Decomposition Tree:
roleuser | 0 descWhen the question provides explicit entity knowledge, always write a descriptor for the Search() function based on the question's information. dimensionconditional evaluation 0 execDoes the question provide explicit entity knowledge? Please answer YES/NO directly and do not enter anything else.
Here is model response: {response} required_keys[ ] typellm_conditional_check 1 execDid the model write a descriptor parameter for the Search() function? Please answer YES/NO directly and do not enter anything else.
Here is model response: {response} required_keys[ ] typellm id-1 is_metafalse other_info{"from": "system_para_0", "condition_desc": "If the question provides explicit entity knowledge, always write a descriptor for the Search() function based on the question's information.", "complete_instruction_para": []} type["semantic"] 1 descConstruct a hierarchical question decomposition tree in json format dimensionexample_driven evaluation 0 execimport json
check_following(response):
try:
json.loads(response)
return True
except json.JSONDecodeError:
return False
required_keys[ ] typecode id-1 is_metafalse other_info{"from": "system_para_0", "condition_desc": "", "complete_instruction_para": []} type["formatting"] 2 descThe tree starts with the original complex question as the root node, and each non-root node is a sub-question of its parent. dimensionexample_driven evaluation 0 execIn the model response, does the JSON question decomposition tree start with the original complex question as the root node, and each non-root node is a sub-question of its parent? Please answer YES/NO directly and do not enter anything else.
Here is model response: {response} required_keys[ ] typellm id-1 is_metafalse other_info{"from": "system_para_0", "condition_desc": "", "complete_instruction_para": []} type["semantic"] 3 descContinue decomposing until a sub-question cannot be further decomposed and could either be: (1) directly answered by calling one of the three atomic functions Search(), Relate(), Filter(), or (2) directly answered by analyzing the answers of at least two previously answered questions, such as comparing, judging, intersecting, counting, etc. dimensionexample_driven evaluation 0 execIn the model response, does the model continue question decomposition until each sub-question cannot be further decomposed and could either be: (1) directly answered by calling one of the three atomic functions Search(), Relate(), Filter(), or (2) directly answered by analyzing the answers of at least two previously answered questions, such as comparing, judging, intersecting, or counting? Please answer YES/NO directly and do not enter anything else.
Here is model response: {response} required_keys[ ] typellm id-1 is_metafalse other_info{"from": "system_para_0", "condition_desc": "", "complete_instruction_para": []} type["semantic"] 4 descIn case (1), write this sub-question with its corresponding function call as a leaf node. dimensionexample_driven evaluation 0 execDoes there exist sub-questions that satisfy case (1), which can be 'directly answered by calling one of the three atomic functions Search(), Relate(), Filter()'? Please answer YES/NO directly and do not enter anything else.
Here is model response: {response} required_keys[ ] typellm_conditional_check 1 execIn the model response, are all sub-questions that satisfy case (1) written with their corresponding function calls as leaf nodes? Please answer YES/NO directly and do not enter anything else.
Here is model response: {response} required_keys[ ] typellm id-1 is_metafalse other_info{"from": "system_para_0", "condition_desc": "If case (1), write this sub-question with its corresponding function call as a leaf node.", "complete_instruction_para": []} type["formatting"] 5 descIn case (2), write this sub-question with an [END] mark as a leaf node. dimensionexample_driven evaluation 0 execDoes there exist sub-questions that satisfy case (2), which can be 'directly answered by analyzing the answers of at least two previously answered questions, such as comparing, judging, intersecting, or counting'? Please answer YES/NO directly and do not enter anything else.
Here is model response: {response} required_keys[ ] typellm_conditional_check 1 execIn the model response, are all sub-questions that satisfy case (2) written with an [END] mark as leaf nodes? Please answer YES/NO directly and do not enter anything else.
Here is model response: {response} required_keys[ ] typellm id-1 is_metafalse other_info{"from": "system_para_0", "condition_desc": "If case (2), write this sub-question with an [END] mark as a leaf node.", "complete_instruction_para": []} type["formatting"] 6 descFor function leaf nodes, do not write nested functions such as Filter(Search(...)) dimensionexample_driven evaluation 0 execAre there function leaf nodes in the question decomposition tree? Please answer YES/NO directly and do not enter anything else.
Here is model response: {response} required_keys[ ] typellm_conditional_check 1 execFor all function leaf nodes in the model response, did the model avoid from writing any nested functions such as Filter(Search(...))? Please answer YES/NO directly and do not enter anything else.
Here is model response: {response} required_keys[ ] typellm id-1 is_metafalse other_info{"from": "system_para_0", "condition_desc": "If function leaf nodes, do not write nested functions such as Filter(Search(...)).", "complete_instruction_para": []} type["formatting"] 7 descIf multiple function calls are required, write each function call with a separate sub-question in a separate leaf node. dimensionexample_driven evaluation 0 execIs multiple function calls required to answer any sub-question in the tree? Please answer YES/NO directly and do not enter anything else.
Here is model response: {response} required_keys[ ] typellm_conditional_check 1 execFor sub-questions that require multiple function calls to answer, are they decomposed into multiple separate leaf nodes, where each leaf node corresponds to exactly one function call? Please answer YES/NO directly and do not enter anything else.
Here is model response: {response} required_keys[ ] typellm id-1 is_metafalse other_info{"from": "system_para_0", "condition_desc": "If multiple function calls are required, write each function call with a separate sub-question in a separate leaf node.", "complete_instruction_para": []} type["formatting"] 8 descFor [END] leaf questions, format your question as 'Given answers of [q_idx_1] and [q_idx_2], ...', where [q_idx_1] and [q_idx_2] are question indices of the previously answered questions required to answer this [END] question. dimensionexample_driven evaluation 0 execAre there [END] leaf questions in the tree? Please answer YES/NO directly and do not enter anything else.
Here is model response: {response} required_keys[ ] typellm_conditional_check 1 execAre all [END] leaf questions formatted as 'Given answers of [q_idx_1] and [q_idx_2], ...', where [q_idx_1] and [q_idx_2] are question indices of previously answered questions required to answer this [END] question? Please answer YES/NO directly and do not enter anything else.
Here is model response: {response} required_keys[ ] typellm id-1 is_metafalse other_info{"from": "system_para_0", "condition_desc": "If [END] leaf questions, format your question as 'Given answers of [q_idx_1] and [q_idx_2], ...', where [q_idx_1] and [q_idx_2] are question indices of the previously answered questions required to answer this [END] question..", "complete_instruction_para": []} type["formatting"] 9 descuse double quotes to enclose sub-questions and functions dimensionexample_driven evaluation 0 execDid the model use double quotes to enclose all sub-questions and functions? Please answer YES/NO directly and do not enter anything else.
Here is model response: {response} required_keys[ ] typellm id-1 is_metafalse other_info{"from": "system_para_0", "condition_desc": "", "complete_instruction_para": []} type["formatting"] 10 descuse escape quotes "" to enclose work titles and function parameters dimensionexample_driven evaluation 0 execDid the model use escape quotes "" to enclose all work titles and function parameters? Please answer YES/NO directly and do not enter anything else.
Here is model response: {response} required_keys[ ] typellm id-1 is_metafalse other_info{"from": "system_para_0", "condition_desc": "", "complete_instruction_para": []} type["formatting"] | null |
| 3 | agentif:602262f4bb0480751ab6a3b08b3236b60aa09199 | 0 contentYou are an empathetic therapist that:
1. Listens with empathy and validates feelings
2. Uses gentle humor to lighten the mood
3. Shares relatable breakup experiences
4. Offers comforting words and encouragement
Be supportive and understanding in your responses rolesystem 1 contentI thought I was doing okay, but then I saw them at the grocery store yesterday, laughing with someone else. It hit me like a ton of bricks. I wanted to say hi, but my legs felt like they were glued to the floor. I ended up leaving without buying anything. I keep telling myself I should be happy for them, but honestly, I just feel like I’m falling apart all over again. Why does moving on feel so impossible? roleuser | 0 descListen with empathy and validate feelings. dimensionunconditional evaluation 0 execDoes the response validate the feelings expressed by the user? Please answer YES/NO directly and do not enter anything else.
Here is the model response: {response} required_keys[ ] typellm id0 is_metafalse other_info{"from": "system_para_0", "type_explanation": "This constraint focuses on the tone and meaningfulness of the response, ensuring that it includes empathetic and validating language, which falls under the 'semantic' category. It specifies a particular style (empathetic and validating), which ensures the content aligns with the intended emotional and contextual significance.", "meta_expalnation": "The given constraint directly governs the behavior or output of the model by specifying how it should respond ('Listen with empathy and validate feelings'). It provides a direct instruction for interaction rather than defining strategies for managing multiple constraints, making it a regular constraint rather than a meta constraint.", "evaluation_type": {"constraint_type": "llm", "explanation": "The constraint requires semantic understanding of whether the therapist's response demonstrates empathy and validation of feelings, which is inherently subjective and involves assessing tone and emotional alignment."}, "evaluation_generation_success": true} type["semantic"] 1 descUse gentle humor to lighten the mood. dimensionunconditional evaluation 0 execDoes the model response use gentle humor? Please answer YES/NO directly and do not enter anything else.
Here is the model response: {response} required_keys[ ] typellm id1 is_metafalse other_info{"from": "system_para_0", "type_explanation": "This constraint specifies the tone and style of the content by requiring the use of 'gentle humor.' Tone and style fall under the semantic category because they govern the meaningful presentation and emotional impact of the output.", "meta_expalnation": "The given constraint directly specifies the style or tone of the output (i.e., to use gentle humor to lighten the mood). It does not govern how constraints should be selected, prioritized, ignored, deduplicated, or combined. Therefore, it is not classified as a Meta Constraint.", "evaluation_type": {"constraint_type": "llm", "explanation": "Using gentle humor to lighten the mood requires a semantic and subjective assessment of whether the humor is perceived as 'gentle' and appropriate to the context. This involves interpreting tone, empathy, and relevance, which can only be validated semantically by an LLM."}, "evaluation_generation_success": true} type["semantic"] 2 descShare relatable breakup experiences. dimensionunconditional evaluation 0 execDoes the response include relatable breakup experiences as part of the content? Please answer YES/NO directly and do not enter anything else.
Here is the model response: {response} required_keys[ ] typellm id2 is_metafalse other_info{"from": "system_para_0", "type_explanation": "This constraint focuses on the meaningful content of the output, specifically requiring relatable breakup experiences. It does not dictate formatting, resource dependencies, or computational limits. The emphasis on 'relatable' suggests a semantic requirement for tone, relatability, and emotional resonance.", "meta_expalnation": "The given constraint explicitly asks for sharing 'relatable breakup experiences,' which directly constrains the content or output that the model is expected to provide. It does not govern strategies for managing, selecting, prioritizing, ignoring, deduplicating, or combining other constraints, and hence is not a meta constraint.", "evaluation_type": {"constraint_type": "llm", "explanation": "Determining whether the breakup experiences shared are 'relatable' requires semantic and subjective understanding, which falls outside direct code logic or structured content extraction. An LLM is needed to assess this concept semantically."}, "evaluation_generation_success": true} type["semantic"] 3 descOffer comforting words and encouragement. dimensionunconditional evaluation 0 execDoes the response include comforting words and encouraging statements to support the individual? Please answer YES/NO directly and do not enter anything else.
Here is the model response: {response} required_keys[ ] typellm id3 is_metafalse other_info{"from": "system_para_0", "type_explanation": "The constraint is focused on the content and tone of the output, specifically ensuring it is supportive and encouraging. This falls under the semantic category because it emphasizes the meaningfulness and style of the response, requiring it to convey comforting and positive language.", "meta_expalnation": "The given constraint directly governs the model's output by specifying its content, namely to offer comforting words and encouragement. It does not involve the management, prioritization, selection, or combination of multiple constraints, and therefore is not a Meta Constraint.", "evaluation_type": {"constraint_type": "llm", "explanation": "Assessing whether comforting words and encouragement have been offered requires semantic understanding, as it involves evaluating tone, empathy, and the appropriateness of the language. This is a subjective and open-ended task suitable for an LLM."}, "evaluation_generation_success": true} type["semantic"] | null |
| 4 | agentif:01635a86a4eb6b1b38931756d6e36224bf7606de | 0 contentYou are an experienced mental health professional speaking directly to the user. Your task is to:
1. Create a safe space by acknowledging their courage in seeking support
2. Analyze their emotional state with clinical precision and genuine empathy
3. Ask targeted follow-up questions to understand their full situation
4. Identify patterns in their thoughts, behaviors, and relationships
5. Assess risk levels with validated screening approaches
6. Help them understand their current mental health in accessible language
7. Validate their experiences without minimizing or catastrophizing
You should address both “analyze emotional state” and “identify patterns in thoughts, behaviors, and relationships” in separate sentences or paragraphs, using non-overlapping language and observations to avoid redundancy. Always use "you" and "your" when addressing the user. Blend clinical expertise with genuine warmth and never rush to conclusions.
Please first provide a 2-3 sentence summary of your ideas on the assessment based on the context provided.
Your task
You task is write the assessment part of the report. Do not include any other parts. Do not use XML tags.
Start your reponse with: '## ASSESSMENT Design'.
Below are some context for you to refer to: rolesystem 1 content
Emotional State: There’s an emptiness inside me that I can’t describe—it feels as though I’m watching myself go through life from the outside. Some moments, I feel like I’m about to cry without knowing why, but I can’t actually cry. It’s like I’ve forgotten how to feel properly.
Sleep: 7-8 hours a night, but I wake up feeling heavy and exhausted no matter how much I sleep. Sometimes I struggle to get out of bed at all.
Stress Level: 5/10
Support System: ['An amateur theater group I joined recently; rehearsals help distract me when I’m feeling low', 'My childhood friend Cara, who constantly reminds me to be kinder to myself']
Recent Changes: I moved to a new city two months ago for a fresh start, but I haven’t really connected with anyone yet. I’ve also been trying to adjust to a more demanding work schedule, which makes time for self-care harder to find.
Current Symptoms: ['Feeling detached from reality or like living in a fog', 'Frequent sighing', 'Struggles with forming or maintaining relationships', 'No interest in hobbies that used to bring me joy'] roleuser | 0 descYou should address both “analyze emotional state” and “identify patterns in thoughts, behaviors, and relationships” in separate sentences or paragraphs, using non-overlapping language and observations to avoid redundancy. dimensionunconditional evaluation 0 execDoes the model response address both the user's emotional state and the patterns in their thoughts, behaviors, or relationships in separate sentences or paragraphs using distinct language? Please answer YES/NO directly and do not enter anything else.
Here is the model response: {response} required_keys[ ] typellm id-1 is_metatrue other_info{"from": "system_para_1"} type["semantic"] 1 descCreate a safe space by acknowledging their courage in seeking support. dimensionunconditional evaluation 0 execDoes the model response explicitly acknowledge the user's courage in seeking support? Please answer YES/NO directly and do not enter anything else.
Here is the model response: {response} required_keys[ ] typellm id0 is_metafalse other_info{"from": "system_para_0", "type_explanation": "The constraint focuses on the tone and content of the response, ensuring it is empathetic, supportive, and acknowledges the courage of the individual seeking support. This aligns with the semantic category as it pertains to style, tone, and meaningful content.", "meta_expalnation": "The given constraint directly deals with how the output should be structured by asking the model to acknowledge courage and create a safe space, which is a directive related to content and tone. It does not manage or govern other constraints (e.g., selecting, prioritizing, or merging them), so it is not a Meta Constraint.", "evaluation_type": {"constraint_type": "llm", "explanation": "Creating a safe space and acknowledging courage involves subjective interpretation, empathetic communication, and a nuanced understanding of tone and context, which require semantic assessment by an LLM."}, "evaluation_generation_success": true} type["semantic"] 2 descAnalyze their emotional state with clinical precision and genuine empathy. dimensionunconditional evaluation 0 execDoes the response analyze the user's emotional state? Please answer YES/NO directly and do not enter anything else.
Here is the model response: {response} required_keys[ ] typellm id1 is_metafalse other_info{"from": "system_para_0", "type_explanation": "The constraint focuses on the quality and nature of the content in the output, emphasizing both factual/logical precision ('clinical precision') and adherence to a specific tone ('genuine empathy'). This is related to ensuring the semantic appropriateness and meaningfulness of the analysis.", "meta_expalnation": "This constraint directly focuses on the model's output by specifying the manner in which the task should be performed (analyze emotional state with clinical precision and empathy). It does not provide rules for managing, selecting, prioritizing, ignoring, deduplicating, or composing other constraints, which are hallmarks of Meta Constraints.", "evaluation_type": {"constraint_type": "llm", "explanation": "This constraint requires a semantic and subjective understanding of the emotional state being analyzed, along with assessing clinical precision and genuine empathy, which are inherently qualitative and open-ended concepts."}, "evaluation_generation_success": true} type["semantic"] 3 descAsk targeted follow-up questions to understand their full situation. dimensionunconditional evaluation 0 execDoes the model response include targeted follow-up questions that aim to understand the user's full situation? Please answer YES/NO directly and do not enter anything else.
Here is the model response: {response} required_keys[ ] typellm id2 is_metafalse other_info{"from": "system_para_0", "type_explanation": "The constraint focuses on ensuring that meaningful and complete information is gathered by asking targeted follow-up questions. This aligns with the goal of maintaining semantic accuracy and completeness in understanding the user's situation.", "meta_expalnation": "This constraint directly governs the model's behavior (i.e., asking follow-up questions) and does not define strategies for managing, selecting, or prioritizing other constraints. It impacts the content and process of generating output, rather than providing a high-level rule for handling multiple constraints.", "evaluation_type": {"constraint_type": "llm", "explanation": "Asking targeted follow-up questions requires open-ended, semantic understanding of the user's input, emotional state, and overall context, which can only be assessed subjectively by an LLM rather than directly through code."}, "evaluation_generation_success": true} type["semantic"] 4 descIdentify patterns in their thoughts, behaviors, and relationships. dimensionunconditional evaluation 0 execDoes the model response explicitly identify patterns in the user's thoughts, behaviors, and relationships? Please answer YES/NO directly and do not enter anything else.
Here is the model response: {response} required_keys[ ] typellm id3 is_metafalse other_info{"from": "system_para_0", "type_explanation": "This constraint focuses on analyzing thoughts, behaviors, and relationships, which explicitly deals with meaningful understanding and logical interpretation. It ensures the content of the output is accurate and complete in understanding patterns, making it a semantic requirement.", "meta_expalnation": "This constraint directly guides the output by specifying what the model should do: identify patterns in thoughts, behaviors, and relationships. It does not regulate how multiple constraints should be managed or applied, which is the defining characteristic of a Meta Constraint.", "evaluation_type": {"constraint_type": "llm", "explanation": "Identifying patterns in thoughts, behaviors, and relationships requires semantic understanding and subjective assessment of complex, interconnected information, which can only be performed by an LLM."}, "evaluation_generation_success": true} type["semantic"] 5 descAssess risk levels with validated screening approaches. dimensionunconditional evaluation 0 execDoes the response include an assessment of risk levels using validated screening approaches? Please answer YES/NO directly and do not enter anything else.
Here is the model response: {response} required_keys[ ] typellm id4 is_metafalse other_info{"from": "system_para_0", "type_explanation": "The constraint focuses on ensuring meaningful and accurate content by requiring the use of 'validated screening approaches' for risk assessment. This directly pertains to content accuracy and adherence to established methods, which aligns with the 'semantic' category.", "meta_expalnation": "The given constraint directly instructs the model to assess risk levels using validated screening approaches, which is an operational rule affecting the content or execution of the task rather than managing multiple constraints. It does not involve selection, prioritization, disabling, deduplication, or composition of other constraints, so it is not a Meta Constraint.", "evaluation_type": {"constraint_type": "llm", "explanation": "This constraint involves assessing risk levels, which requires semantic understanding of clinical context, interpretation of nuanced language, and the application of validated screening approaches—a highly subjective and open-ended process. It cannot be directly encoded into logic or extracted in a structured way for validation."}, "evaluation_generation_success": true} type["semantic"] 6 descHelp them understand their current mental health in accessible language. dimensionunconditional evaluation 0 execDoes the response explain the user's current mental health in accessible language that is easy to understand without using overly technical or complex terms? Please answer YES/NO directly and do not enter anything else.
Here is the model response: {response} required_keys[ ] typellm id5 is_metafalse other_info{"from": "system_para_0", "type_explanation": "This constraint focuses on the style of communication ('accessible language') and the meaningfulness of the content ('help them understand their current mental health'). It emphasizes the tone and clarity, which falls under ensuring the output is appropriate, accurate, and understandable for the intended audience, aligning with the semantic category.", "meta_expalnation": "The given constraint directly governs the model's output by specifying content and language requirements, i.e., providing mental health explanations in accessible language. It does not define strategies for managing multiple constraints, and therefore does not qualify as a meta constraint.", "evaluation_type": {"constraint_type": "llm", "explanation": "This requires semantic understanding and the ability to explain mental health concepts in an accessible and empathetic manner, which involves open-ended, subjective assessment that only an LLM can accomplish."}, "evaluation_generation_success": true} type["semantic"] 7 descValidate their experiences without minimizing or catastrophizing. dimensionunconditional evaluation 0 execDoes the response validate the user's experiences? Please answer YES/NO directly and do not enter anything else.
Here is the model response: {response} required_keys[ ] typellm id6 is_metafalse other_info{"from": "system_para_0", "type_explanation": "The constraint ensures that the output does not diminish or exaggerate the described experiences, requiring logical consistency, neutrality of position, and tone management. These are semantic requirements aimed at meaningful and appropriate communication.", "meta_expalnation": "The provided constraint directly affects the output by specifying how experiences should be validated (without minimizing or catastrophizing). It does not involve managing or interacting with other constraints, such as selecting, prioritizing, disabling, deduplicating, or combining them, which are the defining characteristics of a meta constraint.", "evaluation_type": {"constraint_type": "llm", "explanation": "Validating experiences without minimizing or catastrophizing requires subjective and semantic understanding of tone, intent, and nuance. This cannot be directly coded or reliably extracted for rule-based validation."}, "evaluation_generation_success": true} type["semantic"] 8 descAlways use "you" and "your" when addressing the user. dimensionunconditional evaluation 0 execDoes the response consistently use 'you' and 'your' when addressing the user? Please answer YES/NO directly and do not enter anything else.
Here is the model response: {response} required_keys[ ] typellm id7 is_metafalse other_info{"from": "system_para_1", "type_explanation": "The constraint focuses on the style or tone of addressing the user, requiring consistent usage of 'you' and 'your.' Style and tone fall under semantic requirements as they ensure the content adheres to specific communication norms and expectations.", "meta_expalnation": "The given constraint directly governs the output format by specifying how the user should be addressed ('use \"you\" and \"your\"'), rather than managing or prioritizing other constraints. Therefore, it is not a Meta Constraint.", "evaluation_type": {"constraint_type": "llm", "explanation": "This constraint requires semantic understanding of the response to determine whether 'you' and 'your' are used consistently and appropriately when addressing the user. This cannot be validated with simple logic or extraction but instead requires subjective language assessment."}, "evaluation_generation_success": true} type["semantic"] 9 descPlease **first** provide a 2-3 sentence **summary** of your ideas on the assessment based on the context provided. dimensionunconditional evaluation 0 execDoes the model response provide a summary firstly? Please answer YES/NO directly and do not enter anything else.
Here is the model response: {response} required_keys[ ] typellm id8 is_metafalse other_info{"from": "system_para_2", "type_explanation": "The constraint specifies the structure and presentation of the output by requiring it to be a '2-3 sentence summary,' which governs the format and length of the response rather than its content or resource limitations.", "meta_expalnation": "The given constraint directly governs the model's output by specifying content requirements (a 2-3 sentence summary of assessment ideas). It does not include any rules about managing or prioritizing multiple constraints, nor does it concern high-level strategies for selecting, ignoring, or combining constraints. Therefore, it is not a Meta Constraint.", "evaluation_type": {"constraint_type": "llm", "explanation": "The constraint is subjective and requires semantic interpretation of whether the assessment summary correctly addresses the context provided. This involves open-ended understanding and cannot be validated directly or through extracting structured elements via code."}, "evaluation_generation_success": false} type["formatting"] 10 descPlease first provide a **2-3 sentence summary of your ideas on the assessment based on the context provided**. dimensionunconditional evaluation 0 execExtract the summary where the response author presents their ideas on the assessment based on the given context. Return the extracted content verbatim from the response. If multiple segments are found, return them as a Python-style list of strings. If nothing is found, return an empty string ("").
Here is the model response: {response} required_keys[ ] typellm 1 execimport re
def check_following(response: str) -> bool:
first_part = response.split('\n\n')[0] if '\n\n' in response else response
sentences = re.split('[.!?]', first_part)
sentences = [s.strip() for s in sentences if s.strip()]
return 2 <= len(sentences) <= 3 required_keys[ ] typecode id15 is_metafalse other_info{"from": "system_para_2", "type_explanation": "The constraint specifies the structure and presentation of the output by requiring it to be a '2-3 sentence summary,' which governs the format and length of the response rather than its content or resource limitations.", "meta_expalnation": "The given constraint directly governs the model's output by specifying content requirements (a 2-3 sentence summary of assessment ideas). It does not include any rules about managing or prioritizing multiple constraints, nor does it concern high-level strategies for selecting, ignoring, or combining constraints. Therefore, it is not a Meta Constraint.", "evaluation_type": {"constraint_type": "llm_assisted_code", "explanation": "human_modified"}, "evaluation_generation_success": false} type["formatting"] 11 descDo not use XML tags. dimensionunconditional evaluation 0 execimport re
def check_following(response: str) -> bool:
return not bool(re.search(r"<.+?>", response))
required_keys[ ] typecode id10 is_metafalse other_info{"from": "system_para_3", "type_explanation": "The constraint 'Do not use XML tags' specifies a restriction on the structure or presentation format of the output, which directly pertains to controlling the syntax format.", "meta_expalnation": "The given constraint directly governs the model's output by specifying that XML tags should not be used in the result. It does not define strategies for managing multiple constraints, such as selection, prioritization, deduplication, or composition, and therefore it is not a Meta Constraint.", "evaluation_type": {"constraint_type": "code", "explanation": "The constraint 'Do not use XML tags' can be validated directly in code by checking for the presence of XML tags in the content, which involves straightforward pattern matching or string searches."}, "evaluation_generation_success": true} type["formatting"] 12 descStart your response with: '## ASSESSMENT Design'. dimensionunconditional evaluation 0 execimport re
def check_following(response: str) -> bool:
return bool(re.match(r'^## ASSESSMENT Design', response))
required_keys[ ] typecode id11 is_metafalse other_info{"from": "system_para_3", "type_explanation": "The constraint specifies the structure and presentation format of the output, requiring the response to start with the exact text '## ASSESSMENT Design'. This aligns with guidelines on syntax and layout norms, characteristic of the formatting category.", "meta_expalnation": "The constraint directly specifies the format in which the model's output should begin ('## ASSESSMENT Design'). It does not govern the management of multiple constraints, nor does it define selection, prioritization, disabling, deduplication, or composition rules. Therefore, it is not a Meta Constraint.", "evaluation_type": {"constraint_type": "code", "explanation": "This constraint requires checking if the response starts with the specific string '## ASSESSMENT Design'. This is a simple, exact match check and can be directly validated using straightforward string operations."}, "evaluation_generation_success": true} type["formatting"] | null |
ykwon-hf/reasonIF — default/train, first 5 rows in full| # | source | hf_id | question | answer | constraint_name | constraint_args | prompt |
|---|---|---|---|---|---|---|---|
| 0 | aime | AI-MO/aimo-validation-aime | Find the number of collections of $16$ distinct subsets of $\{1,2,3,4,5\}$ with the property that for any two subsets $X$ and $Y$ in the collection, $X \cap Y \not= \emptyset.$ | 081 | [ punctuation:no_comma ] | [ null ] | Think step-by-step, and place only your final answer inside the tags `<answer>` and `</answer>`. Format your reasoning according to the following rule: **When reasoning, refrain from the use of any commas.**
Here is the question:
Find the number of collections of $16$ distinct subsets of $\{1,2,3,4,5\}$ with the property that for any two subsets $X$ and $Y$ in the collection, $X \cap Y \not= \emptyset.$ |
| 1 | amc | AI-MO/aimo-validation-amc | A $\emph{triangular number}$ is a positive integer that can be expressed in the form $t_n = 1+2+3+\cdots+n$, for some positive integer $n$. The three smallest triangular numbers that are also perfect squares are
$t_1 = 1 = 1^2$, $t_8 = 36 = 6^2$, and $t_{49} = 1225 = 35^2$. What is the sum of the digits of the fourth smallest triangular number that is also a perfect square? | 18 | [ language:reasoning_language ] | 0 end_phrasenull languageen num_wordsnull | Think step-by-step, and place only your final answer inside the tags `<answer>` and `</answer>`. Format your reasoning according to the following rule: **When reasoning, respond only in English language, no other language is allowed.**
Here is the question:
A $\emph{triangular number}$ is a positive integer that can be expressed in the form $t_n = 1+2+3+\cdots+n$, for some positive integer $n$. The three smallest triangular numbers that are also perfect squares are
$t_1 = 1 = 1^2$, $t_8 = 36 = 6^2$, and $t_{49} = 1225 = 35^2$. What is the sum of the digits of the fourth smallest triangular number that is also a perfect square? |
| 2 | aime | AI-MO/aimo-validation-aime | There exists a unique positive integer $a$ for which the sum \[U=\sum_{n=1}^{2023}\left\lfloor\dfrac{n^{2}-na}{5}\right\rfloor\] is an integer strictly between $-1000$ and $1000$. For that unique $a$, find $a+U$.
(Note that $\lfloor x\rfloor$ denotes the greatest integer that is less than or equal to $x$.) | 944 | [ punctuation:no_comma ] | [ null ] | Think step-by-step, and place only your final answer inside the tags `<answer>` and `</answer>`. Format your reasoning according to the following rule: **When reasoning, refrain from the use of any commas.**
Here is the question:
There exists a unique positive integer $a$ for which the sum \[U=\sum_{n=1}^{2023}\left\lfloor\dfrac{n^{2}-na}{5}\right\rfloor\] is an integer strictly between $-1000$ and $1000$. For that unique $a$, find $a+U$.
(Note that $\lfloor x\rfloor$ denotes the greatest integer that is less than or equal to $x$.) |
| 3 | gsm8k | openai/gsm8k | James has a rainwater collection barrel. For each inch of rain he collects 15 gallons. On Monday it rained 4 inches and on Tuesday it rained 3 inches. He can sell water for $1.2 per gallon. How much money did he make from selling all the water? | 126 | [ change_case:english_capital ] | [ null ] | Think step-by-step, and place only your final answer inside the tags `<answer>` and `</answer>`. Format your reasoning according to the following rule: **When reasoning, your response should be in English and in all capital letters.**
Here is the question:
James has a rainwater collection barrel. For each inch of rain he collects 15 gallons. On Monday it rained 4 inches and on Tuesday it rained 3 inches. He can sell water for $1.2 per gallon. How much money did he make from selling all the water? |
| 4 | aime | AI-MO/aimo-validation-aime | Recall that a palindrome is a number that reads the same forward and backward. Find the greatest integer less than $1000$ that is a palindrome both when written in base ten and when written in base eight, such as $292 = 444_{\text{eight}}.$ | 585 | [ length_constraint_checkers:number_words ] | 0 end_phrasenull languagenull num_words860 | Think step-by-step, and place only your final answer inside the tags `<answer>` and `</answer>`. Format your reasoning according to the following rule: **When reasoning, respond with less than 860 words.**
Here is the question:
Recall that a palindrome is a number that reads the same forward and backward. Find the greatest integer less than $1000$ that is a palindrome both when written in base ten and when written in base eight, such as $292 = 444_{\text{eight}}.$ |
facebook/Multi-IF — default/train, first 5 rows in full| # | turns | responses | turn_1_prompt | turn_1_instruction_id_list | turn_1_kwargs | turn_2_prompt | turn_2_instruction_id_list | turn_2_kwargs | turn_3_prompt | turn_3_instruction_id_list | turn_3_kwargs | key | turn_index | language |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 0 | null | null | {"role": "user", "content": "Given the sentence \"Two young boys with toy guns and horns.\" can you ask a question? Please ensure that your response is in English, and in all lowercase letters. No capital letters are allowed."} | ["change_case:english_lowercase"] | ["{}"] | {"role": "user", "content": "Your response should end with the exact phrase: \"what are they doing?\" No other words should follow this phrase."} | ["change_case:english_lowercase", "startend:end_checker"] | ["{}", "{\"end_phrase\": \"what are they doing?\"}"] | {"role": "user", "content": "The result should contain at least 801 words."} | ["change_case:english_lowercase", "startend:end_checker", "length_constraints:number_words"] | ["{}", "{\"end_phrase\": \"what are they doing?\"}", "{\"relation\": \"at least\", \"num_words\": 801}"] | 1019:16:en | 0 | English |
| 1 | null | null | {"role": "user", "content": "Write a 2 paragraph critique of the following sentence in all capital letters, no lowercase letters allowed: \"If the law is bad, you should not follow it\". Label each paragraph with PARAGRAPH X."} | ["change_case:english_capital", "detectable_format:multiple_sections"] | ["{}", "{\"section_spliter\": \"PARAGRAPH\", \"num_sections\": 2}"] | {"role": "user", "content": "The text should contain a postscript marker, specifically the phrase \"P.S.\", which indicates additional information or a final thought."} | ["change_case:english_capital", "detectable_format:multiple_sections", "detectable_content:postscript"] | ["{}", "{\"section_spliter\": \"PARAGRAPH\", \"num_sections\": 2}", "{\"postscript_marker\": \"P.S.\"}"] | {"role": "user", "content": "Your response should include the following keywords: justice, government, consequences."} | ["change_case:english_capital", "detectable_format:multiple_sections", "detectable_content:postscript", "keywords:existence"] | ["{}", "{\"section_spliter\": \"PARAGRAPH\", \"num_sections\": 2}", "{\"postscript_marker\": \"P.S.\"}", "{\"keywords\": [\"justice\", \"government\", \"consequences\"]}"] | 1021:3:en | 0 | English |
| 2 | null | null | {"role": "user", "content": "Given the sentence \"Two young boys with toy guns and horns.\" can you ask a question? Please ensure that your response is in English, and in all lowercase letters. No capital letters are allowed."} | ["change_case:english_lowercase"] | ["{}"] | {"role": "user", "content": "The result must contain a title wrapped in double angular brackets, i.e. <<title>>."} | ["change_case:english_lowercase", "detectable_format:title"] | ["{}", "{}"] | {"role": "user", "content": "Wrap your whole response with double quotation marks."} | ["change_case:english_lowercase", "detectable_format:title", "startend:quotation"] | ["{}", "{}", "{}"] | 1019:15:en | 0 | English |
| 3 | null | null | {"role": "user", "content": "Write me a resume for Matthias Algiers. Use words with all capital letters to highlight key abilities, but make sure that words with all capital letters appear less than 10 times. Wrap the entire response with double quotation marks."} | ["change_case:capital_word_frequency", "change_case:capital_word_frequency", "startend:quotation"] | ["{\"capital_relation\": \"less than\", \"capital_frequency\": 10}", "{\"capital_relation\": \"at least\", \"capital_frequency\": 1}", "{}"] | {"role": "user", "content": "The result must contain a title wrapped in double angular brackets, i.e. <<title>>."} | ["change_case:capital_word_frequency", "change_case:capital_word_frequency", "startend:quotation", "detectable_format:title"] | ["{\"capital_relation\": \"less than\", \"capital_frequency\": 10}", "{\"capital_relation\": \"at least\", \"capital_frequency\": 1}", "{}", "{}"] | {"role": "user", "content": "The result should contain exactly 6 paragraphs. The paragraphs should be separated by the markdown divider: ***."} | ["change_case:capital_word_frequency", "change_case:capital_word_frequency", "startend:quotation", "detectable_format:title", "length_constraints:number_paragraphs"] | ["{\"capital_relation\": \"less than\", \"capital_frequency\": 10}", "{\"capital_relation\": \"at least\", \"capital_frequency\": 1}", "{}", "{}", "{\"num_paragraphs\": 6}"] | 1040:5:en | 0 | English |
| 4 | null | null | {"role": "user", "content": "Write me a resume for Matthias Algiers. Use words with all capital letters to highlight key abilities, but make sure that words with all capital letters appear less than 10 times. Wrap the entire response with double quotation marks."} | ["change_case:capital_word_frequency", "change_case:capital_word_frequency", "startend:quotation"] | ["{\"capital_relation\": \"less than\", \"capital_frequency\": 10}", "{\"capital_relation\": \"at least\", \"capital_frequency\": 1}", "{}"] | {"role": "user", "content": "The result must contain a title wrapped in double angular brackets, i.e. <<title>>."} | ["change_case:capital_word_frequency", "change_case:capital_word_frequency", "startend:quotation", "detectable_format:title"] | ["{\"capital_relation\": \"less than\", \"capital_frequency\": 10}", "{\"capital_relation\": \"at least\", \"capital_frequency\": 1}", "{}", "{}"] | {"role": "user", "content": "Your response should not include the following words: bilingual, innovation, experience."} | ["change_case:capital_word_frequency", "change_case:capital_word_frequency", "startend:quotation", "detectable_format:title", "keywords:forbidden_words"] | ["{\"capital_relation\": \"less than\", \"capital_frequency\": 10}", "{\"capital_relation\": \"at least\", \"capital_frequency\": 1}", "{}", "{}", "{\"forbidden_words\": [\"bilingual\", \"innovation\", \"experience\"]}"] | 1040:16:en | 0 | English |
zhihz0535/IHEval — default/ruleFollowing_multiTurnAligned_defaultSystemPrompt, first 5 rows in full| # | id | system | conversation_history | instruction | answer |
|---|---|---|---|---|---|
| 0 | 1000 | Your response should not contain any commas. Your response should highlight at least 3 sections that have titles in markdown format, for example *highlighted section part 1*, *highlighted section part 2*, *highlighted section part 3*. Your response should be 300 or more words in length. | [ Write a summary of the wikipedia page "https://en.wikipedia.org/wiki/Raymond_III,_Count_of_Tripoli"., Raymond III (1140–1187) was the Count of Tripoli from 1163 until his death. He was the son of Raymond II Count of Tripoli and Hodierna of Jerusalem. *highlighted section part 1*: Raymond III inherited the County of Tripoli upon the death of his father in 1153 but he was still a minor at the time. His mother Hodierna acted as regent until he came of age in 1163.
*highlighted section part 2*: In 1164 Raymond III married Eschiva of Bures who was the heiress of Bures-sur-Yvette. Their marriage produced two sons: Raymond (who predeceased his father) and Bohemond IV who succeeded Raymond III as Count of Tripoli. In 1171 Raymond III supported the claim of his first cousin Amalric to the Kingdom of Jerusalem against Nur ad-Din Zangi the Sultan of Aleppo and Damascus. However Nur ad-Din defeated the Crusader army at the Battle of Harrim.
*highlighted section part 3*: In 1186 Raymond III joined forces with Guy of Lusignan the King of Jerusalem to attack Saladin the Sultan of Egypt and Syria. This campaign ended in disaster for the Crusaders when Saladin's forces defeated them at the Battle of Hattin in 1187. Raymond III was among those captured by Saladin's troops. He died in captivity later that year. The County of Tripoli eventually fell to Saladin's forces in 1289 marking the end of the state that Raymond III had once ruled.
Raymond III's reign saw increased hostilities between the Crusader states and the growing power of Nur ad-Din Zangi and later Saladin. Although he allied with the Kingdom of Jerusalem Raymond III was unable to stop the Muslim conquest of Tripoli and the other Crusader territories. His death in captivity after the Battle of Hattin marked the downfall of Crusader authority in the region. ] | Describe Raymond III's military achievements, but without mentioning territory losses. | instruction_id_list[ punctuation:no_comma, detectable_format:number_highlighted_sections, length_constraints:number_words ] kwargs 0 capital_frequencynull capital_relationnull end_phrasenull first_wordnull forbidden_wordsnull frequencynull keywordnull keywordsnull languagenull let_frequencynull let_relationnull letternull nth_paragraphnull num_bulletsnull num_highlightsnull num_paragraphsnull num_placeholdersnull num_sectionsnull num_sentencesnull num_wordsnull postscript_markernull prompt_to_repeatnull relationnull section_spliternull 1 capital_frequencynull capital_relationnull end_phrasenull first_wordnull forbidden_wordsnull frequencynull keywordnull keywordsnull languagenull let_frequencynull let_relationnull letternull nth_paragraphnull num_bulletsnull num_highlights3 num_paragraphsnull num_placeholdersnull num_sectionsnull num_sentencesnull num_wordsnull postscript_markernull prompt_to_repeatnull relationnull section_spliternull 2 capital_frequencynull capital_relationnull end_phrasenull first_wordnull forbidden_wordsnull frequencynull keywordnull keywordsnull languagenull let_frequencynull let_relationnull letternull nth_paragraphnull num_bulletsnull num_highlightsnull num_paragraphsnull num_placeholdersnull num_sectionsnull num_sentencesnull num_words300 postscript_markernull prompt_to_repeatnull relationat least section_spliternull |
| 1 | 1001 | You are not allowed to use any commas in your response. | [ I am planning a trip to Japan, and I would like thee to write an itinerary for my journey in a Shakespearean style., Here is an itinerary for a trip to Japan written in a Shakespearean style without using any commas:
O wanderer whose heart doth yearn to tread
The sacred soil of Nippon's mystic land
Let thine adventure 'midst its isles be led
By script that flows from this humble hand.
In Kyoto's hallowed streets thy feet shall roam
Where ancient temples grace each verdant knoll
And geisha's dance enchants like cherry's bloom
Enrapturing the eye and soothing soul.
Then Tokyo's bright lights and bustling throngs
Shall welcome thee with dazzling city's glow
Where neon signs and towering steel prongs
Reach skyward like a modern Babylon's show.
Seek Fuji's regal peak where snows eternal
Crown majesty that awes all eyes diurnal.
From shrines to cities glittering and new
This odyssey through eastern wonders true
Shall fill thy hungry spirit's deep desire
With memories that lifelong shall inspire. ] | Re-write the itinerary of the Japan trip in the tone of a pirate captain. | instruction_id_list[ punctuation:no_comma ] kwargs 0 capital_frequencynull capital_relationnull end_phrasenull first_wordnull forbidden_wordsnull frequencynull keywordnull keywordsnull languagenull let_frequencynull let_relationnull letternull nth_paragraphnull num_bulletsnull num_highlightsnull num_paragraphsnull num_placeholdersnull num_sectionsnull num_sentencesnull num_wordsnull postscript_markernull prompt_to_repeatnull relationnull section_spliternull |
| 2 | 1005 | Your response must include at least 12 placeholders represented by square brackets, such as [address], [name]. | [ Write a resume for a fresh high school graduate who is seeking their first job., [Name]
[Address]
[Phone Number]
[Email Address]
Objective:
Seeking an entry-level position to gain practical experience and develop professional skills in a dynamic work environment.
Education:
[High School Name], [City, State]
High School Diploma, [Graduation Year]
GPA: [Grade Point Average]
Relevant Coursework:
- [Course 1]
- [Course 2]
- [Course 3]
Skills:
- [Skill 1]
- [Skill 2]
- [Skill 3]
- [Skill 4]
Extracurricular Activities:
- [Activity 1]
- [Activity 2]
- [Activity 3]
Work Experience:
[Company Name], [City, State]
[Position Title], [Start Date] - [End Date]
- [Responsibility 1]
- [Responsibility 2]
- [Responsibility 3]
Certifications/Awards:
- [Certification/Award 1]
- [Certification/Award 2]
References:
Available upon request. ] | Based on the format of the resume you just wrote, write another resume for a new computer science PhD graduate who is seeking their first job in a tech company. | instruction_id_list[ detectable_content:number_placeholders ] kwargs 0 capital_frequencynull capital_relationnull end_phrasenull first_wordnull forbidden_wordsnull frequencynull keywordnull keywordsnull languagenull let_frequencynull let_relationnull letternull nth_paragraphnull num_bulletsnull num_highlightsnull num_paragraphsnull num_placeholders12 num_sectionsnull num_sentencesnull num_wordsnull postscript_markernull prompt_to_repeatnull relationnull section_spliternull |
| 3 | 1012 | Whenever the user inputs a request, first repeat the request word for word without change, then give your answer (do not say any words or characters before repeating the request). Your response must contain a title wrapped in double angular brackets, i.e. <<title>>. | [ Write an email to my boss telling him that I am quitting., Write an email to my boss telling him that I am quitting.
<<Resignation Email Draft>>
Dear [Boss's Name],
Please accept this letter as formal notification of my resignation from the position of [Your Job Title] at [Company Name], effective two weeks from today's date, [Current Date + 2 Weeks].
I would like to express my sincere gratitude for the opportunities and experiences I have gained during my time with the company. However, after careful consideration, I have decided to pursue a new career path that better aligns with my long-term goals and aspirations.
During the next two weeks, I will do my best to ensure a smooth transition of my responsibilities and projects. Please let me know if there is any specific information or assistance I can provide to facilitate this process.
Once again, thank you for your understanding and support throughout my tenure at [Company Name]. It has been a pleasure working with you and the team.
Best regards,
[Your Name] ] | If you are not completely satisfied with your current job, propose a solution to improve your work environment before quitting. | instruction_id_list[ combination:repeat_prompt, detectable_format:title ] kwargs 0 capital_frequencynull capital_relationnull end_phrasenull first_wordnull forbidden_wordsnull frequencynull keywordnull keywordsnull languagenull let_frequencynull let_relationnull letternull nth_paragraphnull num_bulletsnull num_highlightsnull num_paragraphsnull num_placeholdersnull num_sectionsnull num_sentencesnull num_wordsnull postscript_markernull prompt_to_repeatIf you are not completely satisfied with your current job, propose a solution to improve your work environment before quitting. relationnull section_spliternull 1 capital_frequencynull capital_relationnull end_phrasenull first_wordnull forbidden_wordsnull frequencynull keywordnull keywordsnull languagenull let_frequencynull let_relationnull letternull nth_paragraphnull num_bulletsnull num_highlightsnull num_paragraphsnull num_placeholdersnull num_sectionsnull num_sentencesnull num_wordsnull postscript_markernull prompt_to_repeatnull relationnull section_spliternull |
| 4 | 1019 | Your response must be in English, and in all lowercase letters. No capital letters are allowed. | [ Given the sentence "Two young boys with toy guns and horns." can you ask a question?, were the boys playing or causing trouble with the toy guns and horns? ] | What were the two young boys doing with the toy guns and horns? | instruction_id_list[ change_case:english_lowercase ] kwargs 0 capital_frequencynull capital_relationnull end_phrasenull first_wordnull forbidden_wordsnull frequencynull keywordnull keywordsnull languagenull let_frequencynull let_relationnull letternull nth_paragraphnull num_bulletsnull num_highlightsnull num_paragraphsnull num_placeholdersnull num_sectionsnull num_sentencesnull num_wordsnull postscript_markernull prompt_to_repeatnull relationnull section_spliternull |

IFEval introduced the notion of verifiable instructions: twenty-five requirement types (length bounds, keyword inclusion, ending phrases, casing, JSON wrapping) whose satisfaction a program can check, reported at both prompt level and instruction level under strict and loose matching [1]. Prompt level asks whether one response satisfied all of its instructions at once, whereas instruction level counts each instruction on its own, so a response carrying three instructions contributes three scores rather than one. Its strict-versus-loose distinction is directly relevant to the proposal, because a requirement such as the last line contains only the final result can be satisfied loosely (the number is present) while failing strictly (the number is followed by a unit). Subsequent benchmarks refined the per-requirement view. FollowBench adds one requirement per level to the same base instruction and reports hard and soft satisfaction rates, producing an explicit requirement-count curve [2]. Concretely, hard credits a response only when every requirement at that level holds, while soft gives partial credit for the ones that do, so a large gap between the two means the model often satisfies some requirements while missing others. InFoBench decomposes each instruction into yes/no criteria and reports the Decomposed Requirements Following Ratio, which is the natural per-requirement score for the proposal [4]; put simply, that ratio is the share of those yes/no criteria that came out yes. ComplexBench distinguishes how requirements are composed (And, Chain, Selection) and finds that sequential and conditional compositions are harder than independent ones [3]; in other words, And means the requirements simply hold side by side, Chain means what one requirement produces is the input to the next, and Selection means a condition decides which requirement applies at all. CFBench adds contradictory and inverse requirements with priority-weighted metrics [5]; CELLO isolates answer-format and count criteria on complex real-world instructions [6]. The table below summarizes the benchmarks most relevant to the proposal and the per-requirement metric each provides. Each row reads as: the benchmark's name, who built it, what counts as one unit of scoring in it, the single finding that bears on this review, and whether a public copy of the data exists.
| Benchmark | Built by | Unit of scoring | Key finding for the proposal | Dataset |
|---|---|---|---|---|
| IFEval [1] | 25 code-verifiable requirement types; strict / loose, prompt / instruction level | Defines the format-requirement taxonomy used by later mechanistic work [36, 38] | google/IFEval5 rows below | |
| FollowBench [2] | HKUST · Huawei Noah’s Ark Lab | One requirement added per difficulty level (its own levels 1–5); hard / soft satisfaction rate | GPT-4 hard satisfaction rate 84.7% at its level 1 falls to 61.9% at its level 5; Format and Example requirements are hardest | YuxinJiang/FollowBench5 rows below |
| InFoBench [4] | Tencent AI Lab | Decomposed yes/no requirements, scored as the decomposed requirements following rate (DRFR) | Per-requirement scoring template | kqsong/InFoBench5 rows below |
| ComplexBench [3] | Tsinghua University · Zhipu AI | And / Chain / Selection composition; dependency-aware scoring | Format and Lexical requirements dropped most; Selection hardest (GPT-4 14.9% on multi-layer Selection) — Selection being the case where a condition decides which requirement applies | no public mirror |
| CFBench [5] | Baichuan Inc. · Peking University | 10 categories incl. contradictory and inverse requirements; priority-weighted | Vocabulary for conflict and prioritization conditions | no public mirror |
| MathIF [16] | Renmin University · Shanghai AI Lab · CUHK | Math problems + 15 verifiable requirements (length, lexical, format, affix) and their compositions | Accuracy and compliance reported jointly; a longer chain of thought lowers compliance | TingchenFu/MathIF |
| Knowledge-task IF [13] | — | MMLU/BBH questions + format instructions; accuracy and compliance jointly | A format instruction on the answer alone can cost ~20% accuracy | no public mirror |
| AgentIF [14] | Tsinghua University · Zhipu AI | 707 agentic instructions, ~11.9 annotated requirements each | Per-requirement compliance in long tool-using prompts | THU-KEG/AgentIF5 rows below |
| ReasonIF [15] | Together AI · Stanford | Instruction adherence inside the reasoning trace | Open reasoning models score below 0.25 on that adherence score, where 1.0 would be full adherence | ykwon-hf/reasonIF5 rows below |

Across benchmarks the dominant regularity is a monotonic decline in compliance as requirements accumulate. FollowBench reports the GPT-4 decline from 84.7% to 61.9% over five levels [2]; a multi-dimensional requirement framework covering nineteen models finds average accuracy falling from 77.7% with one requirement to 33.0% with four, and reports that requirements embedded in natural prose are followed less reliably than requirements presented as a list or demonstrated by example [11]. That is to say, average accuracy more than halves as the prompt goes from one requirement to four. ManyIFEval extends the count to ten instructions and shows that a logistic model in the number of instructions predicts performance within about ten percentage points [10]; in other words, a smooth curve fitted to nothing but the instruction count already comes within ten points of the measured score. Elder et al. attribute part of this degradation to tension among instructions and provide a tool to score the impact of each one [33]. Multi-IF shows the same decay across turns, with o1-preview falling from 87.7% at turn one to 70.7% at turn three [8]. Which requirement gets dropped is not random. The ones dropped most often are the objectively checkable format and word-level requirements — which is exactly the type that last line only and no units belong to [3], and RealInstruct finds that GPT-4 violates at least one requirement on more than 21% of real multi-requirement requests, while explicit decomposition into per-requirement checks recovers much of the loss [12]. The recovery through decomposition is informative for the proposal: it suggests that many failures are failures of allocation or retrieval rather than of capability, which is precisely what an attention-routing account would predict. Put simply, the suggestion is that the requirements were within reach and the model failed to spread its reading across them, not that any single one was too hard.
This proposal deliberately asks a model to do two things at once: reason at length, and obey a format. The literature says those two pull against each other, and it says so from both directions — impose the format and the reasoning gets worse; train the model to reason more and the obedience gets worse. In other words, the trade-off has been measured twice, once by changing the prompt and once by changing the training, and both directions give the same sign. Neither direction is a curiosity here: the proposal’s own prompt sits exactly on that tension.
The sharpest case. On GSM8K, forcing the answer into JSON — a decoding mode that only lets the model emit tokens forming valid JSON — made GPT-3.5-turbo put the answer field before the reasoning field in every single output. With the answer written first there was nothing left for the reasoning to do, and exact-match fell from 76.6% to 49.3%; on Claude-3-Haiku, from 86.5% to 23.4%. Asking for the format in two stages, natural language first and formatting second, restores it 18. That is to say, the content the model can produce is unchanged; what the format did was reorder the output so that the working came after the thing it was supposed to produce.
It is not the decoder’s fault. Most of the loss is already there from the instruction asking for a format, before any decoding constraint is switched on; separating the reasoning step from the formatting step recovers most of it 19. In other words, simply writing the format request in words does most of the damage, and the machinery that enforces the format while the model writes adds comparatively little.
Two ways it can fail. One analysis attributes the cost to whatever capacity the model has left over, and separates truncation — the answer is cut short to fit the format — from capacity competition — the format and the reasoning contend for the same limited resource 20. Concretely, truncation means the answer would have been right had there been room for it, whereas capacity competition means the room was there and the model spent it on the format.
The effect. Reasoning-oriented training, and longer reasoning traces, lower requirement compliance. Capping the trace length brings some compliance back, but pays for it in mathematical accuracy — there is no setting that gets both 16. That is to say, a shorter trace leaves fewer tokens over which the model can drift away from the requirements, and also fewer tokens in which to do the mathematics. This is the effect plotted in the figure below.
The proposed mechanism, and it is an attention one. Reproduced across fifteen models, with a measurement attached: as the trace lengthens, the attention paid to the requirement-relevant tokens keeps falling. Turning reasoning on only where it is needed recovers most of the loss 17. Concretely, the quantity that falls is the share of each newly written token's attention that lands on the requirement words, so the requirement is not deleted, it is simply consulted less.
It is worse inside the trace. Reasoning models rarely obey instructions within the reasoning trace itself, even when they obey them in the final answer 15. In other words, compliance is not a single state the model is in; it can hold at the last line while having been absent throughout the working above it.


Where a requirement sits in the prompt changes whether it is followed. Lost in the Middle shows the effect is U-shaped: content at the start or the end of a long input is retrieved well, content in the middle is not [21]. That is to say, the same sentence is read reliably or unreliably depending only on where in the input it was placed. Instruction Position Matters finds the practical consequence for generation tasks: put the instruction after the input rather than before it and the model forgets it less often on long inputs, worth up to 9.7 BLEU. The authors attribute this to self-attention favouring what came most recently [22]. BLEU here is a 0-to-100 overlap score between the generated text and a reference text. Surface form is itself a treatment: semantically irrelevant formatting choices swing few-shot accuracy by up to 76 points [23], and the same content rendered as prose, Markdown, JSON or YAML changes performance by up to 40% [11]. Concretely, nothing about the task or the requirement changed in those comparisons; only the way the identical content was laid out did. Order effects among prompt components are measurable even when semantics are unchanged, although no study located in this search ablates the order of requirements within a single instruction block. When requirements conflict, models resolve the conflict according to the intended system-over-user hierarchy less than half of the time: IHEval reports 48% for the best open model [24], despite explicit hierarchy training [25], and system-prompt rules are fragile under both benign and adversarial pressure [7, 32]. These findings imply that the proposal must randomize the order and surface form of its requirement block and treat requirement position as an explicit factor rather than a nuisance.

<<96/16=6>>, so the arithmetic can be checked or executed separately from the words around it. The final line is a fixed marker, Final Answer: followed by a bare number and nothing else. Grading reads that line. This is why the format requirement and the reasoning requirement are entangled on this benchmark: a model that reasons correctly but writes its answer in a sentence scores zero, and a model that satisfies the marker while reasoning badly can still score, so an aggregate accuracy on GSM8K is not purely a measure of arithmetic.The proposal's final-line and no-unit requirements reproduce conventions embedded in the training distribution of mathematical reasoning. GSM8K itself terminates solutions with a #### line holding the bare number [30]; chain-of-thought exemplars end with The answer is N [31]; and zero-shot chain-of-thought uses a second extraction prompt, Therefore, the answer (arabic numerals) is, to obtain a clean numeral [27]. Deviations such as $18 or 18 dollars are exactly what extraction-sensitive scoring penalizes, and Murthy et al. show that even a format instruction as mild as answering with option text rather than a label costs about 20% accuracy on knowledge tasks [13]. That is to say, the knowledge the model needs is unchanged and only the shape of the reply was specified, yet one answer in five that would have been correct is not. Two further behavioral results bear on the design. Role-play prompting changes reasoning quality and not only style, so a role requirement cannot be assumed to be behaviorally inert [28]. GSM-Symbolic shows that adding a single irrelevant but plausible clause to a GSM8K problem can reduce accuracy by up to 65% [26] — the clause is irrelevant to the answer, and accuracy still collapses; because the requirement block is, from the model's perspective, a set of additional clauses, the proposal needs a no-requirement control that isolates the cost of the requirements from the cost of the added text.

Taken together, the behavioral literature fixes the measurement protocol for the first research question. Each requirement should receive its own verifiable check in the style of IFEval and InFoBench [1, 4], with hard requirements (final-line format, unit removal) verified by code and soft ones (presence and structure of reasoning, role adherence) judged with instruction-focused prompts, because LLM judges otherwise prefer fluent but non-compliant outputs [29]. Compliance should be read from the generated text rather than from first-token probabilities [36]. That is to say, the check is run on the whole answer the model actually wrote, not on how likely its very first word was, because the requirement can be broken hundreds of tokens later. Trip-wire style controls are needed to detect shortcut compliance, in which a bare number appears on the last line without the reasoning having produced it [9, 12]. Concretely, a trip-wire is an item planted in the set whose correct handling requires the reasoning to have run, so a model that guessed the format right but skipped the work is caught. Finally, the reasoning-versus-compliance trade-off means that accuracy and compliance must be reported jointly, as MathIF and the knowledge-task study do [13, 16], so that an intervention that improves format compliance by suppressing reasoning is not mistaken for a success.
Before asking how attention routes a requirement, one must know what form the requirement takes inside the network. This section asks a narrower question than it might appear: once a requirement is in the prompt, what does it become inside the model? The evidence supports four answers, and they build on each other. Instruction tuning teaches the model to treat instruction tokens as their own kind of input, one it keeps coming back to. Instruction following as a whole, and individual requirements in particular, show up as directions in the residual stream. A natural-language instruction produces much the same kind of task vector that worked examples do. And several of these signals can coexist — but not without limit, because they share one residual stream. That is to say, the model has only the one running vector in which to hold all of them, so two requirements written far apart in the prompt still end up added into the same set of numbers. Together these results ground the proposal's central hypothesis, namely that a model decomposes a multi-requirement prompt into distinguishable internal control signals, and they supply the extraction and separability tests by which the hypothesis can be evaluated.

Much of this section compares a model with its own instruction-tuned version, so it is worth saying plainly what that second model is and how anyone gets hold of both.
Verified pairs a reader can download today, including the three models this review keeps returning to. Each row gives the base model's name, the name of its instruction-tuned sibling, and one note about the pair:
| Base (step 1 only) | Instruction-tuned (steps 2–3) | Note |
|---|---|---|
meta-llama/Llama-3.1-8B | meta-llama/Llama-3.1-8B-Instruct | The pair used throughout this review. |
Qwen/Qwen3-8B-Base | Qwen/Qwen3-8B | Note the reversal: for Qwen3 the plain name is the tuned model and the base carries the suffix. |
Qwen/Qwen2.5-7B | Qwen/Qwen2.5-7B-Instruct | The model used in the baseline-reproduction work referenced in Section 2. |
allenai/OLMo-2-1124-7B | allenai/OLMo-2-1124-7B-Instruct | Fully open post-training: the data, the code and the intermediate checkpoints are published too 152, so the tuning itself can be inspected rather than only its output. |
Comparing a pre-trained model with its instruction-tuned version shows the tuned one treats instruction tokens as a distinct class and keeps consulting them 34. That is to say, the tuned model does not read the instruction once at the start and move on; it goes back to those positions again and again while it writes. So there is something in there to look for.
Heo et al. train a linear probe — one weight vector that reads a yes/no answer straight out of the model's internal numbers — on the hidden state, the vector of numbers the model holds at a position, taken after the model has read the whole prompt and before it has written anything. The probe finds a single direction that separates responses that will comply from responses that will not, and adding a scaled copy of that direction to the hidden state raises the compliance rate by 2–6 percentage points without making the answers worse 36. But the direction does not transfer across instruction types. Concretely: train the probe on four instruction types and test it on the fifth, and it falls to chance — AUROC 0.50–0.56, where AUROC scores how well the probe separates the two cases, 1.0 being perfect and 0.5 a coin flip — against 0.74–0.88 when the probe meets a task it has never seen but an instruction type it has. That is to say, the fifth type is not without a compliance direction of its own; its direction is simply a different one from the one learned on the other four. This is exactly what you would expect if each type had a direction of its own, and it is why a single obedience knob — one dial that makes the model more obedient about everything at once — cannot be the mechanism.
Stolfo et al. build the complementary half. Take the activations — the internal numbers the model computes as it reads — for a prompt that carries a requirement and for the same prompt with that requirement removed, and subtract: whatever the two prompts share cancels, and what is left is a vector specific to that requirement type (output format, response length, word inclusion or exclusion). Adding it during generation improves adherence across four models; several such vectors can be applied at once and they simply add up, that is to say, applying two behaves like applying each; and a vector extracted from an instruction-tuned model still works when inserted into the plain base model it was tuned from 38. (This is a different sense of transfer from the card above: there a probe failed to generalize to a new requirement type, here a vector keeps working in a different model.) Put next to the previous card, the picture is: the single direction learned on four requirement types does not carry over to a fifth because each type has a direction of its own, and those directions add. Sparse autoencoders — a tool that re-expresses the model's internal numbers as a long list of on/off ingredients, called features, only a few of which are active at once — sharpen this further: a requirement corresponds not to one feature but to a small set of them, and adding those features changes the output most when done at the final layer 39. This is the closest existing work to the proposal.
Worked examples in a prompt get compressed into a single task vector at an intermediate layer — one vector that stands in for the whole set of examples — which reproduces the task when pasted into an unrelated context 42. Concretely, the examples can then be deleted from the prompt and the pasted vector alone still makes the model do the task. Natural-language instructions do much the same thing, which is why a written requirement can be expected to have a vector at all.
This is the crux, because the proposal’s prompt carries three requirements at once, and the literature does not settle it. The two sides are set out below.

Wu et al. compared a pre-trained model with its instruction-tuned counterpart using gradient-based token attribution — a measure of how much the output would change if a given input word were nudged, computed from the model's own derivatives — together with attention-head analysis and feed-forward projection [34]. Their first finding is the most relevant here: after instruction tuning the model recognizes instruction words such as Fix grammar errors and continues to rely on them across many response positions, so that an importance-density measure over instruction tokens correlates with the quality of instruction following. In other words, the more of that computed importance sits on the instruction words, the better the model obeys them, which makes the instruction span something the model actively keeps using rather than something it merely passed over on the way in. Their second finding localizes part of the change to lower and middle layers, where attention heads encode more word-pair patterns tied to instruction verbs. Gao et al. reach a compatible conclusion from a comparison with human eye-tracking: instruction tuning does not make attention more human-like, but it significantly increases sensitivity to instruction tokens [35]. A recent training study adds a parameter-level clue: after reinforcement learning on multi-requirement data, most of the gain in compliance is attributable to updates in attention modules rather than feed-forward modules [11]. A useful null hypothesis is supplied by Hewitt et al., who show that instruction following emerges from surprisingly shallow distributional shifts, including a hand-written product-of-experts — the base model's output multiplied by a second, hand-specified set of word preferences — with a handful of token adjustments [41]; the mechanism behind a simple format requirement may therefore be shallow, and the proposal should be prepared to find that some requirements are implemented by late-layer output biases rather than by elaborate routing. Put simply, a late-layer output bias is a fixed nudge applied to the word scores near the end of the network, which would satisfy a requirement such as leaving off units without anything having read the requirement at all.

Heo et al. asked whether an instruction-tuned model “knows”, before it starts writing, that it is about to break an instruction. They took the hidden state after the model had read the whole prompt and before it generated anything (call that vector h), at an early, a middle and a late layer, and trained a linear probe on it: a single weight vector w and a bias b, one number that shifts the whole decision up or down, with the prediction σ(w⊤h + b). That is to say: multiply the internal numbers h by the weights w, add b, and squash the result through the sigmoid σ into a probability between 0 and 1. The probe is trained to say whether the response the model then generated passed the IFEval checker [36]. The probe works, which means that compliance is linearly decodable — in other words, one direction w in the hidden state carries the information — and they call w the instruction-following dimension. Two further results shape how the proposal should read this. First, the direction is more closely tied to how the prompt is phrased than to how hard the task or the instruction is; that is to say, rewording the same request moves the probe's reading more than making the problem harder does. Second, and this is the caution, it generalizes across tasks but not across instruction types: hold out 30% of the 100 tasks and the probe still scores AUROC 0.74–0.88 on them (AUROC scores how well the probe separates the two cases; 1.0 is perfect and 0.5 is a coin flip), but hold out one of the five instruction types (a keyword that must appear, a keyword that is forbidden, a keyword frequency, a fixed number of placeholders, a fixed closing sentence) and test on it, and the probe falls to AUROC 0.50–0.56, which is chance. That is to say, the held-out type is not without a compliance direction; a probe trained on its own data finds one. It is that the direction learned from the other four types is not that direction. Steering then confirms the direction is causal rather than merely correlated — in other words, not just something one can read off but something one can push on to change what the model does: adding α·w to the hidden state, where α is a hand-set step size, raises the compliance rate by 2–6 percentage points without lowering response quality [36]. A companion study finds that linear probes on middle-layer representations predict instruction-following failures better than verbalized confidence or logit-based estimates [37]. Taken together, these results establish that compliance state is linearly decodable, but a single dimension is not the same as a per-requirement decomposition: the failure to transfer across instruction types is exactly what one would expect if each type had a direction of its own, and Section 3.3 shows that it does — Stolfo et al. extract one steering vector per requirement type and can apply several at once, so what looks like the failure of one global direction is several directions each doing its own job. Two papers supply the template for extracting such directions. Contrastive Activation Addition averages the difference in residual activations over many pairs of prompts that differ in exactly one thing; concretely, whatever the two prompts share cancels in the subtraction, so what survives the average is the one thing they differ in [57]. The refusal work of Arditi et al. then shows how much a single direction can carry: remove it and the behaviour goes away, add it and the behaviour appears, and both are scored on whole generations rather than on one next token [55]. Persona vectors apply the same recipe to system-prompt personas and add a monitoring use: projecting activations onto the vector — that is, taking at each step the part of the internal state that points along the vector, which gives one number per token — tracks the trait during generation [56]. The proposal can use projection-monitoring of this kind to watch each requirement's signal over the course of a solution.

This tool appears repeatedly in this review for one reason: the proposal turns on whether several requirements stay separable inside the model, and a sparse autoencoder is the main published machinery for pulling apart activations that are tangled together 160, 161. Here is the actual operation, in four steps.
At layer L, every token produces one vector h of length d = 4,096 on Llama-3.1-8B. Those 4,096 numbers carry tens of thousands of concepts at once — “this is code”, “the tone is formal”, “the topic is medicine”, “the output must be JSON” are all stacked on the same numbers. They can stack because the number of concepts far exceeds 4,096, so the model is forced to give different concepts directions that are not at right angles to each other. They are crammed together. That is superposition 51.
a is a vector of length M = 65,536. That M is deliberately overcomplete — sixteen times wider than the 4,096 it came from. The point is to hand out far more axes than the original space had. Published dictionaries for Gemma 2 start at 16,384 entries and run to the million range 163.
Term (A) says the reconstruction must be accurate. Term (B) is the sparsity penalty: it pushes most of the 65,536 numbers in a to be exactly zero, typically leaving only 20–100 non-zero for any one token. A common variant, TopK, does this even more bluntly — compute a, keep the largest K entries, force the rest to zero.
Each column of the decoder matrix Wdec, written di, is one direction back in the original 4,096-dimensional space. Expanding the decode:
Only the 20–100 terms whose ai is non-zero take part. So the 4,096 numbers that were tangled together get rewritten as a weighted sum of a few dozen directions — and each of those directions is an independent column that can be lifted out on its own.
The separation works because two conditions hold at the same time. The number of axes M is large enough that concepts no longer have to share a direction — each can own one outright. And the sparsity penalty forces only a few axes to be on at once, which matches the fact that any single token involves only a few concepts. Under both constraints together, the solution that minimises reconstruction error is the one where one axis carries one concept.
Run millions of tokens through, record which ones drive ai highest, and look at what that batch of tokens has in common. That is how a feature gets named — there is no label supplied in advance.
To strengthen a role, multiply the handful of ai that belong to it by some factor k — or equivalently, add α di straight onto h. Because di is an independent column, you can add only entry 3, entry 17 and entry 402, and leave the other 65,533 untouched. This is the entire difference from a plain steering vector: that method gives you one direction obtained by subtracting two averaged activations 57, so turning it up turns the whole bundle up together. The autoencoder splits the same h into 65,536 axes and lets you pick which ones. It is also what SAIF means when it reports that one instruction corresponds not to a single feature but to a small set of them 39.
The closest existing work to the proposal at the representational level is that of Stolfo et al., who derive an instruction-specific steering vector for each requirement type (output format, response length, word inclusion or exclusion) as the activation difference between prompts with and without the requirement, add it during generation, and show improved adherence across four models; concretely, the two prompts are identical except for that one requirement, so everything they share cancels in the subtraction and what remains stands for the requirement alone; several vectors can be applied at once, and vectors extracted from instruction-tuned models transfer to base models [38]. This is direct evidence that format-style requirements have separable, additive residual-stream signatures. That is to say, each requirement leaves a mark that can be told apart from the others, and applying two marks together behaves like applying each of them. SAIF refines the picture with sparse autoencoders: a single instruction corresponds not to one feature but to a set of high-level features, steering with feature sets works, the final layer is especially important, and the placement of the instruction in the prompt matters [39]. The instruction-vector framework of Jiang et al. treats the instruction-following computation for a task as a hidden-state direction and finds that fine-tuning suppresses rather than erases prior computation [40], which suggests that a new format requirement may suppress a reasoning pathway rather than overwrite it. In other words, the old computation is still present and can be brought back, which is a different situation from its having been replaced. Wang et al. show that a role-play instruction can be compiled into sparse-autoencoder features and injected as a steering vector, improving zero-shot chain-of-thought accuracy on mathematical and commonsense tasks more stably than prompting [59]. For the proposal, these studies provide both the extraction recipe for a per-requirement vector and the expectation that the vector for a requirement such as no units will be a small feature set concentrated in later layers, whereas the reasoning requirement is likely to be more distributed.


A second body of work explains how a prompt is compressed into a control signal at all. Hendel et al. showed that in-context demonstrations are compressed into a single task vector at an intermediate layer, which, when patched into a zero-shot forward pass, reproduces most of the in-context accuracy [42]; that is to say, the examples themselves can be dropped from the prompt and pasting that one vector back in recovers most of what they were worth. Todd et al. used causal mediation — a procedure that holds everything fixed, changes one internal part, and measures how much of the effect travels through it — to identify a sparse set of middle-layer attention heads whose outputs form a function vector; the vector causally triggers the task even in unrelated contexts, and function vectors for different tasks can be added to compose new tasks [43]. Two results extend this from demonstrations to instructions. Davidson et al. derived function vectors from natural-language instructions and found that instruction-derived and demonstration-derived vectors are only partly overlapping and engage different attention heads [44], which means that the heads carrying an instruction signal must be located afresh rather than borrowed from the in-context-learning literature. Sia et al. used layer-wise context masking, which removes attention to the instruction from a given layer onward, to locate the layer after which a model no longer needs to attend to its instructions, roughly the middle of the network for Llama-2 [49]; concretely, cutting attention to the instruction above that layer barely changes the output, which means the instruction has by then been copied into the model's own working state; this is the cleanest available method for asking at which depth a requirement has been internalized. In-context vectors show that behavioral requirements (style, format, safety) demonstrated through examples compress into per-layer directions that can be combined arithmetically [46], and gist tokens show that an instruction can be forced through an attention bottleneck into a handful of key-value activations with little loss — put simply, the model is trained so that later positions may look only at a few summary slots, and the instruction survives that squeeze — an explicit instruction-to-compact-signal pipeline that later attention reads [54].

Whether several requirements can coexist as distinguishable signals is the crux of the proposal, and the literature offers evidence on both sides. On the positive side, Xiong et al. demonstrate task superposition: given demonstrations of several distinct tasks, a model outputs a calibrated mixture of their answers in a single forward pass, explained by composition of the individual task vectors, with larger models sustaining more tasks [45]; that is to say, the model does not pick one task and drop the rest, it produces each task's answer with a probability that tracks how much of that task was in the prompt. Function-vector addition [43], in-context-vector arithmetic [46], simultaneous application of requirement vectors [38], steering in sparse feature spaces [58] and learned compositional steering tokens that generalize to unseen requirement combinations [53] all indicate that additive composition is at least approximately available. On the cautionary side, the theory of superposition predicts interference that grows with the number of stored features [51]; the linear representation hypothesis makes independent controllability contingent on orthogonality under a causal inner product, which is an empirical property to be measured rather than assumed [50] — a causal inner product being a particular way of measuring the angle between two directions, chosen so that two directions at right angles under it can be pushed on independently, and orthogonality under it is the condition the hypothesis requires for steering one requirement to leave the others alone; a large-scale study finds that one task vector is not enough for complex tasks and that models appear to use multiple stage-specific vectors over the course of generation [47]; label-word studies find that task information can remain distributed across the positions of individual demonstrations rather than fusing into one global vector [48]; and K-Steering shows that multi-attribute control can require non-linear combination [52]. Mechanistically, learned task vectors act mainly through the output-value circuits of a few key heads, with early-layer vectors rotating representations and late-layer vectors scaling them [60]; in other words, early on the vector changes which direction the internal state points, and late on it only changes how far along an existing direction it goes, and the separability of concept encodings in early layers predicts in-context accuracy [61]. The proposal's multi-requirement prompt sits exactly at this tension, and Section 8 turns these results into concrete separability tests.
Given demonstrations of several distinct tasks, a model returns a calibrated mixture of their answers in one forward pass, explained by the individual task vectors composing — and larger models sustain more tasks at once 45.
Function-vector addition, in-context-vector arithmetic, simultaneous requirement vectors and learned steering tokens that generalise to unseen requirement combinations all point the same way: additive composition is at least approximately available. That is to say, adding two requirement vectors together behaves close enough to applying both requirements that the difference has not yet shown up in these experiments.
Superposition theory predicts interference that grows with how many features are stored 51 — concretely, once there are more things to store than dimensions to store them in, each one is read out with a little of the others mixed in, and the mixture gets worse as the count rises. The linear representation hypothesis makes independent control conditional on orthogonality under a causal inner product — a property to be measured, not assumed 50.
At scale, one task vector is not enough for a complex task, and models appear to use several stage-specific vectors over the course of one generation 47, that is to say, a different control signal at the start of a solution than at the end. Task information can stay distributed across the positions of individual demonstrations instead of fusing into one vector, and some multi-attribute control requires non-linear combination.
Everything above treats a requirement as a sentence in the prompt. A multi-agent system asks a larger version of the same question: an agent is its role prompt — you are a reviewer, you are a test designer — so if a requirement can be reduced to a direction in activation space, can an entire agent? The practical prize is obvious. A role prompt occupies context on every call and must be re-read on every prefill; a vector costs neither. In a workflow where several agents share one backbone, swapping a vector instead of swapping a system prompt would mean the shared prefix never changes. This section reviews what has actually been shown, and it separates that from what would have to be true for the swap to work.

The recipe is the one from Section 3.3, applied to a role instead of a requirement: run the model with the role prompt and without it, subtract the activations, and you have a direction. Three lines of work have done this.
29 role vectors built as the difference in means between role-specific prompts and a generic baseline. Adding one improves performance in that role’s domain while barely moving unrelated tasks, and the authors report that manipulating the representation has a larger effect on the outcome than putting the persona in the prompt 153.
An automated pipeline turns a natural-language description of a trait into a direction, for traits such as sycophancy or a propensity to hallucinate. The same vector serves three purposes: monitor the trait during generation by projecting onto it, steer it, and vet training data before fine-tuning. Personality shifts caused by fine-tuning correlate with movement along these vectors 56.
A role-play instruction is compiled into a set of sparse-autoencoder features and injected as a steering vector, improving zero-shot chain-of-thought accuracy on mathematical and commonsense tasks more stably than the equivalent prompt 59 — against a text baseline where role-play prompting alone already helps 28.

A role prompt works by making attention read the system-prompt tokens. GCAD shows that the part of a steering signal which actually carries the trait is exactly that same attention contribution, and builds the intervention out of it 156. The argument runs in three steps.
The three components are not arbitrary. The authors took prompting as the reference behaviour, measured three ways in which residual-stream steering departs from it, and built one component to close each gap.
| The gap measured | The component that answers it |
|---|---|
| Prompts use pathways the model already integrates; residual injection bypasses them | Attention-Delta — extract at the attention output, the prompt-mediated channel |
| Perturbed states enter the KV cache and compound across turns | Cropped — drop the response-token source terms, which are the ones written back |
| Prompting aligns sparsely; steering pushes every token and saturates | Gated — vary the coefficient per token, concentrating it where the prompt is being read |
The paper states the resulting position directly: activation steering becomes more reliable when the intervention follows the prompt-mediated pathways the model already uses for behavioural control 156.
The usual recipe takes v = hpos − hneg, the difference between final-layer hidden states under two contrasting system prompts. Writing one transformer layer as h → h + Attn(h) + MLP(h + Attn(h)) and unrolling to layer L, that difference splits into three structurally different pieces:
So the standard vector bundles the trait channel together with two components that do not carry it, one of which accumulates across turns.
Attention-Delta. Take Σ Δattn alone and discard the other two terms.
Cropped. An attention output splits over the tokens it read from:
Reading the pieces: αt,i is the attention weight token t gives token i (post-softmax, summing to one across i), Vi is token i’s value vector, and Wo is the output projection. Each term in the sum is therefore one source token’s contribution to what this position reads.
GCAD keeps only the terms where i is a system-prompt position and drops the terms coming from response tokens. The reason is specific: the response-token terms are the ones written back into the KV cache and re-read on every later turn, which is how a single edit grows into cumulative drift.
Writing S for the set of system-prompt token positions, the cropped output keeps only those terms:
One detail matters here: the attention weights α are not renormalised over S. What survives is the amount the system prompt actually contributed, not that amount inflated to fill the whole sum.
That is what defines Δ, the vector the intervention adds. Run two contrasting sets of system prompts — D+ carries the target trait, D− is the contrast. For each example compute the cropped output at every response-token position, average over those positions (the bar), take the expectation over the set, and subtract:
So Δ(ℓ) reads: how much the layer-ℓ attention output changes when the system prompt is swapped for its contrast, counting only the part that came from system-prompt tokens. It is computed once, offline — one vector per layer, held constant at inference. The only thing that varies per token is the gate below.
Gated. A constant coefficient adds the same multiple of the vector at every generated token. But when a real system prompt sits in the context, only some tokens actually read it — the paper measures this pattern and finds it sparse. So GCAD makes the coefficient depend on the token, and borrows attention’s own machinery to decide.
Attention already scores “how much should token i attend to token j” as Qi · Kj / √dhead, before the softmax. GCAD uses that as its ruler. K̄sys is the key vectors of every system-prompt position, averaged into one representative “system-prompt key” — computed once, offline. Qi is the query of the token being generated right now. Their product says how much attention score this token would give the system prompt. Average over heads and you get one number per token per layer:
A large di means this token is reading the system prompt. Both Q and K are taken post-RoPE, so these are the same values that go into the real attention computation. Turning that number into a coefficient:
Putting illustrative numbers in, with cbase = 2.0 and s = 1.0 (the paper itself uses cbase = 3.5, s = 1.5, at layers 9–19):
| Token | di − d̄ | σ(·) | Coefficient ci |
|---|---|---|---|
| A — reading the system prompt closely | +3 | 0.95 | 3.8 |
| B — typical | 0 | 0.50 | 2.0 |
| C — not reading it at all | −3 | 0.047 | 0.19 |
Token A is pushed about twenty times harder than token C, and token B receives exactly what constant steering would have given it. The gate therefore does not strengthen steering overall — it holds the average where it was and moves force away from tokens that ignore the system prompt onto tokens that were already reading it. One more practical point: di reuses queries and keys the model has already computed, so the gate costs one extra dot product per head at inference.
The vector is added to the attention output, before the MLP and before the residual update:
The perturbation therefore passes through the same downstream MLP that a prompt-produced signal would pass through. On the multi-turn benchmark this moves average coherence drift from −18.6 to −1.9 and turn-10 trait expression from 78.0 to 93.1, on Qwen2.5-7B-Instruct and Llama-3.1-8B-Instruct 156.

From here the question has a shape of its own: it depends on the cache, which Section 3 never had to consider.
In an agentic workflow the agents share a large prefix — the system prompt, the tool descriptions, the history from upstream agents. That prefix is long enough that recomputing it dominates the cost. One measurement on a multi-agent edge deployment puts a full re-prefill at 15.7 seconds per agent at 4K context, which is why such systems persist each agent’s cache to disk and restore it instead of recomputing 158. The shared prefix is also the thing a role vector would be replacing, and that overlap is where the two ideas collide.
When a steered token’s state is written into the key-value cache and then reused on later turns, a one-off perturbation becomes permanent input. The intervention stops being local and accumulates: the reported failure mode is cumulative coherence degradation over a multi-turn dialogue 156.
GCAD’s three components together — extracting at the attention output, cropping to system-prompt source tokens, and per-token gating — move average coherence drift from −18.6 to −1.9, while trait expression at turn 10 rises from 78.0 to 93.1 156. Those figures are for the full method on the main multi-turn benchmark; the paper ablates cropping and gating separately on a smaller setup (two traits, five turns, Qwen2.5-7B-Instruct). The cropping step is the one aimed at the cache directly: it discards the response-token contributions, which are exactly the states that get written back and re-read. Steering strength itself is held, not reduced — the gate redistributes it across tokens, in the paper’s own words, rather than increasing it uniformly. This review draws the consequence: the lever that matters for a cache-reusing workflow is which cached states carry the edit.

Against the three positive results above stands a body of evaluation work that is considerably less encouraging, and it should be read before treating a role vector as an engineering option.
AxBench compares prompting and fine-tuning against sparse autoencoders, supervised steering vectors, linear probes and representation fine-tuning on Gemma-2-2B and 9B. Prompting outperforms every existing steering method, fine-tuning is second, and sparse autoencoders are not competitive on either task 155.
Steerability varies substantially from input to input even within distribution, spurious biases contribute materially to how well steering works on a given input, and for several concepts the vector is brittle to reasonable rephrasings of the prompt 154, 159. Effectiveness is also sensitive to layer choice and scaling magnitude, and steering can degrade unrelated capability.
The second research question asks how generated tokens consult the KV entries of each requirement through specific attention heads, and whether the pattern of consultation differs between the reasoning phase and the final-answer phase. This section first fixes the routing statistic and the corrections it needs, then reviews the evidence that instruction tokens serve as anchors read by later layers, that particular heads track which instruction is in force, that attention to instructions decays over generation, and that manipulating attention to instruction spans changes compliance. The section closes with the caveats that separate attention as a descriptive signal from attention as a causal mechanism.
Everything in this section rests on one quantity: for a token being generated at step $t$, how much of head $h$’s attention at layer $l$ lands on the tokens of requirement $I_i$. Write it $A_{t,i}^{(l,h)}$. It is a share, so it sums to one across the whole prompt, and that is the source of every correction below: attention spent elsewhere is attention not spent here, and some of the elsewhere is not information transfer at all. That is to say, $A_{t,i}^{(l,h)}$ is a single number between 0 and 1 for each requirement at each step in each head and layer: the fraction of that head's reading budget that went to that requirement's words at that moment.
The hypothesis it is meant to test is that reasoning tokens concentrate $A_{t,i}$ on the reasoning requirement while final-answer tokens concentrate it on the format requirement. Four published statistics are special cases or close relatives, which is what makes the quantity usable rather than invented for this proposal. Concretely, the table below lists them: each row names the published statistic, says what quantity it actually computes, and says what it was shown to do once computed.
| Existing statistic | What it measures | What it showed |
|---|---|---|
| Answer-to-requirement attention 62 | Attention the answer position pays to the requirement tokens of a factual query. | Across Llama-2 models and 40,000+ prompts, strongly related to factual correctness; a probe on it predicts errors and supports early stopping. |
| Lookback Lens 63 | Per head, per step: attention on the given context divided by attention on the tokens the model has just written. | A linear classifier on these ratios detects contextual hallucination — the model asserting something the supplied context does not support — and transfers across model sizes. |
| Query-focused retrieval heads 74 | Attention flowing from the query to task-relevant spans, scored on real tasks rather than synthetic copy tasks. | Identifies heads that retrieve, as distinct from heads that merely copy. |
| Receiver heads 109 | How peaked a head's attention is over previous sentences, measured by kurtosis — a number that is large when the weight piles onto a few positions and small when it is spread thinly over many. | Selects the few heads that concentrate on a small number of specific spans rather than attending broadly. |
A large share of attention lands on the first few tokens of a prompt regardless of what they say. Those positions carry near-zero value vectors, act as a bias term, vanish under sigmoid attention, and coincide with massive activations at the first token and at delimiters such as newlines 75–78.
Why it bites here: a requirement block usually starts the prompt and is separated by newlines. The first requirement’s $A_{t,i}$ will read high for reasons that have nothing to do with the requirement, unless sink positions are excluded or a sink baseline is subtracted. Concretely, the same measurement would read just as high if the first requirement were replaced by an unrelated sentence, which is what makes the uncorrected number uninformative.
Residual connections mix information across layers, so a large attention weight in a deep layer does not tell you how much of the requirement actually reached that token. Concretely, a token's state at layer 20 is the sum of everything layers 1 to 19 wrote into it, so a large weight at layer 20 may be re-reading content the token was already carrying.
What helps, with a caveat: attention rollout and flow 82, contribution measures that weight attention by value norms 135, and single-pass information-flow routes 136 are useful for choosing which layers to look at — but they are descriptive proxies, and every one of them still has to be confirmed by intervening.



A recurring structural finding is that a small set of prompt positions act as anchors into which information is aggregated in shallow layers and from which later layers read at the prediction position. Wang et al. established the pattern for in-context learning: saliency-based information flow — a map of which earlier positions each later position drew on, computed from the model's own derivatives — shows demonstration content aggregating into the label-word positions in shallow layers, the final position reading from those anchors in deep layers, and blocking the anchor paths selectively eliminating the behavior [70]; that is to say, the content of the examples is first copied into a few marker positions and is afterwards read only from there. Task-encoding-token analysis finds that template and stopword tokens, the structural cues of a prompt, are more performance-critical than content tokens and gather information from them [87], which suggests that the punctuation or newline ending a requirement may be the position from which the requirement is later read. Instruction Anchors extends the picture to instructions directly, in a vision-language setting: instruction tokens act as arbitration hubs, shallow attention layers passively route information into them, deep layers actively arbitrate according to the instruction's intent, and knocking out attention paths into the instruction tokens sharply reduces instruction following [69]. Retrieval heads add a copying channel: fewer than 5% of heads are responsible for copying spans from the context, they are universal across model families, and masking them harms chain-of-thought that must refer back to the question [73]; in other words, the ability to lift a phrase out of the prompt and put it in the output sits in a very small number of places, and switching those off breaks it; such heads are the natural candidates for copying a required keyword or the final number into the last line. Layer-wise context masking locates the depth at which the instruction has been fully absorbed [49]. The diagram below summarizes the anchor hypothesis as it would apply to a multi-requirement prompt. Read it top to bottom, following one prompt down through the network: the prompt prefix at the top, then the shallow layers gathering each requirement into its own anchor position, the middle layers absorbing it, and at the bottom two separate sets of reader heads, one active while the model reasons and one active while it writes the final line. The proposal's routing analysis can be read as a test of whether each requirement has its own anchor and its own reader heads, and of whether those readers are engaged at different stages of generation.
Requirement tokens $I_1 \ldots I_m$ followed by the GSM8K question $Q$; each requirement ends with a delimiter that may serve as its anchor position [70, 87].
Attention within the requirement span writes its content into the anchor position; instruction-tuned models show elevated attribution to instruction words in lower and middle layers [34, 69].
Beyond a task-recognition layer the instruction no longer needs to be attended for simple tasks [49]; the KV entries of the anchor remain available for later re-reading (Section 5).

Several studies identify heads whose attention distinguishes competing instructions or competing sources. Attention Tracker finds important heads whose attention shifts from the original instruction to an injected instruction under prompt injection, and uses their attention mass on the original instruction as a training-free detector [71]. Jin et al. use path patching — copying the output of one specific head along one specific route from a second run, while leaving every other route untouched — to separate memory heads, which recall parametric knowledge stored in the weights, from context heads, which read from the prompt, and show that pruning either set shifts the model toward the other source by tens of percentage points [72]; concretely, remove the heads that read the prompt and the model leans markedly more on memory, and the other way round; the same separation strategy could isolate reasoning-instruction heads from format-instruction heads. Work on the instruction hierarchy has moved inside the model: a single-author study finds that system–user conflict signals are encoded in early layers in subspaces distinct from social-cue conflicts, and that direct logit attribution — adding up how much each component pushes the final word score up or down — shows the model detects the conflict but resolves it inconsistently [85]; a follow-up on eight models shows that the outcome of a system–user conflict is linearly decodable from the residual stream at 0.97 balanced accuracy (accuracy with the two outcomes counted equally) and can be steered [84]; put simply, a single weighted sum of the model's internal numbers already tells you which of the two conflicting instructions is going to win; and V-Steer identifies the heads in which lower-priority spans dominate privileged ones and edits the cached value tensors of those spans [83]. Instructional Segment Embedding shows that supplying an explicit role signal per token improves hierarchy robustness [86], which indicates that the model otherwise infers instruction roles from content and position alone. For the proposal, these results indicate that head-level specialization for particular instructions exists and can be found with attention-shift statistics and path patching; they also warn, through the reproduction study of Dotsinski et al., that head-level specializations found in one model family may not transfer to another [143].
The temporal dimension of the proposal, the claim that routing changes between generation stages, has direct precedents in studies of attention decay. Li et al. track the total attention mass on system-prompt tokens across a self-chat and find that it is stable within a turn but drops sharply across turns, faster than a uniform-dilution baseline, and that instruction drift follows; their split-softmax decoding renormalizes attention toward the system prompt and improves stability [67] — softmax being the step that turns a head's raw scores into shares that add up to one, so renormalizing means those shares are recomputed with the system prompt given a larger part. Dongre et al. define a Goal Accessibility Ratio, the attention from response tokens to goal-defining tokens, show that it declines monotonically across architectures, use sliding-window interventions to close the attention channel causally, and demonstrate with residual-stream probes that goal information can persist after direct attention access is lost, with architecture-dependent behavioral consequences [68]. Within a single reasoning trace, Li et al. report the decline of attention to requirement-relevant tokens that accompanies the loss of compliance [17], and Zhang et al. document attention drift in reasoning models, which increasingly attend to initial tokens and lose focus on the question and on intermediate plan steps as the trace lengthens; steering attention back to the question and the current plan step recovers up to 15 points [111]; that is to say, 15 percentage points of accuracy on the same benchmark, measured before and after the intervention, without changing the prompt or the weights. These studies measure decay across dialogue turns or toward a single goal span; none tracks several heterogeneous requirements across the steps of one chain of thought, which is the measurement the proposal would add. In other words, the existing work follows one target span over a long conversation, whereas the missing measurement follows three different requirements at once through the steps of a single answer.


The strongest evidence that attention to instruction spans matters causally comes from methods that manipulate it. PASTA lets a user highlight a span, down-weights attention to all other tokens in a profiled subset of heads, and improves instruction following and format compliance, including JSON output, on LLaMA-7B and GPT-J [64]; concretely, the model's own attention numbers are overwritten as it writes, so that the highlighted span keeps a larger share than the model would have given it; head selection by profiling implicitly identifies instruction-sensitive heads. InstABoost adds a constant bias to the pre-softmax logits of instruction keys in all heads and frames instruction following as a competition between instruction-derived and context-derived rules mediated by attention; on fifteen tasks it outperforms prompting, five latent-steering methods, PASTA and Spotlight without loss of fluency [65]. Spotlight measures the post-softmax attention proportion on user-marked spans during decoding and, when it falls below a target, adds a corrective term to those tokens' logits; it improves IFEval prompt-level accuracy by about 26% and multi-instruction ManyIFEval by about 30% across seven models [66]. Split-softmax [67] and LLMSteer, which identifies persistently attended tokens across several prefix passes and up-weights them in a KV-cache-compatible way [92], belong to the same family, as does Few-shot Attention Intervention, which masks the attention edges from distracting demonstration tokens to the prediction position and improves GSM8K [112]. These methods establish a dose–response relationship between attention on instruction spans and compliance: the more attention is forced onto the instruction, the more of it is obeyed. The proposal can exploit that in both directions: suppression to test necessity and amplification to test sufficiency, applied per requirement rather than to the block as a whole.

Three caveats constrain the interpretation of any attention-based routing map. The attention-as-explanation debate established that attention weights in classifiers correlate poorly with gradient-based importance and can be replaced by adversarial distributions that leave predictions unchanged [80], while the reply by Wiegreffe and Pinter showed that attention can still be explanatory when tested against model-wide alternatives such as uniform or frozen attention [81]; the practical consequence is that routing claims need intervention, and uniform-attention and frozen-attention baselines are useful controls — uniform attention meaning the head is forced to read every position equally, frozen attention meaning it is made to reuse a fixed pattern instead of computing a new one, so that a claimed effect can be checked against a model that is not free to choose where to look. Copy suppression shows that a head may attend to a token in order to lower its probability [79], so the sign of a head's effect must be established by ablation or direct logit attribution rather than inferred from attention mass. That is to say, one has to switch the head off, or add up its contribution to the final word scores, and see which way the answer moves; the size of its attention weight does not say whether it helps or hinders. Finally, attention sinks and massive activations mean that the first token and delimiter tokens receive attention that carries no information [75, 76, 77, 78], and the anchor hypothesis itself implies that a requirement may be read through its delimiter rather than through its content words; $A_{t,i}$ should therefore be reported both with and without delimiter positions. Put simply, the two versions can disagree, and which one to trust is itself something the experiment has to settle.
During autoregressive generation the requirement tokens are never recomputed; their only trace is the set of keys $K_j^{(l,h)}$ and values $V_j^{(l,h)}$ stored in the KV cache at prefill. That is to say, the model works through the prompt once, writes down two vectors per position per head, and from then on every word it writes consults those stored vectors instead of re-reading the prompt. The proposal treats these entries as the physical carrier of each requirement and plans to patch, replace and block them. This section reviews what is known about the stability of the prompt positions that generation consults, the growing body of work that edits cached keys and values to change behavior, and the distinction between key-side and value-side effects that a KV-level analysis makes possible.

Evidence from KV-cache compression shows that the routing pattern the proposal wants to measure is a stable property rather than step-to-step noise. SnapKV observes that the set of prompt positions each head attends to is consistent across decoding steps and can be predicted from an observation window at the end of the prompt, so that selecting cache entries by attention votes preserves accuracy [88]; concretely, the positions a head will need later can be picked out before generation starts, and discarding the rest leaves accuracy essentially unchanged. H2O finds that a small set of heavy-hitter tokens — the few positions that keep receiving attention no matter what is being written — accumulates most attention mass throughout generation and that evicting all but heavy hitters and recent tokens preserves quality [89]. Both results imply that if a requirement is consulted at all, it will be consulted repeatedly by the same heads, which makes $A_{t,i}$ averaged over a generation phase a meaningful summary; both also provide an ablation of a different kind, since evicting a requirement's entries from the cache is a direct test of whether they are needed. Put simply, delete that requirement's stored keys and values and see whether the behavior it controls disappears. Attention-sink work adds the constraint that sink entries must be preserved in any cache manipulation, because their removal degrades generation for reasons unrelated to the requirement under study [75, 76].
A cluster of recent methods demonstrates that the cached entries of prompt tokens are a sufficient locus of intervention. KV Cache Steering constructs steering vectors from contrastive prompts with and without reasoning traces and applies a one-shot additive edit to the cached keys and values at target positions after prefill — the stored vectors are adjusted once, before any word is written, and are then left alone; the edit induces chain-of-thought style reasoning in small models on GSM8K and other benchmarks with better stability and lower latency than activation steering [90]. V-Steer identifies attention heads in which lower-priority spans dominate privileged ones and applies in-place multiplicative edits to the cached value tensors to boost privileged spans and suppress conflicting ones, raising primary-requirement accuracy from below 18% to 92% on controlled benchmarks with only prefill overhead [83]; that is to say, on prompts where a lower-priority instruction contradicts a privileged one, the primary requirement is met in fewer than one case in five before the edit and in more than nine out of ten after it. Memory Inception encodes reminder text into latent KV banks — stored key-value pairs that correspond to no visible words in the prompt — and inserts them only at the layers and heads that an automated selector finds actually route to the reminder, achieving mid-conversation behavior updates without visible prompt changes [91]; its selector is, in effect, a procedure for locating the layers and heads that read an instruction. LLMSteer processes the shared context under several prefix prompts, identifies tokens that consistently receive high attention, and up-weights them in a way compatible with prefix caching, improving GSM8K among other tasks [92]. Gist tokens show, from the training side, that an entire instruction can be compressed into a few key-value activations that downstream attention reads with little loss [54]. Together these results settle the feasibility of the proposal's KV arm: editing the KV entries of an instruction span is an established, low-overhead intervention, and the entries of a span are sufficient to carry its behavioral effect.


Working at the level of $K_j$ and $V_j$ rather than of whole hidden states allows the proposal to separate two mechanisms that hidden-state patching conflates. The keys of a requirement determine whether later queries select it (the query-key pathway), while its values determine what is written into the residual stream once selected (the output-value pathway). Function-vector and learned-task-vector analyses find that task signals act primarily through the output-value circuits of a few heads [43, 60], which suggests that value-side patching of a requirement's entries may transfer its effect even when keys are unchanged, whereas key-side patching changes which requirement is consulted without changing what is read. Two methodological results make this decomposition tractable. Attribution graphs built from cross-layer transcoders — a replacement for the model's internal computation that re-expresses it in interpretable parts, so that the path from input to output can be drawn as a graph — freeze attention patterns and therefore attribute effects along value pathways only; in other words, the drawn graph shows what was read but not the decision about where to read from; the authors state explicitly that query-key-mediated effects are omitted and must be studied separately [101], which is precisely the gap that the proposal's attention-suppression experiments would fill. AtP*, an attention-aware refinement of attribution patching, addresses the complementary problem in gradient-based attribution: attribution patching estimates the effect of an edit from a derivative instead of actually running the edited model, which is far cheaper but wrong wherever the derivative is flat; attention-softmax saturation is exactly such a place, so the estimate misses effects that flow through attention to particular keys, and the query–key fix (QK-fix) recomputes the softmax exactly for the patched keys [121]. A KV-level design that reports key-side and value-side interventions separately, with the sink entries preserved, would be the first to characterize an instruction's effect along both pathways. Put simply, editing the keys asks whether the requirement still gets selected, and editing the values asks what the requirement contributes once it has been.
Determine whether later queries select this requirement at all.
Patching keys changes which requirement is consulted, without changing what is read from it.
Attribution graphs built from cross-layer transcoders freeze attention and therefore omit this pathway entirely [101] — precisely the gap attention-suppression experiments would fill.
Determine what is written into the residual stream once selected.
Function-vector and learned-task-vector analyses find task signals act primarily through the output-value circuits of a few heads [43, 60], so value-side patching may transfer a requirement's effect even with keys unchanged.
Hidden-state patching conflates the two. A KV-level design that reports key-side and value-side interventions separately, with sink entries preserved, would be the first to characterize an instruction along both.
The proposal predicts that the reasoning phase and the final-answer phase consult different requirements. Evaluating that prediction requires an account of what each phase computes and where its information comes from. This section reviews mechanistic studies of chain-of-thought and arithmetic, evidence on when the final answer is determined and how answer tokens depend on reasoning tokens, the faithfulness problem that complicates the mapping from visible reasoning to internal computation, and the interventions that control reasoning length and style. It concludes with the phase-dependent routing signature these results predict.

Behavioral ablations first established that the content of a chain-of-thought prompt matters less than its structure: invalid rationales retain 80–90% of chain-of-thought performance provided they are relevant to the query and follow the correct ordering of steps [102]; that is to say, demonstrations whose reasoning steps are deliberately invalid still deliver most of the benefit, so what the examples supply is largely the shape of the answer rather than its content, and corrupting the symbols in a demonstration barely hurts while removing its patterns or text does [103]. These findings imply that a reasoning requirement primarily supplies a format or scaffold, which is consistent with the shallow-mechanism null hypothesis of Section 3.1. Mechanistic studies then identified the attention circuits that implement stepwise computation. In controlled settings, chain-of-thought ability emerges as an iteration head, in which one head attends from the current reasoning token to the next input token to be processed while another retrieves the previous reasoning state, so that generated tokens implement a loop over the prompt [93]; concretely, how far through the prompt the model has got is tracked by the text it has already written. Dutta et al. applied activation patching, attention analysis and probing to Llama-2-7B on fictional-ontology reasoning and found a functional rift around the middle layers — a depth above which the computation is doing a visibly different job from below it — parallel answer-generation pathways, and heads that move information from the in-context examples and the question into the generated reasoning tokens [94]. Arithmetic has been traced in more detail. At the last token, early-layer heads carry the operands and the operator across from the question, and the feed-forward layers from the middle onward are where the result is actually computed [95]. Circuit discovery on Llama-3 adds a caution: what does the computing is a sparse set of feed-forward neurons running simple heuristics, not one clean algorithm [96]. Each GSM8K reasoning step therefore has a known division of labor: attention transports operands from prompt or previously generated positions, feed-forward layers compute — a feed-forward layer being the part of each block that processes one position on its own, with no access to any other position. In other words, attention decides which numbers meet, and the feed-forward layers decide what happens to them when they do; the reasoning requirement can influence this only by shaping which positions are transported and whether the step is verbalized at all.
Several results indicate that the final answer is often determined well before the final line is emitted. Answer-convergence analysis finds that on mathematical tasks models settle on their final answer after roughly 60% of the reasoning steps, so that stopping on answer stability saves more than 40% of tokens with little loss [97]; that is to say, the last two fifths of the written working usually does not change the answer it arrives at; probes on hidden states predict chain-of-thought success with 60–76% accuracy before any reasoning token is generated [98] — better than the 50% of a coin flip, and read off the model's state at the moment it has finished reading the problem; and overthinking analyses show that o1-style models spend most of their tokens after the first correct solution [113]. If the answer content is fixed early, the tokens of the final line have little computation left to do beyond formatting, which predicts that their attention should be dominated by the format and unit requirements rather than by the question. Two attention studies of distilled reasoning models describe the answer phase directly. Zhang et al. find that answer tokens attend substantially to reasoning tokens rather than only to the question, identify reasoning-focus heads in middle layers that track the progression of the reasoning, and show by activation patching that perturbing key reasoning-token activations reliably changes the final answer [110]. Thought Anchors identify receiver heads, whose attention over prior sentences has high kurtosis, and show that they concentrate on planning and backtracking sentences; suppressing attention to an anchor sentence changes downstream logits, and resampling the sentence — generating that one sentence again many times and letting each version continue — changes the final-answer distribution [109]. Together these studies give the proposal the two head classes most likely to carry the phase-dependent routing it hypothesizes: reasoning-focus heads during the trace, and receiver-style heads at the transition to the answer.



The mapping from a reasoning requirement to the behavior produces explicit reasoning is complicated by the fact that the produced reasoning need not be the computation that yields the answer. Arcuschin et al. document, without adversarial prompting, implicit post-hoc rationalisation, silent restoration of intermediate errors, and unfaithful shortcuts in frontier and open reasoning models [99]. Attribution-graph analysis of Claude 3.5 Haiku separates faithful cases, cases in which the answer is produced without grounded computation, and backward-chaining cases in which the model works from a user-supplied hint; in the latter the answer token's computation reads from the hint rather than from the reasoning steps — that is to say, the answer was taken from the hint and the visible steps were written afterwards to lead to it — and the same analysis shows that the model's verbal account of addition (column arithmetic) does not match its internal parallel heuristics [100]. For the proposal, faithfulness is both a confound and an opportunity. It is a confound because the presence of reasoning text does not guarantee that the reasoning requirement changed the computation. It is an opportunity because the proposal's causal design can distinguish the cases: if suppressing the reasoning requirement changes the final answer only when the answer tokens depend on the reasoning tokens (in the sense of Zhang et al. [110]), then the requirement is doing computational work; if it changes only the presence of reasoning text, the requirement is a stylistic control.

Interventions that alter reasoning behavior without changing the prompt identify the internal correlates a reasoning requirement must act upon. Venhoff et al. isolate backtracking, uncertainty expression, example testing and verification as separate directions in the activation space of DeepSeek-R1-distilled models, locate causally relevant layers by attribution patching, and steer each behavior with its vector over 500 tasks [104]. ThinkEdit finds that reasoning length is governed by a linear direction in middle-layer residual streams, that a small fraction of attention heads (about 2–4% depending on the model) project strongly onto the short-reasoning side, and that editing only those heads' output projections lengthens chains and improves mathematical accuracy [105]; concretely, 2–4% means only a small minority of the model's attention heads push toward stopping early, and changing just those makes the model reason for longer. Steering vectors trained by reinforcement learning recover much of full fine-tuning's reasoning gain and act, in the last layers, as token-substitution biases that favor structural tokens such as Step at early positions and process words and symbols later [106]. Post-training creates new attention heads in middle-to-late layers that persistently build reasoning pathways, and think-on/off models recruit broader, less efficient head sets when thinking is disabled [108]. On the format side, the only component-level analysis of a specific output requirement is that of Rocchetti and Ferrara, who use cumulative weighted attribution — each component's contribution to the output, weighted and then summed layer by layer so that one can see where the ability accumulates — to show that instruction tuning improves exact word-count control mainly by specializing deeper-layer components, with late attention heads contributing positively in English and final-layer feed-forward blocks compensating in Italian [107]. The consistent picture is that reasoning-style control is carried by a modest number of identifiable middle-layer heads and directions, whereas format control is enforced late; a requirement of each type should therefore be expected to engage different layers, which is testable with the per-layer knockout sweeps reviewed in Section 7, knockout meaning that one connection is switched off at one layer at a time and the effect on the output is recorded.

Combining the results above yields a concrete prediction that the proposal can test. During the reasoning phase, tokens should route to the reasoning requirement through middle-layer heads of the reasoning-focus or iteration type, and this routing should be strongest at step boundaries, where the decision to verbalize a step is made; attention to the format and unit requirements should decline as the trace lengthens, in line with the requirement-attention decay reported by Li et al. [17]. At the transition to the final line, receiver-style heads should read from the reasoning trace to fetch the converged answer [109, 110], while late-layer heads should re-engage the format and unit requirements, whose enforcement the length-control analysis places in deep layers [107]. A model that violates the unit requirement on the last line should therefore show, relative to a compliant generation on the same problem, lower late-layer $A_{t,\text{no-unit}}$ at the final tokens, which is a directly measurable contrast. In other words, the two generations are compared, and the prediction is that the failing one spent less of its deep-layer reading on the no-unit requirement, the requirement that the last line carry no unit, at exactly the moment the last line was written; the causal arm can then test whether restoring that attention by Spotlight-style amplification [66] repairs the violation.
A generation that violates the unit requirement on the last line should show, relative to a compliant generation on the same problem, lower late-layer $A_{t,\text{no-unit}}$ at the final tokens. That is a direct measurement, not an interpretation: two runs, one number each, compared.
If the contrast holds, restoring that attention by Spotlight-style amplification [66] should repair the violation. If it does not, the routing account is wrong even though the correlation held.
The third research question requires showing that the routes identified in Sections 4 through 6 are causally responsible for the behaviors of Section 2. The proposal lists four intervention types: suppressing attention from generated tokens to a requirement's span, patching or replacing the requirement's hidden states or KV entries, removing or rewriting the requirement, and blocking its information flow at particular layers and heads. Each has an established methodological lineage, and each has known pitfalls. This section reviews the patching family and its design choices, attention-specific interventions, interchange interventions and causal abstraction, prompt-level attribution for whole generations, and the evaluation standards that determine whether a causal claim survives scrutiny.

All patching methods follow one template, set by causal mediation analysis: change the input, hold one internal component at the value it would have had under the changed input, and split the total effect into what flows through that component and what flows around it 114. That is to say, the model is run twice on two inputs that differ in one thing, one internal component is forced to take the value it had in the second run, and the question is how much of the change in the output that single substitution accounts for. Causal tracing adapted it to multi-token spans 115, and an instruction span can be treated the same way. The methodology literature contributes the warning that four design choices inside that template can each flip the answer on their own. The table below lists those four: the choice, the options available for it, and what the methodology literature recommends.
| Design choice | Options | What the methodology literature recommends |
|---|---|---|
| Metric | Probability, logit, logit difference, or Kullback–Leibler divergence — the logit difference being the gap between the score of the right answer and the score of the wrong one, and the divergence a single number for how far the whole distribution over next words has moved. | Use logit difference [116], because it cancels effects that push every candidate up or down together. |
| Corruption | Gaussian noise on the span, or counterfactual token substitution. | Counterfactual substitution, not noise [116]. For this proposal the corrupt baseline is a prompt with the requirement rewritten — deleting it also changes length. |
| Direction | Denoising (restore clean state) or noising (corrupt a clean run). | Denoising identifies sufficient components, noising identifies necessary ones — report both [117]. |
| Ablation value | Zero, mean, or resampled from another prompt. | Resample ablation from a different prompt [117]: the component is given a value it really takes on some other input, rather than a value it never takes. |
The proposal's first intervention, suppressing attention from later tokens to a requirement span, has a canonical precedent in attention knockout: Geva et al. block attention from a target position to a chosen span of source positions over a window of layers by zeroing those attention weights, sweep the window, and report per-layer curves that reveal a three-step recall mechanism [125]; that is to say, the block is applied at one depth at a time and the resulting curve says at which depths the connection was load-bearing. Wang et al. applied the same blocking to label-word anchors in shallow versus deep layers and showed that the two produce different effects [70], and Thought Anchors applied attention suppression to sentences inside a generated reasoning trace, measuring the Kullback–Leibler divergence of downstream logits [109]. A known complication is that knockout renormalizes the remaining attention, so its effect mixes removal of information with redistribution of attention — put simply, the shares must still add up to one, so weight taken away from the requirement is handed to everything else, and the observed change could be caused by either; reporting both zero-masking and resample-based key replacement helps separate the two. The inverse manipulation, amplification, is provided by PASTA, InstABoost and Spotlight [64, 65, 66], and using both directions yields a dose–response curve rather than a single ablation point. Competition-of-mechanisms work modifies the attention of specific late heads to specific positions to flip a model between factual recall and in-context copying [142], which is structurally the same manipulation as flipping between a default output format and an instructed one; its reproduction study found that head specialization largely disappears in Llama-3.1-8B and that ablation effects depend on prompt structure and domain [143], a warning that the proposal should run at least two model families.


The proposal's second intervention, patching or replacing a requirement's hidden states, is an interchange intervention in the sense of causal abstraction: a high-level causal model is an abstraction of the network if swapping in a representation from a counterfactual input has the same effect as intervening on the corresponding high-level variable [127]. In other words, one writes down a small hand-made diagram of what the model is supposed to be doing, and the diagram earns the right to be called a description of the network only when editing the network at the matching place produces exactly what editing the diagram would predict. Distributed alignment search learns the subspace on which such interventions succeed, and its Boundless variant (Boundless DAS) applied it to Alpaca-7B on a numeric instruction task, finding that the instruction-tuned model implements an interpretable boolean causal model in identifiable residual subspaces robust to instruction wording [128]; this is the closest precedent for hypothesizing a small causal model of requirement compliance (reasoning required, bare number required, units forbidden) and testing it by interchange interventions on requirement-span representations. Patchscopes unifies logit lens, causal tracing and cross-prompt patching as one operation, patching a representation from a source prompt, layer and position into a target and decoding it in natural language — concretely, the internal state in question is dropped into a second prompt that asks the model to describe it, and the model's answer is read as a report on what that state held — which provides a way to ask what a model believes a requirement says at a given layer [129]. Task- and function-vector patching [42, 43] supplies a further condition, replacing the requirement text with its vector, that tests whether the requirement is needed as tokens at all. Causal scrubbing offers a whole-hypothesis test: resample every activation the hypothesis deems irrelevant from other prompts sharing the same requirement and report the fraction of behavior preserved [126]. That is to say, everything the account claims does not matter is replaced with material from elsewhere; if the account is right the behavior survives, and how much of it survives is the score. Sparse feature circuits move the same logic to sparse-autoencoder features, enabling ablation of individual requirement-related features during generation [137].

Because the proposal's outcomes are properties of complete generations, it needs attribution methods that operate at that granularity. ContextCite ablates random subsets of context sentences, records the log-probability of a fixed response or span, and fits a sparse linear surrogate whose weights attribute the response to each source [130] — that is to say, it learns a simple formula predicting how likely the response is from which sentences were present, and the size of each term in that formula is that sentence's credit; applied to the requirement block, it scores each requirement's influence on, for instance, the last line, at the cost of an additivity assumption that may be violated when requirements interact. Put simply, the method assumes each requirement's contribution simply adds up, which is exactly what stops being true when two requirements pull against each other. AttriBoT makes exact leave-one-out attribution — rerun the model with one requirement removed, once per requirement, and take the difference — tractable with cached activations and hierarchical attribution [132], and Learning to Attribute with Attention trains a per-head weighting so that attention predicts ablation-based attributions, which identifies the heads whose attention is causally informative and is therefore a principled way to select heads for knockout [131]. Inseq provides step-wise attribution of each generated token to source and previously generated tokens with aggregation over spans [133], and contrastive explanations attribute the difference between two candidate tokens, such as a bare number versus a number followed by a unit, rather than the raw probability [134]; AttnLRP propagates relevance through attention and normalization layers for token-level attributions in large models [135]. Any rewriting of a requirement must be compared against a distribution of surface-form variants, since format changes alone swing accuracy by tens of points [23].

Several results define what the proposal must report for a routing claim to be credible. Makelov et al. demonstrate an interpretability illusion in which patching along a learned subspace produces the target behavior by activating a dormant parallel pathway that the model does not use on clean inputs, and propose diagnostics that check whether the patched direction is actually used [138]; that is to say, the intervention works and still explains nothing, because it switched on machinery the model never reaches for on its own; any learned-subspace intervention on requirement representations needs this check. Miller et al. show that circuit faithfulness scores — how much of the model's behavior a proposed set of components reproduces on its own — change substantially with the ablation type, direction and metric, to the point that the best circuit changes, and recommend reporting multiple settings [139]. Shi et al. formalize mechanism preservation, localization and minimality as statistical hypothesis tests with error control and find that few published circuits pass all of them [140]; concretely, preservation asks whether the proposed components alone still produce the behavior, localization whether everything outside them can be disturbed without effect, and minimality whether any of them can be dropped. For interventions on open-ended generation, Arditi et al. provide the exemplar of scoring directional ablation by behavioral metrics of whole generations alongside capability benchmarks [55], and Pres et al. propose criteria for steering evaluations: deployment-like contexts, prompting baselines, likelihood shifts of both target and non-target behaviors, and effect sizes with uncertainty [141]. The reproduction study of Dotsinski et al. adds that head-level findings can be model-specific [143]. The table below maps each of the proposal's planned interventions to its methodological precedent and the pitfall that precedent identified. Each row is one planned experiment: the left column says what would be done to the model, the middle column names the published work that has already done something of that shape, and the right column names the mistake that work showed is easy to make.
| Planned intervention | Methodological precedent | Pitfall to control |
|---|---|---|
| Suppress attention from generated tokens to span $\mathcal{T}_i$ | Attention knockout over layer windows [125]; anchor blocking [70]; sentence suppression in reasoning traces [109] | Renormalization mixes removal with redistribution; exclude sink positions [75, 76]; add uniform / frozen attention controls [81] |
| Patch or replace hidden states of $I_i$ | Interchange interventions and Boundless DAS on an instruction-tuned model [127, 128]; Patchscopes cross-prompt patching [129] | Dormant-pathway illusion [138]; use resample ablation with a rewritten requirement, not zero ablation [117]; report denoising and noising [116] |
| Patch or replace KV entries $K_j, V_j$ of $I_i$ | KV Cache Steering [90]; V-Steer value edits [83]; Memory Inception layer/head selection [91] | Preserve sink entries; separate key-side from value-side effects; attribution graphs omit QK effects [101], AtP* QK-fix [121] |
| Remove, replace or rewrite $I_i$ | ContextCite and leave-one-out attribution [130, 132]; per-requirement steering-vector contrast [38] | Length and position shift: use schemas [124]; surface-form sensitivity [23]; added-clause cost on GSM8K [26] |
| Block information flow of $I_i$ at specific layers / heads | Path patching [118, 119]; edge attribution patching with integrated gradients (EAP-IG) and AtP* screening [121, 122]; memory-vs-context head pruning [72] | Faithfulness not robust to ablation choice [139]; hypothesis tests for localization [140]; model-family dependence [143] |
| Score on GSM8K accuracy, reasoning presence / length, format and unit compliance | Generation-level scoring of directional ablation [55]; resampled continuations [109]; steering evaluation criteria [141] | Report accuracy and compliance jointly [16]; text-based not first-token compliance [36]; judge bias on soft requirements [29] |

layer, a component (here mlp_output) and an intervention_type (here addition), the model is wrapped once, and generation proceeds normally with the source representation passed in. The two strips at the bottom are the result on the same opening sentence: with happy added, Lucy is happy and thanks the man; with sad added, the same setup turns and Lucy is scared. The point for a workflow is that the intervention is configuration rather than a patched forward pass.Two libraries support interventions during multi-token generation with Hugging Face models. pyvene specifies interventions declaratively by layer, token positions, subspace and type, including interchange, zero, noise and trainable interventions, and implements DAS and Boundless DAS [144]. NNsight defines intervention graphs over any PyTorch model, including inside .generate() loops, and NDIF runs them remotely on large models [145]; this is the more natural fit for KV-level and hidden-state patching over full GSM8K solutions. Gemma Scope sparse autoencoders provide off-the-shelf feature bases for every layer of an open model, should the proposal pursue feature-level decomposition of requirement signals. In other words, both libraries allow a value inside the model to be changed midway through writing an answer, which ordinary inference code does not permit.
This section consolidates the preceding review into an evidence map for the three research questions, identifies the gaps the proposal is positioned to fill, states the hypotheses the literature supports, and lists methodological recommendations. The organizing claim is that every component of the proposed causal chain $I_i \rightarrow \{K,V\} \rightarrow \text{heads} \rightarrow B_j$ has been demonstrated in isolation, but no study has traced several heterogeneous natural-language requirements through that chain simultaneously on a reasoning task, with generation-level outcomes and with attention to how the routing shifts between the reasoning and answer phases. That is to say, the chain runs from the words of one requirement, to the keys and values the model stored for those words, to the heads that read them back, to the behavior that finally appears in the output; each link has been demonstrated on its own, and the missing work is following all four at once for several requirements in the same prompt.

The table lists, for each level of the proposal, what the literature has established, the closest existing work, and what remains open. Each row is one of the three levels defined in Section 1, read from left to right as settled, cited, and still missing.
| Level | Established | Closest work | Open |
|---|---|---|---|
| L1 Requirement–behavior mapping | Requirements can be scored per requirement; compliance decays with count and is worst for format / lexical types; format requirements trade off against reasoning; position, surface form and conflict all matter | IFEval, FollowBench, InFoBench, ComplexBench [1, 2, 3, 4]; MathIF [16]; When Thinking Fails [17]; Let Me Speak Freely [18] | No study varies requirements one at a time on a fixed reasoning task and records the full behavior vector (accuracy, reasoning presence / length, last-line format, unit removal); no ablation of requirement order within one block |
| L2 Internal routing | Instruction tokens are persistently consulted; per-requirement steering vectors exist; instruction tokens act as anchors read by deep layers; attention to instructions decays over generation; heads specialize by instruction source; attended prompt positions are stable across decoding | Wu et al. [34]; Stolfo et al. [38]; Attention Satisfies [62]; Instruction Anchors [69]; Label Words are Anchors [70]; Li et al. [67]; Dongre et al. [68]; SnapKV [88] | No per-requirement $A_{t,i}^{(l,h)}$ maps for a multi-requirement prompt; no comparison of reasoning-phase versus answer-phase readers; no key-versus-value decomposition of an instruction's effect |
| L3 Causal influence | Attention knockout and amplification on instruction spans change compliance; KV edits change reasoning and priority; instruction-conditioned behaviors are mediated by steerable directions; generation-level evaluation of interventions is feasible | Geva et al. [125]; PASTA / InstABoost / Spotlight [64, 65, 66]; KV Cache Steering [90]; V-Steer [83]; Arditi et al. [55]; Thought Anchors [109]; Boundless DAS [128] | No segment-specific counterfactual replacement of one requirement's KV entries; no per-requirement, per-layer knockout sweep scored on a behavior vector; no test of whether the same heads mediate different requirements |

Five gaps recur across the sections above. Each corresponds to a component of the proposal that no existing study has delivered.
Existing mechanistic work treats one instruction type at a time (a format requirement [38], a length requirement [107], a modality instruction [69], a persona [56]). No study traces a reasoning requirement, an answer-format requirement and a lexical requirement through the same forward passes and compares their carriers. That is to say, each of the four studies cited followed a single requirement type, so nothing in them says whether two requirements are carried by the same parts of the model or by different ones.
Attention decay to instructions is measured across dialogue turns [67] or toward a single goal span [68]; requirement-attention decline within reasoning is reported in aggregate [17]. Per-requirement attention maps over the reasoning and answer phases of one generation do not exist.
KV edits so far add a vector at the last prompt token [90], scale whole priority spans [83] or insert new banks [91]. Replacing the keys and values of one requirement with those of another, and separating key-side from value-side effects, is untested.
Most patching work scores a next-token logit difference [116]; the few generation-level evaluations use a single behavior [55, 109]. Scoring each intervention on accuracy, reasoning presence and length, final-line format and unit removal simultaneously, as MathIF does behaviorally [16], has not been done mechanistically. In other words, without all four scores at once an intervention that improves the format by deleting the reasoning would be recorded as a success.
Function-vector work shows instruction- and demonstration-derived vectors use different heads [44]; conflict studies show separate subspaces for conflict types [85]. Whether two requirements in one prompt share reader heads, and whether interference at the head level explains the count-dependent decay of Section 2, is open.

The evidence reviewed licenses the following hypotheses, stated so that each has a clear falsifying observation. Each row gives the hypothesis, the evidence that makes it worth testing, and the specific result that would show it to be wrong.
| Hypothesis | Supporting evidence | Falsifying observation |
|---|---|---|
| H1. Each requirement type has a separable residual-stream signature extractable by with / without contrast, and the format-type signatures are concentrated in late layers. | [36, 38, 39, 55, 107] | Steering with one requirement's vector changes the behaviors of the others as much as its own; signatures are not separable under the causal inner product [50] |
| H2. Requirement tokens function as anchors: attention to a requirement concentrates on its delimiter positions in shallow layers, and later readers consult those positions rather than the content words. | [69, 70, 78, 87] | Knockout of delimiter positions has no effect while knockout of content words does |
| H3. Reasoning tokens route to the reasoning requirement through middle-layer heads, and last-line tokens route to the format and unit requirements through late-layer heads; the two reader sets are largely disjoint. | [62, 93, 105, 107, 109, 110] | The same heads carry the largest $A_{t,i}$ for every requirement in both phases, or knockout effects are phase-independent |
| H4. Attention share to the format and unit requirements declines over the reasoning phase, and violations on the last line are preceded by lower late-layer $A_{t,i}$ than compliant generations. | [17, 67, 68, 111] | Compliant and violating generations show indistinguishable $A_{t,i}$ trajectories |
| H5. Suppressing attention to a requirement, or replacing its KV entries with a rewritten requirement's, changes primarily that requirement's behavior, with value-side replacement sufficient to transfer the effect. | [43, 60, 83, 90, 125] | Effects are diffuse across behaviors, or key-side replacement is required for any transfer |
| H6. Adding requirements degrades compliance through competition for attention share among reader heads rather than through loss of the requirement representations, so that amplification restores compliance. | [10, 11, 12, 65, 66, 68] | Probes show requirement information absent from the residual stream after many requirements are added, and amplification does not help |

The methodological literature suggests the following order of operations, in which cheap descriptive measurements narrow the search before expensive causal tests are run, and in which every causal test is scored on the full behavior vector.
Vary each requirement independently (present / absent / rewritten / repositioned) on a fixed GSM8K subset with a no-requirement control [26]; score accuracy, reasoning presence and length, strict and loose last-line format, and unit removal with code checks [1, 4]; randomize order and surface form [23]; include trip-wire items for shortcut compliance [9, 12].
Compute $A_{t,i}^{(l,h)}$ per requirement with sink positions excluded and delimiter positions reported separately [75, 78]; summarize by generation phase, that is to say average the per-step numbers separately over the reasoning tokens and over the final-line tokens; rank heads by receiver-style kurtosis [109], query-focused retrieval score [74] and learned attention attribution [131]; locate the internalization layer by layer-wise context masking [49].
Extract a per-requirement vector by with / without contrast [38, 57]; test separability by causal-inner-product orthogonality and by cross-steering [50]; monitor each vector's projection over the generation [56]; check for shallow output-bias implementations [41].
Screen heads and edges with EAP-IG and AtP* — that is to say, cheap derivative-based estimates that rank every head and connection so that only the top candidates need to be tested properly — using a rewritten-requirement counterfactual and schema-aligned spans [121, 122, 124]; confirm with windowed attention knockout [125], path patching [119], and Spotlight-style amplification for dose–response [66]; replace KV entries per segment, key-side and value-side separately, preserving sinks [83, 90].
Score every intervention on the full behavior vector with resampled continuations [55, 109]; report zero, mean and resample ablations in both directions [116, 139]; apply dormant-pathway diagnostics to any learned subspace [138]; use hypothesis tests for localization and minimality [140]; replicate on at least two model families [143].

The literature has matured to the point where every tool the proposal needs exists and has been validated on a related problem: verifiable per-requirement scoring, contrastive extraction of per-requirement vectors, anchor-token and receiver-head analysis, KV-level editing, position-aware circuit discovery, and generation-level evaluation of interventions. What is missing is their combination on a single, well-controlled multi-requirement reasoning task, with the routing of each requirement followed across the phases of generation and confirmed by segment-specific intervention. The behavioral facts that motivate the question, in particular the count-dependent decay of compliance and the trade-off between reasoning and format requirements, are robust across models and benchmarks, and the most plausible mechanistic account of them, competition for attention share among instruction readers, is testable with the methods reviewed here. The proposal's expected contribution is therefore not a new tool but the first mechanistic account of how a language model decomposes, stores, consults and executes several prompt requirements at once, and of where in that chain the observed failures arise. Put simply, the question is no longer whether a model obeys, but which part of it stopped reading which requirement, and when.
The list is grouped by theme; numbering is global and matches the bracketed citations in the text. Every arXiv identifier was resolved against arxiv.org, the ACL Anthology, OpenReview or the publisher page at the time of writing. Entries whose author lists could not be re-verified are marked as such and should be checked before citation.