Dynamic MAS Architectures — Synthesis of Two Independent Reviews

The adjudication layer: where the two reviews agreed, where they conflicted, and what the primary sources say. Supersedes both underlying reports where they disagree.

Adjudicated 2026-07-20 10 conflicts resolved ≈54 papers total
This page settles conflicts. Where the Fable and Sol reports disagree on a checkable fact, the verdicts below supersede both.

This document merges and adjudicates two independently-produced reviews of the same research question: DYNAMIC_MAS_LITREVIEW_FABLE.md (Claude Fable 5, backed by a 70-subagent (automated LLM worker) extraction-and-adversarial-verification pipeline over primary sources) and DYNAMIC_MAS_LITREVIEW_SOL.md (GPT-5.6-Sol at ultra reasoning effort, with independent web research). Neither reviewer saw the other's draft. Where the two disagreed on a checkable fact, this synthesis re-fetched the primary source and records the verdict. Date: 2026-07-20.

1. Why two reviews, and what the disagreement pattern shows

Running two independent reviewers over the same brief is a verification instrument: claims both reviewers reach independently are very likely robust, and their disagreements localize exactly the places where the literature's own record is confusing. The comparison's principal result is convergence. Both reviews, without coordination, converged on the same organizing taxonomy — what structure is decided (edges, team, roles, workflow code) crossed with when it is decided (offline per-domain; per-query at inference; per-round during execution) — and on the same four field-level verdicts (Section 4). The disagreements that did occur were almost all about venue presentation tiers and evaluation-hygiene severity, which is itself a finding: the field's public record of "which of these papers got orals" is genuinely hard to reconstruct, because conference virtual sites label the same paper differently on different pages.

The two reviews also had complementary coverage gaps, so the merged corpus is larger than either alone. The Fable review covered four core systems Sol missed entirely — MDAgents (a NeurIPS 2024 Oral squarely in scope, since its whole contribution is per-query structure selection), Puppeteer (NeurIPS 2025), MasRouter (ACL 2025), and AMAS (EMNLP 2025 Industry) — plus a 27-paper verified frontier (graph-diffusion generators, reinforcement-learning stabilization work, test-time adaptation, and the boundary surveys). The Sol review covered three systems the Fable search missed — AGP (Adaptive Graph Pruning, ECAI 2025; arXiv 2506.02951), CARD (conditional graph designer, ICLR 2026 Poster; arXiv 2603.01089), and DySCo (training-free trust-weighted sparse consensus; arXiv 2606.01828) — and consistently recovered finer-grained protocol detail (exact validation/test counts, model snapshot strings, per-component cost accounting).

2. Adjudicated disagreements

Each row below was a genuine factual conflict between the two reports. The verdict column states what the primary evidence shows; both underlying reports remain as written, so this table is the correction layer over them. The two tier conflicts were resolved by fetching the conference pages during synthesis; the remaining rows were resolved from evidence already inside one of the two reports.

(Table content is kept in English; numbers and identifiers are verbatim from the report.)

(scroll right for more columns →)

#QuestionFable saidSol saidVerdict (evidence)
1GPTSwarm ICML 2024 tierOralPoster (citing the poster page + PMLR)Oral — icml.cc/virtual/2024/oral/35447 is headed "Oral — GPTSwarm: Language Agents as Optimizable Graphs" (fetched 2026-07-20). The poster page Sol cited is the poster slot that ICML orals also receive.
2MaAS ICML 2025 tierSpotlightOral (citing oral session 46894)Oral — icml.cc/virtual/2025/session/46894 ("Oral 1A Alignment and Agents") lists MaAS (fetched 2026-07-20). ICML 2025 labels oral papers' poster slots "Spotlight Poster", which is what misled the Fable verifier. Fable's report has been corrected in place.
3DynaSwarm ↔ AMAS relationshipSame lead authors; near-identical method and numbers; AMAS is the published successor (EMNLP 2025 Industry Track)Treated DynaSwarm in isolation as withdrawn; did not cover AMASAdopt Fable's finding — DynaSwarm (Leong & Wu, withdrawn 2025-08-12 for a "content error") and AMAS (Leong, Li, Wu et al., aclanthology.org/2025.emnlp-industry.144) share lead authors, the K=4 candidate-bank + LoRA-selector design, and task scores (accuracy or pass-rate fractions) matching to within ±0.004 on overlapping settings. Cite AMAS; treat DynaSwarm's numbers as superseded.
4MASS venueICLR 2026 reported, tier unverifiedICLR 2026 Poster (OpenReview I05H9RUzHB)Adopt Sol's version — the OpenReview record confirms acceptance as poster.
5AgentDropout ACL 2025 tierNot published by ACL; "tier n/a"Poster (citing an underline.io presentation page)Tier n/a stands — ACL does not publish oral/poster tiers in the Anthology; an underline listing shows a presentation existed, not its tier. Sol's "Poster" is a reasonable default, not a verified fact. Same reasoning applies to Sol's Poster labels for SwarmAgentic, AnyMAC-adjacent EMNLP papers, and DyLAN.
6FlowReasoner venue detailarXiv preprint (OpenReview submission)Adds: Poster at the ICML 2025 "Multi-Agent Systems in the Era of Foundation Models" workshopAdopt Sol's addition — icml.cc/virtual/2025/49324. Both agree there is no verified archival main-venue acceptance.
7ADAS meta-agent model"GPT-4" (as the paper's main text names it)gpt-4o-2024-05-13 (snapshot-level)Minor; unresolved at snapshot level — both agree on the architecture (frozen LLM programmer) and the GPT-3.5 executor; the exact meta-model snapshot differs between paper versions and was not re-fetched. Flagged rather than adjudicated.
8GPTSwarm HumanEval protocolOptimization pools overlap tiny evaluation pools (general concern)Sharper: the online prompt/demonstration optimizer consumed the full 164-problem set that is also reported (first-attempt pass rate rising from 0.76 to 0.88), and the 20 crossword puzzles were both optimized and evaluatedAdopt Sol's more specific reading — it is consistent with Fable's verifier notes and more specific.
9AnyMAC training dataTrain/test disjointness "not asserted"Sharper: where a dataset has no train split, the router trains on the first 80 test items — an explicit optimize-on-test caseAdopt Sol's more specific reading.
10G-Designer per-table numbersOnly abstract-confirmed numbers quoted (table unpinnable at 2nd decimal)Quotes full table (avg 89.84 vs GPTSwarm 87.32; PHP 88.17)Compatible — Sol's figures fall inside the ranges Fable's verifier observed; treat Sol's as the working numbers with Fable's pinning caveat attached.

Two meta-lessons from the adjudication. First, both directions of tier error occurred — Sol assigned GPTSwarm too low a tier, and Fable did the same to MaAS — and both errors trace to the same root cause: conference virtual sites expose multiple pages per paper with different labels. Any survey asserting tiers should cite the oral-session page, not the paper's own virtual page. Second, on evaluation hygiene the reviewer with the finer protocol detail (usually Sol) consistently identified the stronger problem wherever the two reports differed in severity — an indication that closer protocol reading tends to surface additional issues across this corpus.

3. The merged corpus and the unified map

Merging the two reviews (and removing the withdrawn DynaSwarm as a separate entry) gives 29 fully-treated core systems and roughly 25 verified frontier/boundary works — approximately 54 papers total. The unified placement below uses the taxonomy both reviews independently arrived at. Rows are decision-timing regimes; columns are mechanism families; boundary controls sit outside the grid. Papers contributed by only one review are marked (F) — covered only by the Fable review — or (S) — covered only by the Sol review.

(Table content is kept in English; numbers and identifiers are verbatim from the report.)

(scroll right for more columns →)

Timing \ MechanismLearned edge distribution / pruningConditional graph generatorSupernet / router / trained generatorSearch over workflow codeUntrained LLM orchestration
Offline per-domain (structure frozen before deployment)GPTSwarm; AgentPrune; AgentDropout; Heterogeneous Swarms (F)ADAS; AFlow; AgentSquare; MASS; SwarmAgentic; AgentSwift (F); W4S (F)
Per-query at inferenceAMAS (F) (selection from a learned bank)G-Designer; ARG-Designer; AGP (S); CARD (S); GTD (F); RADAR (F); GoAgent (F)MaAS; MAS-GPT; ScoreFlow; MasRouter (F); DAAO (F); FlowBank (F); MetaFlow (F)FlowReasoner (per-query + repair)MDAgents (F); AutoAgents
Per-round during executionSafeSieve (F) (progressive pruning)AnyMAC; Puppeteer (F)EvoMAC; TacoMAS (F)DyLAN (runtime part); Captain Agent (F); AgentNet; DySCo (S); ARMOR-MAD (F); SID (F)

Boundary controls both reviews agree on: MacNet (static topology scaling; ICLR 2025), MAST (failure taxonomy; NeurIPS 2025 Spotlight, (F)), Optima (fixed graph, trained communication; ACL 2025 Findings), plus the position paper and two surveys (F). The map makes the field's shape visible at a glance: the offline column is populated by the 2024 generation, the per-query column is where 2025–2026 activity concentrates, and the per-round row remains thin — both reviews independently called runtime-adaptive topology the smallest and hardest regime, and Sol's DySCo plus Fable's TacoMAS/SID/ARMOR-MAD entries show 2026 work beginning to fill it, largely with training-free controllers.

The verified presentation-tier record, post-adjudication. Oral: GPTSwarm (ICML 2024), MDAgents (NeurIPS 2024), AFlow (ICLR 2025), MaAS (ICML 2025), ARG-Designer (AAAI 2026). Spotlight: G-Designer (ICML 2025), MAST (NeurIPS 2025). Everything else verified sits at poster/accepted level or preprint. The pattern that motivated this review — a concentration of oral and spotlight presentations in this research line — is confirmed, and the orals concentrate exactly on the mechanism novelties: the first learned graph (GPTSwarm), the first query-adaptive template selector at scale (MDAgents), the first systematic code-space search (AFlow), the first supernet (MaAS), and the first from-scratch autoregressive graph generator (ARG-Designer).

4. What both reviews independently concluded (high-confidence findings)

Because these four verdicts were reached by two reviewers working from different source passes, they are the most defensible take-aways of the whole exercise.

Finding 1 — Moderately sparse structures match or outperform dense ones at lower cost. Every pruning-family paper in both reviews, plus the causal information-propagation study (EMNLP 2025), finds that removing a large fraction of edges or context preserves or improves accuracy while cutting 20–95% of tokens. Both reviews also offer the same explanation: extra messages amplify correlated errors and let mistakes propagate, so selective exposure outperforms broadcast.

Finding 2 — Query-adaptivity pays exactly when difficulty varies. The cleanest evidence in both reports is MDAgents' ablation (adaptive structure 81.2% versus 64.2–71.6% for each of its own structures applied uniformly) and MaAS's early-exit economics (6–45% of competitors' inference cost). The mechanism is the same everywhere: not overspending on easy queries funds the hard ones.

Finding 3 — The margins over strong simple baselines are modest, and evaluation protocols are frequently under-documented. Both reviews independently flagged: the repeated six-benchmark set (MMLU, GSM8K, MultiArith, SVAMP, AQuA, HumanEval); deltas of 1–4 points against test sets sometimes as small as 20–254 items; widely under-documented train/test partitions (with GPTSwarm's crossword/HumanEval loops and AnyMAC's first-80-test-items as the clearest optimize-on-test cases); the strong single-agent prompting baseline PHP sitting above or near several MAS systems; and MASS's staged ablation showing prompt optimization contributes roughly three times more than topology optimization in a joint system. Neither review found a dedicated equal-budget comparison against one stronger single model anywhere in the corpus.

Finding 4 — Structural decisions are cheap; the field has learned to make the controller small. GPT-2-Small routers, VGAE generators, LoRA selector heads, MiniLM-embedding controllers, one 32B generator call priced at half an executor call — both reviews note that deciding the structure costs orders of magnitude less than running the agents, which is what makes per-query adaptivity economically sensible at all.

5. Merged implications for KV-cache-level agent communication

Both reviews wrote their implications sections independently, and they interlock rather than repeat — evidence that the connection between this literature and serving-tier communication research is real rather than forced. Read together they yield one combined agenda.

The shared starting observation: the decision-timing taxonomy is a reusability taxonomy for the KV cache (the per-token key/value attention state a transformer accumulates while reading a context, which serving systems reuse to avoid recomputation). Offline-static graphs allow precomputed, amortized cross-agent prefix reuse; per-query graphs make the reuse pattern known only after a cheap controller pass (just-in-time placement); per-round graphs are the hard case, because a message's future readers are unknown when its KV blocks are produced. Fable adds the trend observation that the field's own token-cost pressure is pushing topologies toward sparse and sequential shapes (AnyMAC's cascade, Puppeteer's one-agent-per-step, MaAS's layered sampling) — precisely the monotone-prefix shapes most compatible with KV-cache handoff. Sol adds the systems design that exploits this: represent dialogue state as immutable content-addressed KV blocks on a message DAG, pass references along edges instead of copying transcripts, and use copy-on-write for private continuations.

The shared gap identification: every cost term in this literature prices tokens or calls; none prices the serving tier. Both reviews independently propose the same unclaimed contribution — a topology objective whose cost model distinguishes warm-prefix tokens from cold-prefill tokens and charges for KV transfer/residency — and both note the optimization machinery (policy gradient over sampled structures against a scalar cost) is already standard in this literature, so only the cost model needs replacing. Sol extends this with concrete scheduler-side ideas (treat router probabilities as prefetch signals; let early-exit confidence drive cache retention) and a joint evaluation metric set (cache-hit ratio, bytes moved, prefill/decode split, peak live KV, time-to-first-token under a memory budget). Fable frames the same endpoint from the venue side: no paper in the ~55 reports TTFT, prefill/decode breakdowns, or cache behavior, so a model-fixed, harness-fixed, serving-tier comparison of dynamic-versus-static topologies fills a slot the accuracy venues have left empty.

One caution both reviews raise in different words: AnyMAC's learned context gates are the closest existing abstraction to KV selection, but everything in this corpus communicates text through APIs. Dynamic-topology research and KV-level communication research are currently non-overlapping literatures; the bridge is open in both directions, and a study that connects them can draw on both fields' established baselines.

6. Recommended reading order and division of labor between the three documents

A suggested entry order for the combined corpus: begin with the Fable review's Sections 1–2 for the framing and the master table; then read four papers in full — GPTSwarm (the origin of learned edges), MaAS (the supernet and cost-aware per-query sampling), AFlow (code-space search, and the strongest cost result), and MDAgents (proof that adaptivity needs no training) — then the Sol review's per-paper "Evaluation" paragraphs for the practice of reconstructing evaluation protocols in detail, and finally MAST for why any of this matters operationally. The Fable review's Section 9 is the map of the 2026 frontier; this synthesis's Section 2 is the correction layer to apply on top of both reports; and any citation of presentation tiers should use the adjudicated list in Section 3, which supersedes both underlying reports where they conflict.

The following uncertainties remain unresolved by either review: the ADAS meta-model snapshot (row 7 above); ACL/EMNLP/COLM presentation tiers (not published by those venues); the exact provenance of ARG-Designer's reference graphs and of FlowReasoner's ~1,400 training traces relative to its test benchmarks (both flagged as the corpus's most significant unresolved contamination questions, by Sol and Fable respectively); and every 2026 preprint's eventual venue. These are marked in the underlying reports and should be re-checked before any of those specific claims is cited in a paper.