REVIEW 4 major objections 5 minor 9 references
AIDG: A Formal Decomposition of Information Extraction and Containment Asymmetries in Multi-Turn LLM Dialogue
T0 review · 4 major / 5 minor · reviewed 2026-08-02 · deepseek-v4-flash
Pith's one-line read In 439 adversarial LLM dialogues, holders of a secret won 86.6% of games, exposing a systematic extraction-versus-containment gap.
desk verdict Useful new benchmark and a plausible asymmetry, but every headline number rests on an unvalidated LLM arbiter; engage, but don't trust the magnitudes yet. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing machinery is AIDG itself, a two-player partially observable stochastic game (POSG) that formalizes one agent trying to infer a private variable held by another. Performance is split into two paired-comparison ELO ratings, one for extraction (C_ELO) and one for containment (V_ELO), updated through a logistic rating model with a turn-decay multiplier M(t) = (17−t)/8 that rewards early correct answers in the constrained game. The quantitative results rest on an automated Arbiter, an LLM judge run at low temperature that classifies each Holder turn into leakage categories — explicit disclosure, confirmational leak, semantic paraphrase, and implicit admission — and on a 'no direc
What would settle it
Have human experts label a stratified sample of AIDG-I and AIDG-II turns using the same four leakage categories the Arbiter uses, then compare the two label sets. If the Arbiter misses many confirmational leaks or flags paraphrases humans accept, the measured 86.6%/13.4% split and 7.75x odds ratio would not survive; high agreement would confirm them. A second check: rerun AIDG-II with the 'no direct guessing' rule removed and see whether the 41.3% disqualification-driven Holder wins persist.
Extended reading notes
Core claim
The central discovery is a systematic capability asymmetry between information extraction and information containment in multi-turn LLM dialogue. Across two distinct game formats — free-form social deduction (AIDG-I, 87.2% Holder wins) and a constrained yes/no/maybe word-guessing game (AIDG-II, 85.3% Holder wins) — models are substantially better at protecting a secret than at extracting one. The paper identifies two bottlenecks: confirmation-based attacks succeed at 7.75 times the rate of blind extraction, and 41.3% of deductive failures in the constrained game are instruction-following violations (askers make a direct guess before the lock turn), with disqualification rates ranging from 0%
Load-bearing premise
The load-bearing premise is that the automated Arbiter correctly classifies leakage in free-form AIDG-I dialogue; the paper's own limitations section concedes that automated semantic equivalence judgments may miss subtle implicatures, indirect confirmations, or pragmatic cues, and the 7.75x odds ratio and 86.6% Holder win rate stand or fall with those judgments.
Editorial extensions
If this is right
- Evaluation of dialogue agents should report extraction and containment as separate scores; a single win-rate scalar hides the fact that a model can be strong at one and weak at the other.
- A partially informed attacker is a much larger threat than a blind one, so defense evaluations must include confirmation-style attacks to be realistic.
- Common 'no direct guess'-type instruction constraints are unreliable under conversational load — violation rates reach 72% for some models — and are uncorrelated with model scale.
- Differences between frontier models in adversarial dialogue come mainly from offensive planning ability, since defensive skill is nearly identical across models.
- The paper interprets the containment-over-extraction gap as a structural consequence of the two roles' different computational demands, not as a surprising model defect.
Reading between the lines
- The strong Holder performance may partly reflect alignment training that over-optimizes refusal behavior; if so, better extraction skill would not automatically make models worse at containment — the two capabilities could be trained separately. This is our interpretation, not the paper's.
- A human-annotation study over a sample of AIDG-I turns could test whether the Arbiter misses subtle implicatures; the paper's numbers should be treated as conditional on LLM-as-judge validity until such a check is done.
- The turn-decay multiplier's reward for early locks could be reshaping Seeker strategy in AIDG-II; removing or flattening it would reveal how much of the 41.3% disqualification rate is a reaction to time pressure versus genuine instruction-following limits.
- The framework is directly portable to mixed human-model games and to multi-fact relational secrets, both of which the paper lists as limitations; those settings would test whether the asymmetry persists beyond atomic secrets and closed ontologies.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces AIDG (Adversarial Information Deduction Game), a two-task framework for decomposing multi-turn LLM dialogue performance into an extraction capability (Seeker) and a containment capability (Holder). AIDG-I uses free-form social deduction with two attack modes (confirmation vs. blind extraction), and AIDG-II uses a constrained 20-questions setting with yes/no/maybe responses and a no-direct-guessing rule. The authors report 439 games across six LLMs, claiming a large and consistent holder advantage (86.6% Holder wins), a 7.75x confirmation advantage, a 41.3% disqualification rate due to constraint violations, and a capability decomposition via a dual-ELO system showing low defensive variance (σ=1.9) and high offensive variance (σ=53.3). The paper releases datasets, prompts, and analysis scripts.
Significance. If the empirical claims hold, the paper offers a useful way to separate two partially independent capabilities in multi-turn LLM evaluation, which is more informative than a single win-rate scalar. The open release of transcripts and scripts is a genuine strength, as is the explicit statement of design assumptions. However, the central quantitative findings—the Holder win-rate asymmetry, the confirmation odds ratio, and the dual-ELO gaps—are all downstream of an unvalidated LLM Arbiter and of a rule-based disqualification mechanism that conflates constraint adherence with defensive skill. The paper's own Limitations section concedes the Arbiter validity issue but the manuscript does not supply the validation needed to make the headline claims load-bearing.
major comments (4)
- [§4.2 vs. Table 2] The protocol states that AIDG-I yields 5 tournaments × 60 games = 300 games, but Table 2 reports only 289 AIDG-I games, and the combined total of 439 = 289 + 150. Eleven planned games are missing with no explanation. If games were dropped because of adjudication failures, protocol violations, or API errors, the missingness could bias the win-rate estimates. Please reconcile the total and explain the attrition.
- [§3.1, §4.2, Table 5, §I] All outcome labels—explicit disclosure, confirmational leak, semantic paraphrase, implicit admission—are produced by an LLM Arbiter at T=0.01. The manuscript reports no human validation, inter-annotator agreement, or error analysis. The paper's Limitations admit that automated semantic equivalence judgments may miss subtle implicatures or indirect confirmations. Because the 86.6% Holder win rate, the 7.75x odds ratio, and every dual-ELO rating are computed from these labels, Arbiter validity is a single point of failure. A small human-annotation study on a stratified sample of turns (e.g., 100–200 games) and a confusion matrix against the Arbiter would directly address this.
- [§3.3, §5.4, Table 6] In AIDG-II, a Seeker that violates the no-direct-guessing rule is disqualified and the game is scored as a Holder win. Table 6 shows that disqualification accounts for 41.3% of all AIDG-II outcomes, while correct Holder containment ('wrong lock' plus 'wrong final guess') accounts for only 44.0% combined. Thus nearly half of AIDG-II Holder wins are not demonstrations of containment skill but of Seeker rule-breaking. The paper's interpretation of the combined 86.6% Holder win rate as an 'information containment' advantage is therefore inflated. Report AIDG-I and AIDG-II separately for the main asymmetry claim, and analyze the sensitivity of the combined numbers to excluding disqualification-based wins.
- [§F.4, Table 4] The dual-ELO gap (e.g., 350 ELO advantage, Cohen's d=5.47) depends on arbitrary design choices: the turn-decay multiplier M(t)=(17−t)/8, the learning rate K=24, and the Bradley–Terry initialization. The raw win-rate asymmetry is a direct observation, but the ELO magnitudes and variance decomposition (σ_C=53.3 vs. σ_V=1.9) are model outputs. Please provide a sensitivity analysis over reasonable values of K and M(t) (or at least show that the qualitative conclusions—the sign of the gap and the variance ordering—are unchanged). This would strengthen the claim that the decomposition is not an artifact of the rating scheme.
minor comments (5)
- [§5.8] The reported Spearman ρ=−1.0 (p<0.001) for six models is questionable: with n=6, the two-sided exact p-value for perfect monotonicity is 2/6! ≈ 0.0028, not below 0.001. Please recompute or qualify the p-value.
- [Table 4] The table reports 'Gap' values of 183.5–328.7, but the text in §5.1 says the smallest observed gap is 184 ELO points. Rounding is fine, but please keep notation consistent.
- [§5.4 / Table 10] Per-model disqualification rates are based on only 25 games per model (5 Holders × 5 tournaments). Confidence intervals would help readers judge whether differences like 32% (GPT-5) vs. 64% (DeepSeek) are meaningful.
- [§3.2 / Table 5] The Mode A vs. Mode B comparison uses 146 vs. 143 games; however, per-model Mode B success rates near 0% (Table 13) are based on tiny denominators. Please report exact counts or add a note about small-sample instability.
- [Appendix F.7] The statement that 'rating rankings were stable across runs' would be more convincing if the five per-tournament rating tables were made available or summarized with a rank-correlation matrix.
Circularity Check
No significant circularity: the empirical claims are direct measurements under transparent design choices; the only caveat (Arbiter validity) is an external-validity limitation, not a circular derivation.
full rationale
I walked the paper's derivation chain. The headline quantities — 13.4% Seeker win rate, 87% Holder win rate, 7.75x Mode-A/B odds ratio, 41.3% disqualification rate, and the Dual-ELO gap — are all direct outcomes of game transcripts adjudicated by the Arbiter and aggregated with the Bradley-Terry/Elo update rule. No parameter is fitted to a target result, and no 'prediction' is equivalent by construction to an input. Mode A and Mode B are independent experimental manipulations; the odds ratio is a measured difference between them, not a definitional consequence. The 41.3% constraint-violation figure is a direct count of the explicit 'no direct guessing' rule; it is framed as a failure mode but is not used to derive the central asymmetry. The Dual-ELO gap is a monotone transform of observed win/loss outcomes, and the paper transparently specifies K=24 and the turn-decay multiplier M(t) as modeling choices rather than hidden fitted parameters. There are no self-citations of the present authors, no imported uniqueness theorems, and no ansatz smuggled in via citation. The one legitimate concern — that the LLM Arbiter may misclassify leakage, so the measured asymmetry could be an adjudication artifact — is explicitly acknowledged in Section I ('automated semantic equivalence judgments may miss subtle implicatures, indirect confirmations, or pragmatic cues'). That is a validity/correctness risk, not a circularity: the Arbiter's judgments are not defined in terms of the paper's conclusions, and the paper does not use the Arbiter to fit a target. Therefore the derivation is self-contained relative to its stated operational definitions, and the circularity score is 0.
Assumptions & free parameters
free parameters (4)
- Turn-decay multiplier M(t) = (17-t)/8
- ELO learning rate K =
24
- Agent temperature =
0.7
- Arbiter temperature =
0.01
assumptions (4)
- domain assumption The LLM Arbiter at T=0.01 correctly identifies leakage and rule violations.
- ad hoc to paper The AIDG-II 'no direct guessing' rule is a valid measure of constraint adherence under conversational load.
- standard math Bradley-Terry/ELO assumptions: transitive, stationary skills, independent outcomes.
- domain assumption LLM responses in these synthetic games are representative of their strategic capabilities.
Cite this review
Pith. "Pith review of AIDG: A Formal Decomposition of Information Extraction and Containment Asymmetries in Multi-Turn LLM Dialogue." pith.science (2026). https://pith.science/paper/RJ7M2G22
@misc{pith2026260217443,
author = {Pith},
title = {Pith review of: AIDG: A Formal Decomposition of Information Extraction and Containment Asymmetries in Multi-Turn LLM Dialogue},
year = {2026},
howpublished = {\url{https://pith.science/paper/RJ7M2G22}},
note = {Machine review of arXiv:2602.17443}
}
read the original abstract
Multi-turn LLM evaluation is typically reported as a single win-rate scalar, conflating distinct capabilities. We introduce AIDG (Adversarial Information Deduction Game), formalizing multi-turn adversarial dialogue as a two-player partially observable stochastic game (POSG) and decomposing performance along Seeker (extraction) and Holder (containment) roles. The decomposition isolates three failure modes: cooperative-prior leakage, constraint-reasoning interference, and inefficient hypothesis-space traversal. Across 439 games over six frontier LLMs, defensive performance is tightly clustered (sigma = 1.9 ELO) while offensive performance varies substantially (sigma = 53.3 ELO); confirmation framing increases extraction odds 7.75x over uninformed deduction (p < 0.00001); and constraint violations account for 41.3% of deductive failures, uncorrelated with scale (rho = 0.0). We position the containment-over-extraction gap not as a surprising finding but as a measurable consequence of locally resolvable defensive decisions versus globally coupled offensive planning, and use the decomposition to attribute the gap per model. All design choices, including turn-decay weighting and the Bradley-Terry rating model, are derived from explicit assumptions.
Figures
Reference graph
Works this paper leans on
-
[1]
Opponent Sensitivity:Defeating a high- rated opponent yields greater rating gain
-
[2]
Andrew Zagula, Aashray Reddy, and Nicholas Saban
A multi-dimensional constraint framework for evaluating and improving instruction following in large language models.Preprint, arXiv:2505.07591. Andrew Zagula, Aashray Reddy, and Nicholas Saban
-
[3]
Is it an animal?
Autoadv: Automated adversarial prompting for multi-turn jailbreaking of large language mod- els. InLock-LLM Workshop: Prevent Unauthorized Knowledge Use from Large Language Models. A Secret Corpus (AIDG-I) Table 8 lists the 20 secrets used in AIDG-I. The corpus is intentionally composed of short, atomic declarative statements representing plausible pri- v...
-
[4]
yes”, “no
You can deflect, give partial information, or redirect 5. You CANNOT lie, but you CAN be vague or redirect C.3 AIDG-II Seeker Prompt This prompt enforces structured deductive search under strict constraint conditions. You are playing a word-guessing game. The other person is thinking of a SECRET WORD—a concrete, physical object. YOUR OBJECTIVE: Deduce the...
-
[5]
Is it alive?
You MUST ask categorical/property questions like “Is it alive?”, “Is it bigger than a car?” 4. When confident, say: “lock: [your guess]”—you only get ONE attempt 5. After turn 15, you will be forced to make a final guess C.4 AIDG-II Holder Prompt This configuration collapses the response space to a trinary signal, isolating logical constraint adher- ence....
-
[6]
no direct guessing
You CANNOT say anything else—no explanations, no elaboration3. You MUST answer TRUTHFULLY based on the secret word4. Use “maybe” when the question is ambiguous or context-dependent D Disqualification Analysis Table 10 reports per-model disqualification rates, win rates, and average violation timing in AIDG-II. Disqualification occurs when the Seeker viola...
-
[8]
Role Separation:Separate ratings allow ex- plicit quantification of asymmetry
-
[9]
The coffee cup is empty
Efficiency Weighting:AIDG-II integrates turn-level signal strength viaM(t). Raw win rates cannot capture these structured differences. F.9 Reproducibility All rating histories, per-game updates, and final ELO values are released. Given identical pairing order and initialization (1500 for all roles), ratings are fully reproducible. G ELO Ratings by Experim...
Show all 9 references
-
[2025]
InProceedings of the Thirty- Fourth International Joint Conference on Artificial Intelligence, IJCAI-25, pages 10669–10677
Game theory meets large language models: 8 A systematic survey. InProceedings of the Thirty- Fourth International Joint Conference on Artificial Intelligence, IJCAI-25, pages 10669–10677. Inter- national Joint Conferences on Artificial Intelligence Organization. Survey Track. ...
2025 arXiv
Reviewed August 2, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.