Pith. sign in

REVIEW 2 major objections 4 minor 65 references

Agent Behavioral Contracts II: Certifying Compositional Reliability Without Assuming Independence

T0 review · 2 major / 4 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read Two instances of one model co-fail on 90.0% of missions where either fails, so the independence assumption behind compositional reliability bounds gives way, and the paper responds with a certificate that assumes no dependence structure.

desk verdict A serious, unusually transparent empirical and theoretical package on correlated failures in multi-agent systems; the 90% co-failure finding is solid modulo one legitimate scorer-validity concern, and the moment-set certificate is a real contribution. read the letter →

arxiv 2608.12895 v1 pith:JWJI2PM6 submitted 2026-08-13 cs.AI cs.MA

classification cs.AIcs.MA MSC 62N0590C0562F4060G42
keywords multi-agentreliabilitycorrelatedfailureconditionalindependencecompositionalboundsredundancyover-creditmoment-setcertificatelinearprogramminganytime-validinference
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Compositional reliability bounds for multi-agent systems multiply component reliabilities, a step licensed by a conditional-independence assumption that is routinely stated and rarely tested. This paper tests it and finds it false where it matters most: two instances of one model, composed in a two-agent handoff, co-fail on 90.0% of the missions on which either fails, in a preregistered evaluation of 18,000 missions scored by deterministic code with no LLM judge. The error is signed and runs against the operator: positive dependence inflates joint failure above the independence product, so redundant paths are over-credited exactly when their components share a model. Dropping the assumption leaves a vacuous bound, and fitting a dependence model is provably worse, so the paper constructs a finite-sample certificate that assumes no dependence structure: a linear program over the joint distribution constrained by measured co-execution moments, sound and sharp for the information supplied, and monotone in the moment family. A careful reader should care because standard compositional certification is quietly wrong in the redundancy case, and this paper replaces it with a guarantee that degrades honestly.

What carries the argument

Two objects carry the argument. Diagnostically, the signed compositional gap: for two components, the true joint failure probability exceeds the independence product by exactly $Cov(h_1, h_2)$, the covariance of the two hard verdicts, so positive dependence always over-credits redundancy and the direction of the error is fixed without further estimation. Constructively, the moment-set certificate: a linear program over the $2^m$ cells of the joint law on $\{0,1\}^m$, minimizing the all-success probability subject to the constraint that measured co-execution moments — single-stage successes, pairwise co-successes, and optionally triple co-successes — lie inside a Bonferroni–Clopper–Pearson box around their empirical values; the box covers by the union bound, the true law is feasible whenever it covers, and the LP minimum is the certified floor. A companion betting e-process, $E_R = \prod_{r \le R}\bigl(1 + \lambda_r (y_r - p_0)\bigr)$, gives an anytime-valid certificate whose null constrains only a conditional mean, so it needs no independence assumption at all.

What would settle it

Rescore a random sample of the confirmatory missions with a contract set and scoring code written independently of the paper's authors, and compare the same-model co-failure rate of 90.0% and the log-odds contrasts; if the co-failure pattern moves materially the headline is a scoring artifact, while reproduction would confirm genuine model dependence. A sharper variant runs the substitution with two models matched on marginal failure rate but from different lineages, separating shared competence from shared inductive bias.

Watch

Extended reading notes

Core claim

The paper's claim is that the conditional-independence condition C5 of compositional contract theory fails for same-model composition, that the failure is large and signed, and that it can be certified around without any dependence assumption. In the preregistered confirmatory arm, two instances of mistral-small-24b in a two-agent handoff co-fail on 2177 of the 2418 missions on which either fails — 90.0%, against the 14.6% the independence product predicts — with log odds ratio 6.66 (95% CI [6.38, 7.00]) and $\phi = 0.916$. Substituting a different model into the second agent reduces the association significantly in six of six contrasts across three topologies; substituting a different vendor, with the model already different, does not, a registered hypothesis reported as a null. The paper further shows that the assumption-free Fréchet–Hoeffding sandwich, which brackets a joint probability from marginals alone, is vacuous — its certified floor is zero whenever mean component reliability falls below $1 - 1/m$ — and that a certificate built on a fitted dependence model loses coverage of the truth as the sample grows, because the identification gap is $O(1)$ while the bootstrap haircut is $O(n^{-1/2})$. Its remedy is a linear program over the joint law on $\{0,1\}^m$, minimized over all distributions whose co-execution moments lie in a Bonferroni–Clopper–Pearson box: valid with no dependence assumption, sharp for the moments supplied, and monotone in the moment family, and on four-stage data enriching ten moment functionals to fourteen narrows the identified interval by 85.7% and lifts the certified floor from 0.2455 to 0.4116.

Load-bearing premise

The headline result stands on the deterministic gold scoring code measuring contract compliance correctly, even though the same authors wrote both the contracts and the scorer; a systematic scoring bias would make the co-failure counts and arm contrasts artifacts rather than true model dependence.

Editorial extensions

If this is right

  • At the measured dependence the independence product over-credits redundancy by a wide margin: the two agents would co-fail on 14.6% of missions if independent and actually co-fail on 36.3%, so a dashboard that multiplies reliabilities reports a reassuring number whose governing assumption the data reject.
  • Model identity, not vendor identity, is the axis on which to choose redundancy: substituting a different model reduced the association significantly in six of six contrasts across three topologies, while substituting a different vendor, with the model already different, produced no consistent reduction.
  • A certificate built on a fitted dependence model is worse than no certificate: its bootstrap lower bound loses coverage of the true reliability as the sample grows, so more data narrows an interval around the wrong target with no visible symptom.
  • Operators can tighten the certified floor without any dependence assumption by measuring more co-execution moments: on four-stage data, enriching ten moment functionals to fourteen narrowed the identified interval by 85.7% and lifted the floor from 0.2455 to 0.4116, and under pre-allocated Bonferroni spending the tightening is monotone by construction.
  • Marginal-bounded dependence statistics — Jaccard, $\phi$, Kendall's $\tau_a$ — can reverse the apparent ordering of conditions when the compared agents fail at different rates, so a marginal-free statistic such as the log odds ratio should be reported alongside them.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the co-failure rate transfers beyond the retail and financial task domains tested, any safety case built on same-model redundant review is effectively relying on a single point of failure: a reviewer and a writer drawn from one model supply far less independent evidence than the design assumes, and the honest redundancy count may be one rather than two.
  • The coverage-collapse theorem is not specific to agents: any bootstrap interval built on a misspecified parametric model with a fixed identification gap will, past some sample size, sit entirely off the truth, which cautions against model-based certificates in copula-based risk and reliability practice generally.
  • Because the anytime-valid e-process needs no independence assumption, it is the one certificate in the paper that survives the failure the paper documents; a natural deployment pattern is a continuously re-earned certificate that a team may watch and stop on at bounded type-I cost.
  • A testable extension suggested by the paper's own limitation section: run the substitution design with two models matched on marginal failure rate but from different lineages, which would separate the contribution of shared competence from shared inductive bias to the co-failure signal.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 4 minor

Summary. The paper tests the conditional-independence condition C5 that licenses multiplicative compositional reliability bounds for multi-agent systems. In a preregistered campaign of 18,000 two-agent handoff missions scored by deterministic code, two instances of mistral-small-24b co-fail on 90.0% of missions on which either fails (log OR = 6.66, 95% CI [6.38, 7.00]); substituting a different model significantly reduces the association in six of six contrasts, while substituting a different vendor, with the model already different, does not. The paper then develops a finite-sample reliability certificate based on a linear program over measured co-execution moments with a Bonferroni-Clopper-Pearson box, proves it sound and sharp, adds an anytime-valid e-process certificate, and proves that bootstrap bounds on a fitted dependence model lose coverage of the true reliability as the sample grows. The evaluation also reports a negative control, cross-backend replication, and an ablation of the i.i.d. assumption. The paper is unusually transparent about its own limitations, including the post-hoc choice of the log odds ratio for H2 and the repository-based rather than external preregistration timestamp.

Significance. If the empirical finding holds, it is significant: it provides controlled evidence that the conditional-independence assumption fails for same-model composition, with a signed, practically important consequence for redundant agent designs. The theoretical machinery is a genuine contribution: the moment-set LP certificate of Theorem 5.2 is a sound finite-sample, copula-agnostic bound, sharp for the supplied moments, and the coverage-collapse result of Theorem 4.2 is a useful warning about fitted-dependence certificates. The paper ships complete proofs, released code, deterministic scoring, and regenerable statistics, which are strengths. The anytime-valid certificate is standard machinery but is applied cleanly and with careful empirical checking. The main risks are empirical rather than theoretical: the headline co-failure numbers depend entirely on the authors' own deterministic scorer, and the substitution manipulation confounds model identity with model capability; both are acknowledged but not resolved.

major comments (2)
  1. [Sections 10.1, 10.2.1, 11.3] The headline co-failure estimate (J = 0.9003, log OR = 6.66) and the arm contrasts of Table 3 are produced entirely by the authors' deterministic gold scorer. Because mission identifiers embed the condition name and each arm draws its own missions, a scorer rule that misfires on particular mission or output types can inflate co-failure counts in one arm more than another, so the Section 11.3 defense that the same contracts and scorer are used in all arms does not fully secure the differential claim. The internal negative control of Table 4 holds the compared pair fixed and therefore cannot validate the scorer's behavior across different model identities. An independent scorer or human rescoring of a stratified subsample is needed before the 90% co-failure finding is relied on.
  2. [Sections 10.1, 10.2.2, 11.3] The manipulation confounds model identity with model capability: the same_vendor substitution uses ministral-8b, a weaker model, and the different_vendor arm uses gemma-3-12b-it. The six significant same_model-versus-substitution contrasts, and the Section 10.9 conclusion that the operative variable is the model not the vendor, could in part be driven by capability-correlated failure modes rather than by shared model weights. The marginal-free statistics remove the effect of different marginal failure rates but not the effect of capability-correlated failure modes, and the negative control does not address this because the control pair is never substituted. A design with models matched on failure rate, or a direct manipulation of shared weights, is required to support the causal attribution in the abstract.
minor comments (4)
  1. [Equation (25), Definition 3.17] The quantity 2(p11p00 − p10p01) is labelled Kendall's τa, but the conventional sample tau-a for binary variables is a factor of 2 larger (approximately 4(p11p00 − p10p01) with a finite-population correction). Please state the scaling convention explicitly, since the registered H1 threshold is stated on τa.
  2. [Abstract and Section 10.2.2] The phrase 'a registered hypothesis reported as a null' for the same_vendor versus different_vendor comparison should be qualified: the registered H2 was a three-level ordering on τa, and the vendor-level comparison is a post-hoc decomposition of that ordering. The paper's Section 11.3 disclosure is accurate, but the abstract's wording invites the reading that the vendor null was itself preregistered.
  3. [Section C.4] The preregistration timestamp rests on repository history rather than on an external registry. The paper discloses this, but given the confirmatory role of the registration, an independent timestamp or third-party registration would materially strengthen the claim.
  4. [Table 3 caption] The verdict label 'conflict' in the SV−DV rows is not defined in the caption; please state that it means the marginal-sensitive and marginal-free statistics disagree in sign or significance, as explained in the text.

Circularity Check

0 steps flagged · score 0.0 of 10

No circular derivation: the empirical co-failure result is a direct preregistered measurement and the certificate theorems are proved from stated moment constraints.

full rationale

I walked the derivation chain from the measured co-failure tables (Table 2) to the headline 90.0% overlap, the arm contrasts (Table 3), and the moment-set and anytime-valid certificates (Theorems 5.2, 6.1, 7.1). The co-failure percentage is arithmetic on n11/(n11+n10+n01) from logs scored by deterministic code; it is not a fitted parameter renamed as a prediction. The LP certificate minimizes all-success over a Clopper-Pearson moment box and is proved sound; sharpness and monotonicity are proven, and the Bonferroni allocation that makes monotonicity hold by construction is explicitly disclosed (Proposition 6.2). Theorem 4.2's coverage-collapse result is a delta>0 misspecification argument, and E3's witness is explicitly constructed as the LP minimizer that is moment-indistinguishable from the Gaussian law; this is a valid illustration of the theorem, not an independent empirical prediction, and the paper says so. The only self-citation is to the v1 paper for the ABC framework definitions and carried-forward single-agent evidence (Sections 2.1, 10.8); the framework formulas are restated in full in Section 3, and the v1 empirical evidence is not used to derive any new result, so it is not load-bearing circularity. The paper itself flags the main validity threats: the gold scorer shares an author with the contracts (Section 11.3), the H2 estimator was chosen after seeing the reversal (Section 11.3), the preregistration timestamp rests on repository history rather than an external registry (Section C.4), and the Meta breadth arm suffered outcome-dependent attrition (Section 10.6). These are correctness and validity concerns, not circularity: none of them makes a stated prediction equal to an input by construction.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The central claims rest on standard probability tools, a preregistered empirical design, and the assumption that the deterministic scorer faithfully measures contract compliance. No new physical entities are introduced.

assumptions (4)
  • domain assumption Missions in each arm are i.i.d. draws from the generator distribution.
    The Tier-1 certificate (Theorem 5.2) and all bootstrap intervals assume i.i.d. missions. Section 10.7 tests and finds small design effects (DEFF up to 1.148), but the proof requires the assumption.
  • domain assumption The deterministic gold scoring code correctly evaluates contract compliance.
    Section 10.1 states scoring is by deterministic gold code; Section 11.3 notes the scorer shares an author with the contracts. The co-failure measurements depend on the scorer's correctness.
  • domain assumption The v1 ABC framework (Definitions 3.1-3.13) is taken as given, including the (p, delta, k)-satisfaction notion and the drift dynamics.
    The paper builds on the author's earlier framework and does not re-test single-agent contract enforcement; Section 10.8 carries those results forward.
  • domain assumption Composition rules (series, quorum, parallel) are deterministic; aggregator nodes are code, not models.
    Section 11.2 states that a model-based aggregator would introduce a third correlated component not covered by the analysis.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Agent Behavioral Contracts II: Certifying Compositional Reliability Without Assuming Independence." pith.science (2026). https://pith.science/paper/JWJI2PM6

@misc{pith2026260812895,
  author       = {Pith},
  title        = {Pith review of: Agent Behavioral Contracts II: Certifying Compositional Reliability Without Assuming Independence},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/JWJI2PM6}},
  note         = {Machine review of arXiv:2608.12895}
}
read the original abstract

Compositional reliability bounds for multi-agent systems multiply component reliabilities, a step licensed by a conditional-independence assumption that is routinely stated and rarely tested. We test it. Two instances of one model, in a two-agent handoff, co-fail on 90.0% of the missions on which either fails (log OR 6.66, 95% CI [6.38, 7.00]; phi 0.916), in a preregistered evaluation of 18,000 missions scored by deterministic code with no LLM judge. Substituting a different model reduces the association in six of six contrasts; substituting a different vendor, model already different, does not -- a registered hypothesis reported as a null. The error is signed and runs against the operator: positive dependence inflates joint failure above the independence product, so redundancy is over-credited exactly when components share a model. The assumption-free alternative is often vacuous, and fitting a dependence model is worse: we prove a bootstrap bound on a fitted model's functional loses coverage of the truth as n grows, the identification gap being O(1) while the bootstrap haircut is O(n^{-1/2}). More data makes such a certificate worse, with no visible symptom. We give a finite-sample certificate assuming no dependence structure: a linear program over the joint, over a Bonferroni-Clopper-Pearson box around measured co-execution moments. It is sound, sharp for the information supplied, and monotone in the moment family. Enriching ten moment functionals to fourteen narrows the identified interval by 85.7% and lifts the certified floor from 0.2455 to 0.4116. A companion anytime-valid certificate holds type-I error at 0.0471 under optional stopping. Common dependence statistics are marginal-bounded and can reverse an apparent ordering of conditions when the compared agents fail at different rates. Contracts, scoring code, analysis scripts, and the preregistration are released.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

65 extracted references · 50 canonical work pages

  1. [1]

    Constitutional AI : Harmlessness from AI feedback

    Yuntao Bai, Saurav Kadavath, Sandipan Kundu, et al. Constitutional AI : Harmlessness from AI feedback. arXiv preprint arXiv:2212.08073, 2022

  2. [2]

    Barlow and Frank Proschan

    Richard E. Barlow and Frank Proschan. Statistical Theory of Reliability and Life Testing. Holt, Rinehart and Winston, 1975

  3. [3]

    Rustan M

    Mike Barnett, K. Rustan M. Leino, and Wolfram Schulte. The Spec\# programming system: An overview. In CASSIS, pages 49--69, 2004

  4. [4]

    International AI safety report 2026

    Yoshua Bengio et al. International AI safety report 2026. arXiv preprint arXiv:2602.21012, 2026

  5. [5]

    Can AI agents agree? arXiv preprint arXiv:2603.01213, 2026

    Fr\'ed\'eric Berdoz, Leonardo Rugli, and Roger Wattenhofer. Can AI agents agree? arXiv preprint arXiv:2603.01213, 2026

  6. [6]

    Optimal inequalities in probability theory: A convex optimization approach

    Dimitris Bertsimas and Ioana Popescu. Optimal inequalities in probability theory: A convex optimization approach. SIAM Journal on Optimization, 15 0 (3): 0 780--804, 2005

  7. [7]

    Agent behavioral contracts: Formal specification and runtime enforcement for reliable autonomous AI agents

    Varun Pratap Bhardwaj. Agent behavioral contracts: Formal specification and runtime enforcement for reliable autonomous AI agents. arXiv preprint arXiv:2602.22302, 2026

  8. [8]

    Bonferroni

    Carlo E. Bonferroni. Teoria statistica delle classi e calcolo delle probabilit\`a. Pubblicazioni del R. Istituto Superiore di Scienze Economiche e Commerciali di Firenze, 8: 0 3--62, 1936

Show all 65 references
  1. [9]

    An Investigation of the Laws of Thought

    George Boole. An Investigation of the Laws of Thought. Walton and Maberly, 1854

  2. [10]

    LangChain

    Harrison Chase. LangChain . https://github.com/langchain-ai/langchain, 2022

  3. [11]

    Clarke, Orna Grumberg, and Doron A

    Edmund M. Clarke, Orna Grumberg, and Doron A. Peled. Model Checking. MIT Press, 1999

  4. [12]

    C. J. Clopper and E. S. Pearson. The use of confidence or fiducial limits illustrated in the case of the binomial. Biometrika, 26 0 (4): 0 404--413, 1934

  5. [13]

    Abstract interpretation: A unified lattice model for static analysis of programs by construction or approximation of fixpoints

    Patrick Cousot and Radhia Cousot. Abstract interpretation: A unified lattice model for static analysis of programs by construction or approximation of fixpoints. In POPL, pages 238--252, 1977

  6. [14]

    Tenenbaum, and Igor Mordatch

    Yilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum, and Igor Mordatch. Improving factuality and reasoning in language models through multiagent debate. In ICML, 2024

  7. [15]

    Bootstrap methods: Another look at the jackknife

    Bradley Efron. Bootstrap methods: Another look at the jackknife. The Annals of Statistics, 7 0 (1): 0 1--26, 1979

  8. [16]

    Endres and Johannes E

    Dominik M. Endres and Johannes E. Schindelin. A new metric for probability distributions. IEEE Transactions on Information Theory, 49 0 (7): 0 1858--1860, 2003

  9. [17]

    Ernst, Jeff H

    Michael D. Ernst, Jeff H. Perkins, Philip J. Guo, Stephen McCamant, Carlos Pacheco, Matthew S. Tschantz, and Chen Xiao. The Daikon system for dynamic detection of likely invariants. Science of Computer Programming, 69 0 (1--3): 0 35--45, 2007

  10. [18]

    Sur les tableaux de corr\'elation dont les marges sont donn\'ees

    Maurice Fr\'echet. Sur les tableaux de corr\'elation dont les marges sont donn\'ees. Annales de l'Universit\'e de Lyon, Section A, 14: 0 53--77, 1951

  11. [19]

    The statistical crisis in science

    Andrew Gelman and Eric Loken. The statistical crisis in science. American Scientist, 102 0 (6): 0 460--465, 2014

  12. [20]

    Safe testing

    Peter Gr\"unwald, Rianne de Heide, and Wouter Koolen. Safe testing. Journal of the Royal Statistical Society B, 86 0 (5): 0 1091--1128, 2024

  13. [21]

    Best possible inequalities for the probability of a logical function of events

    Theodore Hailperin. Best possible inequalities for the probability of a logical function of events. The American Mathematical Monthly, 72 0 (4): 0 343--359, 1965

  14. [22]

    Multi-agent risks from advanced AI

    Lewis Hammond et al. Multi-agent risks from advanced AI . arXiv preprint arXiv:2502.14143, 2025

  15. [23]

    C. A. R. Hoare. An axiomatic basis for computer programming. Communications of the ACM, 12 0 (10): 0 576--580, 1969

  16. [24]

    ur Angewandte Mathematik der Universit\

    Wassily Hoeffding. Ma stabinvariante K orrelationstheorie. Schriften des Mathematischen Instituts und des Instituts f\"ur Angewandte Mathematik der Universit\"at Berlin, 5: 0 179--233, 1940

  17. [25]

    Probability inequalities for sums of bounded random variables

    Wassily Hoeffding. Probability inequalities for sums of bounded random variables. Journal of the American Statistical Association, 58 0 (301): 0 13--30, 1963

  18. [26]

    MetaGPT : Meta programming for a multi-agent collaborative framework

    Sirui Hong, Mingchen Zhuge, Jonathan Chen, et al. MetaGPT : Meta programming for a multi-agent collaborative framework. In ICLR, 2024

  19. [27]

    Howard, Aaditya Ramdas, Jon McAuliffe, and Jasjeet Sekhon

    Steven R. Howard, Aaditya Ramdas, Jon McAuliffe, and Jasjeet Sekhon. Time-uniform, nonparametric, nonasymptotic confidence sequences. The Annals of Statistics, 49 0 (2): 0 1055--1080, 2021

  20. [28]

    Counterfactual graph for multi-agent LLM calibration

    Jiatan Huang, Mingchen Li, Ziming Li, Sunjae Kwon, Hong Yu, and Chuxu Zhang. Counterfactual graph for multi-agent LLM calibration. arXiv preprint arXiv:2605.30653, 2026

  21. [29]

    The distribution of the flora in the alpine zone

    Paul Jaccard. The distribution of the flora in the alpine zone. New Phytologist, 11 0 (2): 0 37--50, 1912

  22. [30]

    Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan

    Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan. SWE-bench : Can language models resolve real-world GitHub issues? In ICLR, 2024

  23. [31]

    Maurice G. Kendall. A new measure of rank correlation. Biometrika, 30 0 (1--2): 0 81--93, 1938

  24. [32]

    Survey Sampling

    Leslie Kish. Survey Sampling. John Wiley and Sons, New York, 1965

  25. [33]

    Specifying Systems: The TLA+ Language and Tools for Hardware and Software Engineers

    Leslie Lamport. Specifying Systems: The TLA+ Language and Tools for Hardware and Software Engineers . Addison-Wesley, 2002

  26. [34]

    Rustan M

    K. Rustan M. Leino. Dafny: An automatic program verifier for functional correctness. In LPAR, pages 348--370, 2010

  27. [35]

    Encouraging divergent thinking in large language models through multi-agent debate

    Tian Liang, Zhiwei He, Wenxiang Jiao, et al. Encouraging divergent thinking in large language models through multi-agent debate. In EMNLP, 2024

  28. [36]

    Divergence measures based on the S hannon entropy

    Jianhua Lin. Divergence measures based on the S hannon entropy. IEEE Transactions on Information Theory, 37 0 (1): 0 145--151, 1991

  29. [37]

    AgentBench : Evaluating LLMs as agents

    Xiao Liu, Hao Yu, Hanchen Zhang, et al. AgentBench : Evaluating LLMs as agents. In ICLR, 2024

  30. [38]

    Shay Seiya McDonnell, Avantika Singh, Quoc-Viet Pham, Vratislav Havlik, and Gregory M. P. O'Hare. Harnessing disagreement: Detecting correlated agreement blindness in multi-agent triage. arXiv preprint arXiv:2607.19899, 2026. Accepted, PAAMS 2026

  31. [39]

    Applying ``design by contract''

    Bertrand Meyer. Applying ``design by contract''. Computer, 25 0 (10): 0 40--51, 1992

  32. [40]

    Roger B. Nelsen. An Introduction to Copulas. Springer, 2nd edition, 2006

  33. [41]

    Nosek, Charles R

    Brian A. Nosek, Charles R. Ebersole, Alexander C. DeHaven, and David T. Mellor. The preregistration revolution. Proceedings of the National Academy of Sciences, 115 0 (11): 0 2600--2606, 2018

  34. [42]

    Training language models to follow instructions with human feedback

    Long Ouyang, Jeffrey Wu, Xu Jiang, et al. Training language models to follow instructions with human feedback. In NeurIPS, 2022

  35. [43]

    O'Brien, Carrie J

    Joon Sung Park, Joseph C. O'Brien, Carrie J. Cai, Meredith Ringel Morris, Percy Liang, and Michael S. Bernstein. Generative agents: Interactive simulacra of human behavior. In UIST, 2023

  36. [44]

    ChatDev : Communicative agents for software development

    Chen Qian, Wei Liu, Hongzhang Liu, et al. ChatDev : Communicative agents for software development. In ACL, 2024

  37. [45]

    VerifyMAS : Hypothesis verification for failure attribution in LLM multi-agent systems

    Hezhe Qiao, Hanghang Tong, Ee-Peng Lim, Bing Liu, and Guansong Pang. VerifyMAS : Hypothesis verification for failure attribution in LLM multi-agent systems. arXiv preprint arXiv:2605.17467, 2026

  38. [46]

    Governed capability evolution: Lifecycle-time compatibility checking and rollback for AI -component-based systems

    Xue Qin, Simin Luan, John See, Zeyd Boukhers, Cong Yang, and Zhijun Li. Governed capability evolution: Lifecycle-time compatibility checking and rollback for AI -component-based systems. arXiv preprint arXiv:2604.08059, 2026

  39. [47]

    FALAT : Tracing failures in LLM agent trajectories via dependency-guided search

    Md Nakhla Rafi, Md Ahasanuzzaman, Dong Jae Kim, Zhijie Wang, and Tse-Hsun Chen. FALAT : Tracing failures in LLM agent trajectories via dependency-guided search. arXiv preprint arXiv:2606.00765, 2026

  40. [48]

    Game-theoretic statistics and safe anytime-valid inference

    Aaditya Ramdas, Peter Grünwald, Vladimir Vovk, and Glenn Shafer. Game-theoretic statistics and safe anytime-valid inference. Statistical Science, 38 0 (4): 0 576--601, 2023

  41. [49]

    NeMo guardrails: A toolkit for controllable and safe LLM applications with programmable rails

    Traian Rebedea, Razvan Dinu, Makesh Sreedhar, Christopher Parisien, and Jonathan Cohen. NeMo guardrails: A toolkit for controllable and safe LLM applications with programmable rails. In EMNLP System Demonstrations, 2023

  42. [50]

    Statistical methods related to the law of the iterated logarithm

    Herbert Robbins. Statistical methods related to the law of the iterated logarithm. The Annals of Mathematical Statistics, 41 0 (5): 0 1397--1409, 1970

  43. [51]

    Testing by betting: A strategy for statistical and scientific communication

    Glenn Shafer. Testing by betting: A strategy for statistical and scientific communication. Journal of the Royal Statistical Society A, 184 0 (2): 0 407--431, 2021

  44. [52]

    Game-Theoretic Foundations for Probability and Finance

    Glenn Shafer and Vladimir Vovk. Game-Theoretic Foundations for Probability and Finance. Wiley, 2019

  45. [53]

    Reflexion : Language agents with verbal reinforcement learning

    Noah Shinn, Federico Cassano, Edward Berman, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao. Reflexion : Language agents with verbal reinforcement learning. In NeurIPS, 2023

  46. [54]

    Simmons, Leif D

    Joseph P. Simmons, Leif D. Nelson, and Uri Simonsohn. False-positive psychology: Undisclosed flexibility in data collection and analysis allows presenting anything as significant. Psychological Science, 22 0 (11): 0 1359--1366, 2011

  47. [55]

    Fonctions de r\'epartition \`a n dimensions et leurs marges

    Abe Sklar. Fonctions de r\'epartition \`a n dimensions et leurs marges. Publications de l'Institut de Statistique de l'Universit\'e de Paris, 8: 0 229--231, 1959

  48. [56]

    The one-sided barrier problem for G aussian noise

    David Slepian. The one-sided barrier problem for G aussian noise. Bell System Technical Journal, 41 0 (2): 0 463--501, 1962

  49. [57]

    \'Etude critique de la notion de collectif

    Jean Ville. \'Etude critique de la notion de collectif. Gauthier-Villars, Paris, 1939

  50. [58]

    Sequential tests of statistical hypotheses

    Abraham Wald. Sequential tests of statistical hypotheses. The Annals of Mathematical Statistics, 16 0 (2): 0 117--186, 1945

  51. [59]

    Self-consistency improves chain of thought reasoning in language models

    Xuezhi Wang, Jason Wei, Dale Schuurmans, et al. Self-consistency improves chain of thought reasoning in language models. In ICLR, 2023

  52. [60]

    Estimating means of bounded random variables by betting

    Ian Waudby-Smith and Aaditya Ramdas. Estimating means of bounded random variables by betting. Journal of the Royal Statistical Society B, 86 0 (1): 0 1--27, 2024

  53. [61]

    AutoGen : Enabling next-gen LLM applications via multi-agent conversation

    Qingyun Wu, Gagan Bansal, Jieyu Zhang, et al. AutoGen : Enabling next-gen LLM applications via multi-agent conversation. In COLM, 2024

  54. [62]

    ReAct : Synergizing reasoning and acting in language models

    Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. ReAct : Synergizing reasoning and acting in language models. In ICLR, 2023

  55. [63]

    Udny Yule

    G. Udny Yule. On the association of attributes in statistics: With illustrations from the material of the childhood society. Philosophical Transactions of the Royal Society A, 194: 0 257--319, 1900

  56. [64]

    Rethinking the reliability of multi-agent system: A perspective from Byzantine fault tolerance

    Lifan Zheng, Jiawei Chen, Qinghong Yin, Jingyuan Zhang, Xinyi Zeng, and Yu Tian. Rethinking the reliability of multi-agent system: A perspective from Byzantine fault tolerance. arXiv preprint arXiv:2511.10400, 2025

  57. [65]

    Xu, Hao Zhu, et al

    Shuyan Zhou, Frank F. Xu, Hao Zhu, et al. WebArena : A realistic web environment for building autonomous agents. In ICLR, 2024

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.