{"id":"047c2dea-db4a-4215-b707-817b978562f4","arxiv_id":"2411.09168","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":4,"one_line_summary":"Artificial agents that model other agents' beliefs, preferences, and constraints can, in principle, raise the collective intelligence of hybrid human-AI groups, though the paper's quantitative illustration is not correctly derived.","lead":"This paper argues that humans work better in groups when they can infer each other's mental states, and predicts that AI systems with a similar mental-modeling ability will boost human-AI team performance. It reviews supporting evidence and offers a small toy model, but the toy model's key 1-bit calculation does not follow from the paper's own formula, so the demonstration is not reliable.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Section 4.2's 1-bit gain is not phi: with iid st, all time-delayed mutual informations vanish, so the toy model does not demonstrate ToM-enhancement under Eq. 1.","rationale":"Read in good faith, the paper is a perspective with a conditional central claim: because human ToM predicts group performance, AIs with ToM may improve hybrid collective intelligence. That claim does not require a proof, but the paper chooses to offer one in Section 4, and the strength of that offering depends entirely on the phi calculation in Eq. 1. My stress-test confirms the reader's weakest assumption: the 1-bit value in Section 4.2 is not a time-delayed excess mutual information under the paper's own definition. With an iid common signal, there is no predictive information from the past, so phi must be zero; the 1 bit is simply the entropy of the signal entering all agents simultaneously. This is a correctness risk rather than a stylistic objection. I also checked whether the piKL section or the human ToM review could carry the central claim independently. The human-ToM-to-group-performance evidence (Woolley et al.) is real and makes the hypothesis plausible, but it does not establish the AI extension; the piKL architecture is a sketch with no experiments and is presented as a formal interpretation, not a validation. Thus the load-bearing quantitative support is the toy model, and it is miscalculated. The right result is not rejection: the perspective is broad, readable, and the hypothesis is reasonable, but acceptance should be conditional on recomputing or replacing Section 4.2 and on empirical testing of AI ToM in human-AI groups. That is exactly the reader's verdict, so I recommend UNCHANGED.","tokens_in":29501,"tokens_out":9266,"duration_ms":105782,"concrete_test":"Analytically re-derive Eq. 1 for the Section 4.2 second scenario at tau=1 with st iid uniform: since x1_t = st/4 and (x2_t,x3_t) are the Nash outcome of the resulting game, X_t depends on t only through st, and st is independent of s_{t-1}; hence the joint TDMI and all individual TDMIs are zero and phi = 0. Also run a JIDT simulation of the stated dynamics for T >= 10^5 at tau = 1, 2, 3 and report the three terms of Eq. 1. If the reported 1-bit value can only be reproduced by setting tau = 0 or by reporting H(st), then Section 4.2's demonstration is invalid and must be replaced with a signal that has temporal structure (e.g., a Markov st) before the claim can be supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 4.2 is the only quantitative support for the central hypothesis, and it does not compute the paper's own measure. In the second scenario, A2 and A3 are at their Nash equilibrium and their actions are a deterministic function of A1's action, which in turn is the iid signal st: x1_t = st/4 and, depending on the sign of x1, (x2_t,x3_t) = (1,1) or (-1,-1). Because st is independent across time, X_t = (x1_t,x2_t,x3_t) is a memoryless function of st. Therefore for every tau > 0, I(X_t; X_{t-tau}) = 0 and each individual I(x_i,t; x_i,t-tau) = 0, so phi(X;tau) from Eq. 1 is 0, not 1 bit. The reported 1 bit is the Shannon entropy of the injected signal, H(st) = H(x2_t) = H(x3_t), or equivalently the tau=0 total correlation of x2 and x3; it is instantaneous common input, not emergent time-delayed information exchange. The model also assumes A1 knows st, the game, and can set the preference parameter c = x1; any purely reactive controller with that information would create the same correlation, so even if the arithmetic were fixed, the example would not isolate a ToM mechanism. Appendix B confirms the utilities used, but no tau or simulation is reported for Eq. 1 in Section 4.2.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper argues that human collective intelligence (CI) is substantially enabled by Theory of Mind (ToM) at the individual level, and it hypothesizes that ToM-equipped AI agents will likewise enhance the collective intelligence of human-AI groups. The argument is built on a broad interdisciplinary review: liquid/solid brains, social network topology, the BPC model of game theory, and psychological evidence such as Woolley et al.'s finding that group performance correlates with ToM. The paper introduces a time-delayed mutual information measure phi(X;tau) (Eq. 1) as a quantitative proxy for CI, applies it to a monkey-computer game dataset, and then presents a three-agent toy model in Section 4 in which an agent A1 uses ToM to manipulate the payoffs of two other agents, claiming that phi increases from 0 to 1 bit. A separate Section 5 sketches an RL/ToM architecture based on piKL regularisation.","tokens_in":29806,"tokens_out":3674,"duration_ms":114939,"significance":"If the central hypothesis is accepted, the paper offers a useful conceptual bridge between social psychology, complex systems theory, and AI design, and it gives a concrete, if minimal, information-theoretic target for measuring emergent CI. The review component is valuable and well grounded, particularly in the empirical ToM literature and in the established liquid/solid brain framework. The paper is ambitious and interdisciplinary, and it explicitly aims to reframe AI as agential participants in a social ecology rather than as tools. However, the single quantitative demonstration of the paper's own measure, Section 4.2, does not compute phi as defined, so the numerical support for the central claim is currently absent. The manuscript's value is therefore mainly in its synthesis and hypothesis formulation rather than in its proof-of-concept model.","major_comments":[{"comment":"The claimed value of 1 bit for the second scenario is inconsistent with Eq. (1). In that scenario x1_t = st/4 and, because A2 and A3 are at their Nash equilibrium, (x2_t,x3_t) = (st,st). Hence X_t is a memoryless function of the iid signal st. For every tau > 0, I(X_t; X_{t-tau}) = 0 and each single-agent TDMI I(x_i,t; x_i,t-tau) = 0, so phi(X;tau) = 0 by the paper's own definition. The reported 1 bit corresponds to H(st), or equivalently to the instantaneous tau = 0 correlation between x2 and x3; it is injected environmental information, not time-delayed excess mutual information. No value of tau or simulation procedure is reported for this computation, so the reader cannot verify any alternative reading. This undermines the only quantitative illustration supporting the central hypothesis.","section":"Section 4.2, Eq. (1)"},{"comment":"Even if the arithmetic of phi were corrected, the toy model would not isolate a ToM mechanism. A1 sets the preference parameter c equal to x1 = st/4 and thereby selects, at each time step, whether the A2-A3 game is the Prisoner's Dilemma or the Harmony game. Any reactive controller that observes st and knows the payoff structure would produce exactly the same joint state process; no inference of hidden beliefs, preferences, or constraints is required. To support the ToM claim, the model needs either a control condition in which the same information is transmitted without BPC modelling or a condition in which A1 must infer the hidden BPC parameters from observed behaviour. As it stands, the 1-bit gain is built into the scenario by construction rather than demonstrated as an emergent effect of ToM.","section":"Section 4.1, Eqs. (8)-(9)"}],"minor_comments":[{"comment":"The expression for mutual information is written as a sum over i = 1 to n, but the summands are not defined and the notation conflates the joint-state sum with the agent index; it should be written as a sum over joint configurations with the appropriate probabilities.","section":"Eq. (2)"},{"comment":"The text states 'for the second scenario it is 1 bit' without specifying the time delay tau, the sample length, or the estimator used; since Eq. (1) depends critically on tau, this omission is important even beyond the conceptual issue raised above.","section":"Section 4.2"},{"comment":"The piKL discussion is presented as a formal framework, but it is not connected to the preceding phi-based analysis and is never used in any simulation or quantitative result; the authors should clarify whether it is intended as a proposal for future work or as part of the present argument.","section":"Section 5"},{"comment":"The manuscript would benefit from a careful proofread: examples include 'discpline' (Section 1.2), 'n the following section' (Section 3.2), 'Battencourt' for Bettencourt (Eq. 1 citation), 'conspicifics' for conspecifics (Section 1.1), and 'understating' for understanding (Section 2.2).","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The quantitative error in Section 4.2 is central to the paper's self-presentation as a demonstration, but the broader synthesis and the literature review have independent value. The authors could revise the toy model to genuinely exhibit time-delayed information flow or, alternatively, reframe the paper as a perspective that explicitly avoids quantitative proof. Because the current claim is load-bearing and not merely cosmetic, I recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper reads like a good review stuck to a faulty demo. The review of how human ToM supports collective intelligence, drawing on Woolley et al. and the liquid/solid brain literature, is competent and worth reading. The central hypothesis—that AIs with ToM could similarly boost hybrid human-AI teams—is plausible and already in the literature. But the only quantitative support, the toy model in Section 4.2, misapplies the paper's own measure. With an iid signal and all three agents deterministic functions of the current st, every time-delayed mutual information in Eq. 1 is zero; phi is 0, not 1 bit. The reported 1 bit is just the entropy of the injected signal, an instantaneous common cause. This is not a small slip: it is the one place the paper uses its measure to back its argument.\n\nCredit where due: the monkey-computer example in Section 3 computes phi on real data and gets plausible nonzero values, so the measure itself is fine. The review of ToM orders, BPC models, and niche construction is broad and mostly accurate. The piKL architecture in Section 5, though, is a sketch with no validation.\n\nThe hypothesis is not new; the authors cite Westby and Riedl (2023), Mao et al. (2024), and others making similar points. What is new is the attempted information-theoretic illustration, and that illustration fails as written.\n\nIf the toy model is corrected or removed and the claims are scaled back to the review evidence, this could be a useful perspective piece. As it stands, the load-bearing example is wrong. I'd send it to review because the review material deserves referee time, and the error is fixable, but I'd flag Section 4.2 prominently and expect major revision. I wouldn't cite it in my own work until the demonstration is fixed.\n\nReading group? Maybe, as a cautionary example of how easily a time-delayed mutual information measure can be misapplied. Not a must-read.","headline":"The review is solid but the toy model miscomputes its own measure: the claimed 1-bit CI is just the injected signal's entropy, so the central demonstration fails.","tokens_in":30369,"tokens_out":7092,"would_cite":false,"duration_ms":62771,"reading_group":"maybe","serious_thinker":"no","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This review argues that equipping AI agents with a Theory of Mind—the ability to infer others' beliefs, preferences, and constraints—will let them enhance collective intelligence in human groups in the same way people do.","keywords":["collective intelligence","theory of mind","human-AI collaboration","time-delayed mutual information","game theory of mind","beliefs preferences constraints","social niche construction","multi-agent systems"],"falsifier":"Recompute $\\phi$ in the Section 4 model from the stated assumptions: $s_t$ independent and uniformly random, $x^1_t=s_t/4$, and $x^2_t,x^3_t$ fixed at Nash choices; then each time-delayed mutual information in Equation (1) is zero whenever the agent variables are deterministic functions of the signal or constants, so the joint term equals the self-predictability of the exogenous signal, giving $\\phi=0$. If that calculation is right, the paper's one-bit illustration collapses. A separate empirical falsifier: in a controlled experiment, compare group-level $c$ or $\\phi$ for teams with a ToM-equipped AI mediator versus a generic AI support tool; no difference in group-level intelligence scores would refute the central hypothesis.","tokens_in":29267,"feed_emoji":"🧠","tokens_out":10335,"duration_ms":100676,"temperature":0.7,"pith_summary":"This paper argues that the same psychological capacity that raises human group performance—Theory of Mind, the ability to infer other agents' beliefs, preferences, and constraints—should be built into artificial agents, and that such AIs will then enhance collective intelligence in hybrid human–AI groups much as people do. It reviews evidence that groups perform better when members score highly on Theory-of-Mind measures, and it interprets that result through a game-theoretic model of hidden mental states. To make the claim quantitative, it proposes a measure of collective intelligence, the excess time-delayed mutual information $\\phi$, and illustrates with a three-agent model in which one ToM-equipped agent reconfigures the incentives of two others, raising $\\phi$ from zero to one bit. The wider point is that agential AI should be seen as inhabitants of a social ecology—choosing, conforming to, and constructing socio-cognitive niches—rather than as tools. If the claim holds, it gives designers a concrete target: AIs that read and rewrite interpersonal incentive structures should measurably lift group information processing.","feed_headline":"AI theory of mind could lift collective intelligence, review argues","feed_subtitle":"Agents that model others' beliefs and preferences could boost group performance the way human teammates do, a review proposes.","key_machinery":"The load-bearing object is $\\phi(X;\\tau)=I(X_t;X_{t-\\tau})-\\sum_i I(X^i_t;X^i_{t-\\tau})$, the excess time-delayed mutual information: how much predictive information the joint state of all agents carries beyond the sum of each agent's own predictive information. The paper uses this as a proxy for collective intelligence, in the spirit of but not identical to integrated information theory. The complementary machinery is the BPC model (beliefs, preferences, and constraints), which turns a Theory of Mind into a tractable game-theoretic inference: an agent models the hidden utility co-factors that drive others' choices. In the minimal model, agent A1 knows the BPC of agents A2 and A3 and sets $x_1 = \\frac{1}{4}s_t$, switching their game between Prisoner's Dilemma and Harmony so that cooperation occurs exactly when the environmental signal says it is valuable; the paper reports $\\phi=0$ before the intervention and $\\phi=1$ bit after, illustrating how one ToM agent can restructure an interaction network at short time scales.","core_discovery":"The paper's central claim is that Theory of Mind is the individual-level mechanism that lets agents take causal, goal-directed control of collective information processing, and that AIs equipped with it will enhance collective intelligence in ways similar to human contributions. To support this, the paper distinguishes zeroth- to fourth-order ToM, associates human social intelligence with second-order ToM (attributing hidden beliefs, preferences, and constraints to others), and links the collective-intelligence factor $c$ of groups to members' ToM capability. It then formalises ToM-agent interaction through game-theoretic utilities in which one agent shifts another pair's effective game from Prisoner's Dilemma to Harmony, creating a hypergraph of interaction that the paper measures as one bit of $\\phi$. The intended upshot is that hybrid human–AI collectives will become more intelligent not because AI computes better, but because AI can participate in the social construction of shared goals.","pith_inferences":["Editorial inference: the BPC-based definition implies a sharp boundary test—an AI that only imitates observable behaviour should not raise collective intelligence as much as one that infers and alters hidden preferences, because the latter changes the effective game rather than merely predicting moves.","Editorial inference: the framing places LLM-based assistants and ToM-equipped agents on different rungs; if so, benchmarks for human–AI collaboration should measure group-level $\\phi$ or $c$, not the standalone capability of the model.","Editorial inference: the same $\\phi$ measurement could be applied directly to human team time series, offering a way to test whether ToM training or team composition shifts collective intelligence independently of individual IQ.","Editorial inference: viewed as a design brief, the paper suggests that alignment of ToM-AI should be evaluated at the level of the collective's information processing, since the incentives an AI manipulates become part of the group's effective game."],"forward_implications":["If the claim is correct, the design target for AI in human groups shifts from better tools to agential social actors that infer and modify the beliefs, preferences, and constraints of human teammates.","The $\\phi$ measure offers a computable objective for such agents: an AI contribution counts as intelligence-enhancing to the extent it raises the excess time-delayed mutual information of the whole human–AI system.","The framework predicts that ToM-equipped AI will matter most in fluid, reconfigurable social settings—the 'liquid brain' end of the spectrum—where fast, targeted interventions in who interacts with whom matter.","Hybrid human–AI teams containing a ToM-equipped agent should display higher group-level collective intelligence (the $c$ factor) than teams using AI purely as an information-processing tool.","AI will need the full repertoire of niche choice, niche conformance, and niche construction—selecting, adapting to, and reshaping its social niche—rather than being deployed into fixed contexts."],"supporting_citations":[{"why":"Supplies the definition of Theory of Mind as inference of others' knowledge, beliefs, and desires that the whole argument builds on.","marker":"[51]"},{"why":"Provides the key empirical result: a group-level collective-intelligence factor $c$ that is predicted by members' Theory-of-Mind performance.","marker":"[54]"},{"why":"Extends the ToM–collective-intelligence link to online and face-to-face groups, strengthening the empirical base.","marker":"[55]"},{"why":"Supplies the Beliefs-Preferences-Constraints model used to formalise ToM as game-theoretic inference.","marker":"[57]"},{"why":"Provides the 'Game Theory of Mind' recursive reasoning framework the paper adapts.","marker":"[56]"},{"why":"Demonstrates an AI reaching human-level play in a social-deduction game by modelling other players' beliefs and goals, supporting the feasibility of ToM-equipped agents.","marker":"[44]"},{"why":"Shows an algorithm cooperating with humans and machines at human-comparable levels through cheap talk and adaptable strategies, supporting AI agency in mixed groups.","marker":"[41]"},{"why":"Provides the integrated-information formalism from which Equation (1)'s $\\phi$ definition is taken.","marker":"[32]"},{"why":"The software toolkit used to compute the time-delayed mutual information values reported in Table 1.","marker":"[112]"}],"fun_headline_variants":["AI with theory of mind can boost group intelligence","Theory of mind in AI may enhance collective smarts","AI that models minds could improve team performance","How AI with theory of mind lifts collective intelligence"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that the $\\phi$ measure is a valid and correctly computed proxy for collective intelligence; if in the three-agent toy model the reported one bit is just the entropy of the external signal, because the two Nash-locked agents contribute no time-delayed mutual information, then the quantitative illustration does not actually show that Theory of Mind increases collective intelligence.","fun_headline_variants_meta":{"raw":{"variants":["AI with theory of mind can boost group intelligence","Theory of mind in AI may enhance collective smarts","AI that models minds could improve team performance","How AI with theory of mind lifts collective intelligence"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000116,"raw_usage":{"total_tokens":1108,"prompt_tokens":1011,"completion_tokens":97,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":627,"completion_tokens_details":{"reasoning_tokens":38}},"tokens_in":627,"tokens_out":97,"duration_ms":2324,"temperature":1.0,"reasoning_tokens":38,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T20:57:50.038526+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute $\\phi$ in the Section 4 model from the stated assumptions: $s_t$ independent and uniformly random, $x^1_t=s_t/4$, and $x^2_t,x^3_t$ fixed at Nash choices; then each time-delayed mutual information in Equation (1) is zero whenever the agent variables are deterministic functions of the signal or constants, so the joint term equals the self-predictability of the exogenous signal, giving $\\phi=0$. If that calculation is right, the paper's one-bit illustration collapses. A separate empirical falsifier: in a controlled experiment, compare group-level $c$ or $\\phi$ for teams with a ToM-equipped AI mediator versus a generic AI support tool; no difference in group-level intelligence scores would refute the central hypothesis.","supporting_citations":[{"cited_title":"The rules of information aggregation and emergence of collective intelligent behavior","cited_arxiv_id":null,"evidence_quote":"The software toolkit used to compute the time-delayed mutual information values reported in Table 1."}],"review_version":1}