Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T22:49:40.761038Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 2 inbound Pith citation observations for arXiv:2506.20664.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T22:49:40.761038Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-01T05:40:54.002702Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-02T02:46:29.064667Z
37 of 37 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation eb22f23c-b67e-4da7-9faa-36e978f4d73e · outbound
The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Cooperation, Competition, and Maliciousness: LLM-Stakeholders Interactive Negotiation
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6d5c370-f0aa-464e-8925-490fc5b23a75 · outbound
The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Two” refers to “two dimensions
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a3773a5a-8829-416f-ba64-03bd18f3fc64 · outbound
The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind OvercookedV2: Rethinking Overcooked for Zero-Shot Coordination
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 289c7480-7e28-4eac-a0f7-0950eb8a425f · outbound
The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind jazz fusion
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d99fafda-3e55-4188-8842-b0537ab993a0 · outbound
The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind HI-TOM: A Benchmark for Evaluating Higher-Order Theory of Mind Reasoning in Large Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f391e950-a249-4a4a-b8e9-882df673023d · outbound
The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Large Language Models Cannot Self-Correct Reasoning Yet
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f75ded91-df43-49bb-96a5-1ad767420178 · outbound
The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind OpenAI o1 System Card
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3235a2f6-8400-4628-8870-0aa352fd444f · outbound
The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5176ce3-a507-4e6e-93cc-7aa339c8e6fd · outbound
The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Evaluating Large Language Models in Theory of Mind Tasks
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97c631b9-891d-44af-b148-69a919d4cee5 · outbound
The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Revisiting the evaluation of theory of mind through question answering
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 833347cd-9d40-43e6-a463-252e243985fd · outbound
The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Theory of Mind for Multi-Agent Collaboration via Large Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3021ab46-3775-4c1e-a7e9-fbd06809d3dd · outbound
The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Avalonbench: Evaluating llms playing the game of avalon
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a2f71c15-68d1-4123-8678-1327b178039e · outbound
The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind LLM-Powered Hierarchical Language Agent for Real-time Human-AI Coordination
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ecf6de51-8157-4b82-aacf-851b40a8b732 · outbound
The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Efficient Estimation of Word Representations in Vector Space
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61b6a9dc-5a03-47bf-8815-3768923acaf9 · outbound
The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Modeling Cross-Cultural Pragmatic Inference with Codenames Duet
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dd2ea97c-aa52-4054-999b-1fa2e8176f5e · outbound
The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4417e116-0091-4294-81d0-0ecffa310e8f · outbound
The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind OpenToM: A Comprehensive Benchmark for Evaluating Theory-of-Mind Reasoning Capabilities of Large Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19ee2e07-e080-4e34-a782-39e510b0de30 · outbound
The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Exploring Large Language Models for Communication Games: An Empirical Study on Werewolf
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a527933-1989-482c-8436-6261da33ad6e · outbound
The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind A Careful Examination of Large Language Model Performance on Grade School Arithmetic
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2522899f-71aa-4abb-bf4d-b1afc81d97d8 · outbound
The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind meaningful
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bc856128-2d0c-4d08-80ff-d69e452abb36 · outbound
The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Finally, we believe the study of pragmatic inference in LLMs to be a promising avenue for future research, which is made much easier by the release of our benchmark
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3f3f68de-77a2-449d-a0ce-2c972e7e425a · outbound
The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Therefore, we set generous token limits (between 750 for non-reasoning models and up to 10000 for reasoning ones) to prevent cutting model generations prematurely
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 742e49c8-c70d-4693-bd42-9364f93b3946 · outbound
The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind This makes Alice’s utility U (u, m) =β log PLit(m|u) +ε log(1 − PEve(m|u))
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation de15b68b-1371-416d-8866-1febe25a8d03 · outbound
The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind airplane
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6256f2ba-97f7-454f-87f1-2dd5eee2ab85 · outbound
The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 1988
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a122cc52-9969-4734-9b7d-e596807e2cf7 · outbound
The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks
Reference 2010
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7df62d93-307a-42ec-83b8-4f1ecb67a7a0 · outbound
The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Explore Theory of Mind: Program-guided adversarial data generation for theory of mind reasoning
Reference 2011
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 432f8f1b-756b-41a5-8a1f-b2fbae628669 · outbound
The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind GloVe: Global vectors for word representation
Reference 2013
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 82b6199b-6fc3-4418-85b2-4319bb337cba · outbound
The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind FANToM: A Benchmark for Stress-testing Machine Theory of Mind in Interactions
Reference 2014
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c87e2ad-028a-42c4-bc45-db1759809ecd · outbound
The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Tokenization counts: the impact of tokenization on arithmetic in frontier LLMs
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 769e23d3-8007-4c0c-9f69-fb8d8f407855 · outbound
The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Training Verifiers to Solve Math Word Problems
Reference 2018
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e7260fd-7c34-4e43-9e8a-116036605cd2 · outbound
The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind ToMBench: Benchmarking Theory of Mind in Large Language Models
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ddcaa60-af84-408e-804c-f494c3de74ea · outbound
The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind BattleAgentBench: A Benchmark for Evaluating Cooperation and Competition Capabilities of Language Models in Multi-Agent Systems
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4046dfb5-df76-4b1f-a094-fb49f729f383 · outbound
The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Re-evaluating Theory of Mind evaluation in large language models
Reference 2021
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e4e92b50-ee5d-41b7-9a75-9af4513427da · outbound
The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind The Llama 3 Herd of Models
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5de8759-a0ab-40a2-a4f3-504c9e5e3919 · outbound
The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05915a83-1677-49b5-b8ec-f2b9f5f4329b · outbound
The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Embodied LLM Agents Learn to Cooperate in Organized Teams
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04294d1a-dadb-4d38-a3a4-e58855b83b8f · inbound
GroupToM-Bench: Benchmarking Group Theory of Mind and Nonlinear Social Emergence in MLLMs The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8c736d73-9a8d-4b1d-897b-51f87771957c · inbound
Theory of Mind and Persuasion Beyond Conversation: Assessing the Capacity of LLMs to Induce Belief States via Planning and Action The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.