Pith. sign in

Paper Citation Record · LEDGER

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind

As of 7 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 2 inbound Pith citation observations for arXiv:2506.20664.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.20664 v1

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:49:40.761038Z

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-01T05:40:54.002702Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T02:46:29.064667Z

Reference resolution

37 of 37 outbound references displayed

  • verified exact2
  • verified fuzzy9
  • unresolved24
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation eb22f23c-b67e-4da7-9faa-36e978f4d73e · outbound

This paper cites Cooperation, Competition, and Maliciousness: LLM-Stakeholders Interactive Negotiation.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Cooperation, Competition, and Maliciousness: LLM-Stakeholders Interactive Negotiation

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:37.942223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:37.942223Z digest=sha256:5070cfd24917be5e12962491f4a568cdf39953d58574c660b9464f54a5df294f

Observation d6d5c370-f0aa-464e-8925-490fc5b23a75 · outbound

This paper cites Two” refers to “two dimensions.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Two” refers to “two dimensions

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:42.287485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:49:40.238004Z digest=sha256:15d1a3bee9f3b675e44c104f45ae55e511db8d9f1fccc9335a26b5120a393f74

Observation a3773a5a-8829-416f-ba64-03bd18f3fc64 · outbound

This paper cites OvercookedV2: Rethinking Overcooked for Zero-Shot Coordination.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind OvercookedV2: Rethinking Overcooked for Zero-Shot Coordination

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:38.289724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:38.289724Z digest=sha256:befed8f061880126592e528b5131f3af3c2d7e3b47aa0ffa5a9dbac5932164de

Observation 289c7480-7e28-4eac-a0f7-0950eb8a425f · outbound

This paper cites jazz fusion.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind jazz fusion

Reference 8

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T22:49:42.045426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:49:40.452292Z digest=sha256:036add810c06b0bd604c905bc4fe53fb2288f1e2d11ac342af2151dbbffabdc1

Observation d99fafda-3e55-4188-8842-b0537ab993a0 · outbound

This paper cites HI-TOM: A Benchmark for Evaluating Higher-Order Theory of Mind Reasoning in Large Language Models.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind HI-TOM: A Benchmark for Evaluating Higher-Order Theory of Mind Reasoning in Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:38.482038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:38.482038Z digest=sha256:e1c6583349ac4a69aa6106a869c2119c63b2d14e81af24064c3fd39d71ff16f9

Observation f391e950-a249-4a4a-b8e9-882df673023d · outbound

This paper cites Large Language Models Cannot Self-Correct Reasoning Yet.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Large Language Models Cannot Self-Correct Reasoning Yet

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:38.615228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:38.615228Z digest=sha256:685b41a022ef3433bb531b404bbc0b3b4bda92e03c70566bb4738b29b58f403c

Observation f75ded91-df43-49bb-96a5-1ad767420178 · outbound

This paper cites OpenAI o1 System Card.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind OpenAI o1 System Card

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:38.679533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:38.679533Z digest=sha256:4676a102f68c94557e664906bf15fcb0d77928d068c3fba40a916eb2b20119ea

Observation 3235a2f6-8400-4628-8870-0aa352fd444f · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:38.741025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:38.741025Z digest=sha256:d44e32f579fecd74bee19b7fb33b159059f4a64d9409c7eb1899dec1d1ac5d9a

Observation b5176ce3-a507-4e6e-93cc-7aa339c8e6fd · outbound

This paper cites Evaluating Large Language Models in Theory of Mind Tasks.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Evaluating Large Language Models in Theory of Mind Tasks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:38.859398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:38.859398Z digest=sha256:3ca45e5e3f85f3fb4b096667b2468002eee2a58bce7d8c78a5317dd3771bbbb2

Observation 97c631b9-891d-44af-b148-69a919d4cee5 · outbound

This paper cites Revisiting the evaluation of theory of mind through question answering.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Revisiting the evaluation of theory of mind through question answering

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:42.342665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:49:38.925060Z digest=sha256:b8dd5372d83d7e3cc3fb60287006f641447296df5ede64676c99e81dab3c9b8b

Observation 833347cd-9d40-43e6-a463-252e243985fd · outbound

This paper cites Theory of Mind for Multi-Agent Collaboration via Large Language Models.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Theory of Mind for Multi-Agent Collaboration via Large Language Models

Reference 17

Resolution
malformed identifier
no resolver link, observed 2026-08-06T22:49:38.986397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:38.986397Z digest=sha256:90445dede203ed27e57570d80271d55afd13a978c5ae2e0c83d5e0014a0c15ae

Observation 3021ab46-3775-4c1e-a7e9-fbd06809d3dd · outbound

This paper cites Avalonbench: Evaluating llms playing the game of avalon.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Avalonbench: Evaluating llms playing the game of avalon

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:42.331673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:49:39.053674Z digest=sha256:dc50e2a20f9d69ac68dfdd655839374af1a579ae2e16840e04ed7d98e48ffded

Observation a2f71c15-68d1-4123-8678-1327b178039e · outbound

This paper cites LLM-Powered Hierarchical Language Agent for Real-time Human-AI Coordination.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind LLM-Powered Hierarchical Language Agent for Real-time Human-AI Coordination

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:39.116735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:39.116735Z digest=sha256:dec2185060c434e7140b4cf5234e9eee9ed408c5246e7ed678b1979f145774fd

Observation ecf6de51-8157-4b82-aacf-851b40a8b732 · outbound

This paper cites Efficient Estimation of Word Representations in Vector Space.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Efficient Estimation of Word Representations in Vector Space

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:39.189471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:39.189471Z digest=sha256:b2a1765b74872864e86c3c539b9917eefaef5efc07048bc44ee44e1648ed2be3

Observation 61b6a9dc-5a03-47bf-8815-3768923acaf9 · outbound

This paper cites Modeling Cross-Cultural Pragmatic Inference with Codenames Duet.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Modeling Cross-Cultural Pragmatic Inference with Codenames Duet

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:49:41.096852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:49:39.355554Z digest=sha256:845decd967cd6c281d4f3f65ec5e52c61a5fac898720072ab59d1009848f1544

Observation dd2ea97c-aa52-4054-999b-1fa2e8176f5e · outbound

This paper cites Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:39.387074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:39.387074Z digest=sha256:eb5d982fc15d50cc8e0edf5faa871675d3f90baa2ec2e5a8bb40c90e326f1072

Observation 4417e116-0091-4294-81d0-0ecffa310e8f · outbound

This paper cites OpenToM: A Comprehensive Benchmark for Evaluating Theory-of-Mind Reasoning Capabilities of Large Language Models.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind OpenToM: A Comprehensive Benchmark for Evaluating Theory-of-Mind Reasoning Capabilities of Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:39.672207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:39.672207Z digest=sha256:e32552c0d9dd90471869e4a30b53fcbfb1dab632f2086205a458f2d0bfe47b8e

Observation 19ee2e07-e080-4e34-a782-39e510b0de30 · outbound

This paper cites Exploring Large Language Models for Communication Games: An Empirical Study on Werewolf.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Exploring Large Language Models for Communication Games: An Empirical Study on Werewolf

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:39.847874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:39.847874Z digest=sha256:d1e44c31d855113e1ea79c990f5b4fc5fe8fa742bf1a43fbba3818e99e11166c

Observation 2a527933-1989-482c-8436-6261da33ad6e · outbound

This paper cites A Careful Examination of Large Language Model Performance on Grade School Arithmetic.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind A Careful Examination of Large Language Model Performance on Grade School Arithmetic

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:39.921570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:39.921570Z digest=sha256:2383685bba9f063838a61982bed2493d25678091f8562492deb3594134043c2d

Observation 2522899f-71aa-4abb-bf4d-b1afc81d97d8 · outbound

This paper cites meaningful.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind meaningful

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:42.311032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:49:40.034128Z digest=sha256:19338c22529a29836eaad84024222ecf8a1b6750155165f6b686799c5b18c565

Observation bc856128-2d0c-4d08-80ff-d69e452abb36 · outbound

This paper cites Finally, we believe the study of pragmatic inference in LLMs to be a promising avenue for future research, which is made much easier by the release of our benchmark.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Finally, we believe the study of pragmatic inference in LLMs to be a promising avenue for future research, which is made much easier by the release of our benchmark

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:42.299944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:49:40.149128Z digest=sha256:c5c7840d878e3fd87d592027935b7d12fded7536b88e7e1dcff7be006c2dbaf3

Observation 3f3f68de-77a2-449d-a0ce-2c972e7e425a · outbound

This paper cites Therefore, we set generous token limits (between 750 for non-reasoning models and up to 10000 for reasoning ones) to prevent cutting model generations prematurely.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Therefore, we set generous token limits (between 750 for non-reasoning models and up to 10000 for reasoning ones) to prevent cutting model generations prematurely

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:42.218772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:49:40.300874Z digest=sha256:38561be6ba8b885f78ebef5bce1e03dabf8885d04d5a3c80dafb38297e1eaefc

Observation 742e49c8-c70d-4693-bd42-9364f93b3946 · outbound

This paper cites This makes Alice’s utility U (u, m) =β log PLit(m|u) +ε log(1 − PEve(m|u)).

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind This makes Alice’s utility U (u, m) =β log PLit(m|u) +ε log(1 − PEve(m|u))

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:41.972989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:49:40.662165Z digest=sha256:acd277edaf09c8dc8cbf7b5cefa5ac13f216b8bda30b88dc8b7b4da08514fad8

Observation de15b68b-1371-416d-8866-1febe25a8d03 · outbound

This paper cites airplane.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind airplane

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:41.731822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:49:40.761038Z digest=sha256:fc6f45b4f2ac5a12d417abe2fae1d922c133afea287cf8a2be539f158f9131a1

Observation 6256f2ba-97f7-454f-87f1-2dd5eee2ab85 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 1988

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:38.360777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:38.360777Z digest=sha256:b650756e3b9a3dd85b7e0cfe713bec1633d4390bd9e214c7f91aac55c07d880d

Observation a122cc52-9969-4734-9b7d-e596807e2cf7 · outbound

This paper cites Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 2010

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:39.394970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:39.394970Z digest=sha256:1c8acad44d3487076dae84ec20c309ac33ee442e35abcf963ff937cbff5f4b21

Observation 7df62d93-307a-42ec-83b8-4f1ecb67a7a0 · outbound

This paper cites Explore Theory of Mind: Program-guided adversarial data generation for theory of mind reasoning.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Explore Theory of Mind: Program-guided adversarial data generation for theory of mind reasoning

Reference 2011

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:39.274909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:39.274909Z digest=sha256:8c760994b579ec48dfcad64067d68c1b4bc303e061af7e5dc97db1ce1301988e

Observation 432f8f1b-756b-41a5-8a1f-b2fbae628669 · outbound

This paper cites GloVe: Global vectors for word representation.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind GloVe: Global vectors for word representation

Reference 2013

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:42.321600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:49:39.230130Z digest=sha256:b49a43e96492f0b206c429887b6f2da6c3c443d460f356e5ea24198a910ca69d

Observation 82b6199b-6fc3-4418-85b2-4319bb337cba · outbound

This paper cites FANToM: A Benchmark for Stress-testing Machine Theory of Mind in Interactions.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind FANToM: A Benchmark for Stress-testing Machine Theory of Mind in Interactions

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:38.805574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:38.805574Z digest=sha256:1a6d83355b9e3080cb1c35cf1c7fc8bf2715989589aaf662febec3ba6c379609

Observation 9c87e2ad-028a-42c4-bc45-db1759809ecd · outbound

This paper cites Tokenization counts: the impact of tokenization on arithmetic in frontier LLMs.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Tokenization counts: the impact of tokenization on arithmetic in frontier LLMs

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:39.391218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:39.391218Z digest=sha256:ad3a852b5c0b2d5c7c0d7afd7253d567b0dfcf215f4871bd548ce90e3321b452

Observation 769e23d3-8007-4c0c-9f69-fb8d8f407855 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Training Verifiers to Solve Math Word Problems

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:38.149003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:38.149003Z digest=sha256:de36b4260c9511808fe019813d0774d864761ab0376222246b72b9c97835aa09

Observation 4e7260fd-7c34-4e43-9e8a-116036605cd2 · outbound

This paper cites ToMBench: Benchmarking Theory of Mind in Large Language Models.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind ToMBench: Benchmarking Theory of Mind in Large Language Models

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:37.997496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:37.997496Z digest=sha256:e9ad2ce973a0aec73a08fbb8617d08c7469a87534b6921709429cf77ec54040f

Observation 5ddcaa60-af84-408e-804c-f494c3de74ea · outbound

This paper cites BattleAgentBench: A Benchmark for Evaluating Cooperation and Competition Capabilities of Language Models in Multi-Agent Systems.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind BattleAgentBench: A Benchmark for Evaluating Cooperation and Competition Capabilities of Language Models in Multi-Agent Systems

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:39.475102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:39.475102Z digest=sha256:887f79903ca87901ddad7517751e10d7857eefff7bb99d849d9db0b06ff09d61

Observation 4046dfb5-df76-4b1f-a094-fb49f729f383 · outbound

This paper cites Re-evaluating Theory of Mind evaluation in large language models.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Re-evaluating Theory of Mind evaluation in large language models

Reference 2021

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:49:41.405397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:49:38.555674Z digest=sha256:fac72007bce835d90f7e501a4cc6bf989da1c0e4fdde3e3931fd70de9cd8a5d4

Observation e4e92b50-ee5d-41b7-9a75-9af4513427da · outbound

This paper cites The Llama 3 Herd of Models.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind The Llama 3 Herd of Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:38.231197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:38.231197Z digest=sha256:3814e5de2c4b73a36eb5df3b4a7949c94b5d81fb8fa7a2086a2f0d46c44aae24

Observation b5de8759-a0ab-40a2-a4f3-504c9e5e3919 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:38.076422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:38.076422Z digest=sha256:49a94ccb9b5f701fb7e79b4eebcbc72ae8e87869bdfadb5856221b9f00d498e5

Observation 05915a83-1677-49b5-b8ec-f2b9f5f4329b · outbound

This paper cites Embodied LLM Agents Learn to Cooperate in Organized Teams.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Embodied LLM Agents Learn to Cooperate in Organized Teams

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:38.421893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:38.421893Z digest=sha256:2e20d31ba853044e28d6d3b56b19faf359f5c8bccd6d2014a46bde5989782153

Pith citing papers

Observation 04294d1a-dadb-4d38-a3a4-e58855b83b8f · inbound

GroupToM-Bench: Benchmarking Group Theory of Mind and Nonlinear Social Emergence in MLLMs cites this paper.

GroupToM-Bench: Benchmarking Group Theory of Mind and Nonlinear Social Emergence in MLLMs The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:46:29.066605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T10:37:45.062718Z digest=sha256:395487f990861f85d73852054ab8934b37fef61f237a20a3d07adcccbf5b8c87

Observation 8c736d73-9a8d-4b1d-897b-51f87771957c · inbound

Theory of Mind and Persuasion Beyond Conversation: Assessing the Capacity of LLMs to Induce Belief States via Planning and Action cites this paper.

Theory of Mind and Persuasion Beyond Conversation: Assessing the Capacity of LLMs to Induce Belief States via Planning and Action The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T10:15:44.730832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-07-01T05:40:54.002702Z digest=sha256:7d62d737b62c9eb894e4108b5bbb2d1e89db3781d25a5c21936c8b632203b7b3