{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2024:H2F2R2KIBALL3EXMINQXZC73LD","short_pith_number":"pith:H2F2R2KI","schema_version":"1.0","canonical_sha256":"3e8ba8e9480816bd92ec43617c8bfb58c4ee2ea2748115f1b302b6867b87ce63","source":{"kind":"arxiv","id":"2402.15052","version":2},"attestation_state":"computed","paper":{"title":"ToMBench: Benchmarking Theory of Mind in Large Language Models","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.AI"],"primary_cat":"cs.CL","authors_text":"Bosi Wen, Gongyao Jiang, Guanqun Bi, Jincenzi Wu, Jinfeng Zhou, Mengting Hu, Minlie Huang, Yaru Cao, Yunghwei Lai, Zexuan Xiong, Zhuang Chen","submitted_at":"2024-02-23T02:05:46Z","abstract_excerpt":"Theory of Mind (ToM) is the cognitive capability to perceive and ascribe mental states to oneself and others. Recent research has sparked a debate over whether large language models (LLMs) exhibit a form of ToM. However, existing ToM evaluations are hindered by challenges such as constrained scope, subjective judgment, and unintended contamination, yielding inadequate assessments. To address this gap, we introduce ToMBench with three key characteristics: a systematic evaluation framework encompassing 8 tasks and 31 abilities in social cognition, a multiple-choice question format to support aut"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2402.15052","kind":"arxiv","version":2},"metadata":{"license":"http://creativecommons.org/licenses/by/4.0/","primary_cat":"cs.CL","submitted_at":"2024-02-23T02:05:46Z","cross_cats_sorted":["cs.AI"],"title_canon_sha256":"4998f904c344b87d47219fead92f19282b8b00514954fa0a070c98e2297b942c","abstract_canon_sha256":"868bb85a17c6c9afc858e505759cd39ddeee4bd012c51e130df4064997d7a0de"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T09:45:44.894776Z","signature_b64":"nNGtxaB7izDPXIIxAkH5RfUBUcUx25yuDb7PbhG6xJXqYW25DwyMAZMm7diC8grkcCzZMkTjIgLtP16T2Y7dBg==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"3e8ba8e9480816bd92ec43617c8bfb58c4ee2ea2748115f1b302b6867b87ce63","last_reissued_at":"2026-07-05T09:45:44.894291Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T09:45:44.894291Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"ToMBench: Benchmarking Theory of Mind in Large Language Models","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.AI"],"primary_cat":"cs.CL","authors_text":"Bosi Wen, Gongyao Jiang, Guanqun Bi, Jincenzi Wu, Jinfeng Zhou, Mengting Hu, Minlie Huang, Yaru Cao, Yunghwei Lai, Zexuan Xiong, Zhuang Chen","submitted_at":"2024-02-23T02:05:46Z","abstract_excerpt":"Theory of Mind (ToM) is the cognitive capability to perceive and ascribe mental states to oneself and others. Recent research has sparked a debate over whether large language models (LLMs) exhibit a form of ToM. However, existing ToM evaluations are hindered by challenges such as constrained scope, subjective judgment, and unintended contamination, yielding inadequate assessments. To address this gap, we introduce ToMBench with three key characteristics: a systematic evaluation framework encompassing 8 tasks and 31 abilities in social cognition, a multiple-choice question format to support aut"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2402.15052","kind":"arxiv","version":2},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2402.15052/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2402.15052","created_at":"2026-07-05T09:45:44.894362+00:00"},{"alias_kind":"arxiv_version","alias_value":"2402.15052v2","created_at":"2026-07-05T09:45:44.894362+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2402.15052","created_at":"2026-07-05T09:45:44.894362+00:00"},{"alias_kind":"pith_short_12","alias_value":"H2F2R2KIBALL","created_at":"2026-07-05T09:45:44.894362+00:00"},{"alias_kind":"pith_short_16","alias_value":"H2F2R2KIBALL3EXM","created_at":"2026-07-05T09:45:44.894362+00:00"},{"alias_kind":"pith_short_8","alias_value":"H2F2R2KI","created_at":"2026-07-05T09:45:44.894362+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":7,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.24267","citing_title":"Pigeonholing: how bad prompts hurt models, causing collapse and mistakes","ref_index":1,"is_internal_anchor":false},{"citing_arxiv_id":"2606.21315","citing_title":"Social World Model for Lifelong Social Intelligence","ref_index":2,"is_internal_anchor":false},{"citing_arxiv_id":"2605.20506","citing_title":"Reinforcing Human Behavior Simulation via Verbal Feedback","ref_index":156,"is_internal_anchor":false},{"citing_arxiv_id":"2605.20423","citing_title":"OSCToM: RL-Guided Adversarial Generation for High-Order Theory of Mind","ref_index":18,"is_internal_anchor":false},{"citing_arxiv_id":"2604.20749","citing_title":"Where and What: Reasoning Dynamic and Implicit Preferences in Situated Conversational Recommendation","ref_index":75,"is_internal_anchor":false},{"citing_arxiv_id":"2604.17989","citing_title":"AIT Academy: Cultivating the Complete Agent with a Confucian Three-Domain Curriculum","ref_index":22,"is_internal_anchor":false},{"citing_arxiv_id":"2605.03149","citing_title":"Are you with me? A Framework for Detecting Mental Model Discrepancies in Task-Based Team Dialogues","ref_index":6,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/H2F2R2KIBALL3EXMINQXZC73LD","json":"https://pith.science/pith/H2F2R2KIBALL3EXMINQXZC73LD.json","graph_json":"https://pith.science/api/pith-number/H2F2R2KIBALL3EXMINQXZC73LD/graph.json","events_json":"https://pith.science/api/pith-number/H2F2R2KIBALL3EXMINQXZC73LD/events.json","paper":"https://pith.science/paper/H2F2R2KI"},"agent_actions":{"view_html":"https://pith.science/pith/H2F2R2KIBALL3EXMINQXZC73LD","download_json":"https://pith.science/pith/H2F2R2KIBALL3EXMINQXZC73LD.json","view_paper":"https://pith.science/paper/H2F2R2KI","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2402.15052&json=true","fetch_graph":"https://pith.science/api/pith-number/H2F2R2KIBALL3EXMINQXZC73LD/graph.json","fetch_events":"https://pith.science/api/pith-number/H2F2R2KIBALL3EXMINQXZC73LD/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/H2F2R2KIBALL3EXMINQXZC73LD/action/timestamp_anchor","attest_storage":"https://pith.science/pith/H2F2R2KIBALL3EXMINQXZC73LD/action/storage_attestation","attest_author":"https://pith.science/pith/H2F2R2KIBALL3EXMINQXZC73LD/action/author_attestation","sign_citation":"https://pith.science/pith/H2F2R2KIBALL3EXMINQXZC73LD/action/citation_signature","submit_replication":"https://pith.science/pith/H2F2R2KIBALL3EXMINQXZC73LD/action/replication_record"}},"created_at":"2026-07-05T09:45:44.894362+00:00","updated_at":"2026-07-05T09:45:44.894362+00:00"}