{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2025:IBG7XPTFQGHLPPR45SRE2PDNVA","short_pith_number":"pith:IBG7XPTF","schema_version":"1.0","canonical_sha256":"404dfbbe65818eb7be3ceca24d3c6da80358a0dbc1c26a6f7ef3eb8d47e3314b","source":{"kind":"arxiv","id":"2504.10839","version":1},"attestation_state":"computed","paper":{"title":"Rethinking Theory of Mind Benchmarks for LLMs: Towards A User-Centered Perspective","license":"http://creativecommons.org/licenses/by-nc-sa/4.0/","headline":"","cross_cats":["cs.AI"],"primary_cat":"cs.HC","authors_text":"Hong Shen, Jodi Forlizzi, Maarten Sap, Qiaosi Wang, Xuhui Zhou","submitted_at":"2025-04-15T03:44:43Z","abstract_excerpt":"The last couple of years have witnessed emerging research that appropriates Theory-of-Mind (ToM) tasks designed for humans to benchmark LLM's ToM capabilities as an indication of LLM's social intelligence. However, this approach has a number of limitations. Drawing on existing psychology and AI literature, we summarize the theoretical, methodological, and evaluation limitations by pointing out that certain issues are inherently present in the original ToM tasks used to evaluate human's ToM, which continues to persist and exacerbated when appropriated to benchmark LLM's ToM. Taking a human-comp"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2504.10839","kind":"arxiv","version":1},"metadata":{"license":"http://creativecommons.org/licenses/by-nc-sa/4.0/","primary_cat":"cs.HC","submitted_at":"2025-04-15T03:44:43Z","cross_cats_sorted":["cs.AI"],"title_canon_sha256":"855723536812f9fca8dc5a753c8f24d6e8c7c3f989e16e1169164f771035b905","abstract_canon_sha256":"c91d42d368af1e03312a7e55054302bafc359936fc08c3f27b77db6e68799fa1"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T10:49:24.999833Z","signature_b64":"t1ZnpqFFo+z0laK7cjfmw2Bw5Jmjuy/xOjgk470L7b3DcRktEs/c6fH1SeMPAZIH4/aof6cY4q4xNf0RbX1vDA==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"404dfbbe65818eb7be3ceca24d3c6da80358a0dbc1c26a6f7ef3eb8d47e3314b","last_reissued_at":"2026-07-05T10:49:24.999378Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T10:49:24.999378Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Rethinking Theory of Mind Benchmarks for LLMs: Towards A User-Centered Perspective","license":"http://creativecommons.org/licenses/by-nc-sa/4.0/","headline":"","cross_cats":["cs.AI"],"primary_cat":"cs.HC","authors_text":"Hong Shen, Jodi Forlizzi, Maarten Sap, Qiaosi Wang, Xuhui Zhou","submitted_at":"2025-04-15T03:44:43Z","abstract_excerpt":"The last couple of years have witnessed emerging research that appropriates Theory-of-Mind (ToM) tasks designed for humans to benchmark LLM's ToM capabilities as an indication of LLM's social intelligence. However, this approach has a number of limitations. Drawing on existing psychology and AI literature, we summarize the theoretical, methodological, and evaluation limitations by pointing out that certain issues are inherently present in the original ToM tasks used to evaluate human's ToM, which continues to persist and exacerbated when appropriated to benchmark LLM's ToM. Taking a human-comp"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2504.10839","kind":"arxiv","version":1},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2504.10839/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2504.10839","created_at":"2026-07-05T10:49:24.999439+00:00"},{"alias_kind":"arxiv_version","alias_value":"2504.10839v1","created_at":"2026-07-05T10:49:24.999439+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2504.10839","created_at":"2026-07-05T10:49:24.999439+00:00"},{"alias_kind":"pith_short_12","alias_value":"IBG7XPTFQGHL","created_at":"2026-07-05T10:49:24.999439+00:00"},{"alias_kind":"pith_short_16","alias_value":"IBG7XPTFQGHLPPR4","created_at":"2026-07-05T10:49:24.999439+00:00"},{"alias_kind":"pith_short_8","alias_value":"IBG7XPTF","created_at":"2026-07-05T10:49:24.999439+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":6,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.23840","citing_title":"Embodied Explainability and Ontological Obstacles: Why We Struggle to Explain the Answers of Large Language Models (LLMs)","ref_index":116,"is_internal_anchor":false},{"citing_arxiv_id":"2606.04184","citing_title":"GroupToM-Bench: Benchmarking Group Theory of Mind and Nonlinear Social Emergence in MLLMs","ref_index":10,"is_internal_anchor":false},{"citing_arxiv_id":"2606.02973","citing_title":"Chatbots Output Meaningful (but Problematic) Language","ref_index":74,"is_internal_anchor":false},{"citing_arxiv_id":"2606.31916","citing_title":"Theory of Mind and Persuasion Beyond Conversation: Assessing the Capacity of LLMs to Induce Belief States via Planning and Action","ref_index":3,"is_internal_anchor":false},{"citing_arxiv_id":"2605.09716","citing_title":"Medical Model Synthesis Architectures: A Case Study","ref_index":19,"is_internal_anchor":false},{"citing_arxiv_id":"2604.20443","citing_title":"DialToM: A Theory of Mind Benchmark for Forecasting State-Driven Dialogue Trajectories","ref_index":37,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/IBG7XPTFQGHLPPR45SRE2PDNVA","json":"https://pith.science/pith/IBG7XPTFQGHLPPR45SRE2PDNVA.json","graph_json":"https://pith.science/api/pith-number/IBG7XPTFQGHLPPR45SRE2PDNVA/graph.json","events_json":"https://pith.science/api/pith-number/IBG7XPTFQGHLPPR45SRE2PDNVA/events.json","paper":"https://pith.science/paper/IBG7XPTF"},"agent_actions":{"view_html":"https://pith.science/pith/IBG7XPTFQGHLPPR45SRE2PDNVA","download_json":"https://pith.science/pith/IBG7XPTFQGHLPPR45SRE2PDNVA.json","view_paper":"https://pith.science/paper/IBG7XPTF","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2504.10839&json=true","fetch_graph":"https://pith.science/api/pith-number/IBG7XPTFQGHLPPR45SRE2PDNVA/graph.json","fetch_events":"https://pith.science/api/pith-number/IBG7XPTFQGHLPPR45SRE2PDNVA/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/IBG7XPTFQGHLPPR45SRE2PDNVA/action/timestamp_anchor","attest_storage":"https://pith.science/pith/IBG7XPTFQGHLPPR45SRE2PDNVA/action/storage_attestation","attest_author":"https://pith.science/pith/IBG7XPTFQGHLPPR45SRE2PDNVA/action/author_attestation","sign_citation":"https://pith.science/pith/IBG7XPTFQGHLPPR45SRE2PDNVA/action/citation_signature","submit_replication":"https://pith.science/pith/IBG7XPTFQGHLPPR45SRE2PDNVA/action/replication_record"}},"created_at":"2026-07-05T10:49:24.999439+00:00","updated_at":"2026-07-05T10:49:24.999439+00:00"}