Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:45:21.376172Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 100 of 164 outbound references and 7 inbound Pith citation observations for arXiv:2505.23713.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:45:21.376172Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T19:16:48.882711Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T10:39:44.817622Z
100 of 164 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 4c19d23c-e099-4afe-8365-8597d00d0577 · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models What can large language models do in chemistry? a comprehensive benchmark on eight tasks.Advances in Neural Information Processing Systems, 36:59662–59688, 2023
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c07aa0af-befa-41b6-a881-62f4cba94803 · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Unveiling the power of language models in chemical research question answering.Communications Chemistry, 8(1):4, 2025
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c93a395a-dc3f-4c91-8ab2-24335af568d2 · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Medtrinity-25m: A large-scale multimodal dataset with multigranular annotations for medicine, 2025
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10d83bd2-6924-4e6f-a2db-8a76e44650ed · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Pre-trained multimodal large language model enhances dermatological diagnosis using skingpt-4.Nature Communications, 15(1):5649, 2024
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef91d5c4-f0e5-44ea-96bf-9ae94dfc3a25 · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Evalu- ating and mitigating bias in ai-based medical text generation.Nature Computational Science, pages 1–9, 2025
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73aa2aa5-0385-4117-b5ef-d1edfe8096e6 · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Llm-mod: Can large language models assist content moderation? InExtended Abstracts of the CHI Conference on Human Factors in Computing Systems, pages 1–8, 2024
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0214642d-ba8c-4bd2-8a1e-d138d1e93ce2 · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models ShieldGemma: Generative AI Content Moderation Based on Gemma
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87e8edcc-6e5b-41f2-8c4f-92e4b614862b · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Scaling up llm reviews for google ads content moderation
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0b1784b-7f58-45fa-b2fa-a2a1da2ac8bd · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Hate Personified: Investigating the role of LLMs in content moderation
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ae55d95-54c4-4bc4-9ce5-a2aefbd9603e · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Autonomous agents for collaborative task under information asymmetry
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7a85e38-eebe-4b4f-b1f6-b7b7a774a32a · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Halc: object hallucination reduction via adaptive focal-contrast decoding
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8ecb7a4-1bdb-4242-8599-1271f87ff4f9 · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models LLM-Assisted Content Analysis: Using Large Language Models to Support Deductive Coding
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b54ee77-7f94-4d69-9951-7ae33443b3b5 · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models The Stepwise Deception: Simulating the Evolution from True News to Fake News with LLM Agents
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64aad1e1-38c5-4a3f-b89d-1146e8187355 · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Decoding echo chambers: LLM- powered simulations revealing polarization in social networks
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8eebf24c-cb28-4630-a5e7-9ac7f70642e6 · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Safewatch: An efficient safety-policy following video guardrail model with transparent explanations
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53d70317-cbd7-4854-a6df-03a57914247f · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Theory of Mind for Multi-Agent Collaboration via Large Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d09d512-727a-42b6-9f28-256481b93c14 · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Exploring Large Language Models for Word Games:Who is the Spy?
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2cfbf2c1-199c-4e48-a2b9-febad1800e59 · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Chatbot arena: An open platform for evaluating llms by human preference
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5bba9b19-3104-4b81-8280-0cfa30e6eb40 · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models On the Trustworthiness of Generative Foundation Models: Guideline, Assessment, and Perspective
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9f38520-a65f-4128-a3f0-10133692af8c · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Trusteval: A dynamic evaluation toolkit on trustworthiness of generative foundation models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c47236ff-3e92-407a-b1a9-6bdb42ef4fcc · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Evaluating large language models in theory of mind tasks.Proceedings of the National Academy of Sciences, 121(45):e2405460121, 2024
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1861385d-1771-442f-b1f7-b6ffd1243484 · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models MJ-Bench: Is Your Multimodal Reward Model Really a Good Judge for Text-to-Image Generation?
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ced0f817-bd1d-4b6e-8ce2-9bc40924d527 · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Cross-lingual pitfalls: Automatic probing cross-lingual weakness of multilingual large language models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 144b460c-365a-4c3b-9346-af7033c27931 · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Injecting Domain-Specific Knowledge into Large Language Models: A Comprehensive Survey
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0ce336a-c6c3-40bb-a732-555597637bcf · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Honestllm: Toward an honest and helpful large language model.Advances in Neural Information Processing Systems, 37:7213–7255, 2024
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b705d457-b9ad-4bf6-aedb-6512f0c9a1e6 · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Vuldetect- bench: Evaluating the deep capability of vulnerability detection with large language models, 2024
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53675f87-0002-40fc-ac33-6f073a12c82b · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Datagen: Unified synthetic dataset generation via large language models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb390568-9bb4-43d4-9110-5414a5a2fd84 · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Shaping the safety boundaries: Understanding and defending against jailbreaks in large language models, 2024
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c5fecf5-1d5a-4114-beb0-125e884bc2b3 · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Nqe: N-ary query embedding for complex query answering over hyper-relational knowledge graphs
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c05919c-aa1e-4ef6-8b37-46ce0ca263c2 · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68b15ff9-d70c-4251-b8c0-e09a4dad781e · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Gta: Graph theory agent and benchmark for algorithmic graph reasoning with llms, 2025
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96fdb768-209a-4ab6-be35-47b046e434cc · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models SocialIQA: Commonsense Reasoning about Social Interactions
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2caac88c-83d4-43a6-b5e8-d34fda8f6c67 · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Atomic: Anatlasofmachinecommonsense for if-then reasoning
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 562d762f-b506-4b2a-93db-7968ba8248e3 · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models CREAK: A Dataset for Commonsense Reasoning over Entity Knowledge
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4eb93cc-e80b-4bc4-b8cb-90455851fc98 · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models GoEmotions: A Dataset of Fine-Grained Emotions
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 892c49eb-0b48-4ca7-ba06-f328200e6c82 · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models CommonGen: A constrained text generation challenge for generative commonsense reasoning
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fae1bdab-2a69-466e-af42-439d37bc5b3a · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Clever Hans or Neural Theory of Mind? Stress Testing Social Reasoning in Large Language Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f772664-b5eb-427f-8021-856e067cfbc8 · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Evaluating Theory of Mind in Question Answering
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d50980b8-64d9-48ee-b6e8-489b317aa38a · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Social Chemistry 101: Learning to Reason about Social and Moral Norms
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9e34c8e-6b11-4a98-bdea-0ba7aeae428c · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Aligning ai with shared human values.Proceedings of the International Conference on Learning Representations (ICLR), 2021
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5096160d-dc35-4a12-bd7d-d52ca33d0792 · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models DebateBench: A Challenging Long Context Reasoning Benchmark For Large Language Models
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ed687835-2afc-4ef2-b5bd-8032534ebd3b · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Understanding social reasoning in language models with language models.Advances in Neural Information Processing Systems, 36:13518–13529, 2023
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bed73b27-ddcb-4714-b0e6-b02bd592d5fa · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Mastering the Game of No-Press Diplomacy via Human-Regularized Reinforcement Learning and Planning
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 160bc2c0-759b-4f4d-9147-24b6198c4e06 · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models AvalonBench: Evaluating LLMs Playing the Game of Avalon
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6df2bdba-089d-4aee-bafa-a3f3e4da6171 · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Werewolf Arena: A Case Study in LLM Evaluation via Social Deduction
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8346a62d-881e-42ce-8837-f77d96954f15 · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Exploring Large Language Models for Communication Games: An Empirical Study on Werewolf
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7afb8966-e0ae-4455-b482-01614fbc96dd · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Does the chimpanzee have a theory of mind?Behavioral and brain sciences, 1(4):515–526, 1978
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2771bbea-358a-45a0-8a3c-0386c43f734a · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Oxford University Press, 2014
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation caa6c6f4-d336-4015-bb62-7ccfabd777d7 · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models MIT press, 1999
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e9c919f-6e1e-4fc7-ba27-efe5618d6d4e · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Social cognition in humans.Current biology, 17(16):R724–R732, 2007
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e27229f5-1234-43d0-bed7-c944e88d070f · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Counterfactual thinking.Psychological bulletin, 121(1):133, 1997
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42f63aa8-154a-4dfd-a993-13a2ecee78e4 · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models CLOMO: Counterfactual Logical Modification with Large Language Models
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 06583048-38d9-4073-8ab5-b8fd7f349246 · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models LLM as a Mastermind: A Survey of Strategic Reasoning with Large Language Models
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57e784be-1ab8-44fe-a2cb-722ecdff9b39 · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models What is agency?American journal of sociology, 103(4):962– 1023, 1998
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5ac645c-90ae-47fd-a850-83a416fdf9b5 · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Penguin, 2014
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 531f0aae-4831-4879-8414-49c086fa0760 · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Cambridge University Press, 2008
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7d3c459-8d16-4d8c-9bac-7aa76bc442e6 · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Generative agents: Interactive simulacra of human behavior
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34db5b5e-464e-4998-a1af-daab964ecc81 · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Shieldagent: Shielding agents via verifiable safety policy reasoning.arXiv preprint arXiv:2503.22738, 2025
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3664a79-20f0-44d4-b9a6-75cf166c320a · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Fusing heterogeneous data: A case for remote sensing and social media.IEEE Transactions on Geoscience and Remote Sensing, 56(12):6956–6968, 2018
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5da8a377-711a-4360-b89c-e46dd09446d9 · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models How noisy social media text, how diffrnt social media sources? InProceedings of the sixth international joint conference on natural language processing, pages 356–364, 2013
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6d3cc52-1d5d-4440-a460-5bdb623a1096 · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Social-ecological systems as complex adaptive systems.Ecology and Society, 23(4), 2018
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70afb51a-c228-4a3c-900a-458fff4c341f · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models The science of fake news.Science, 359(6380):1094–1096, 2018
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9039d21-c660-4f93-bfb5-837990d164fc · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Agentpoison: Red-teaming llm agents via poisoning memory or knowledge bases.Advances in Neural Information Processing Systems, 37:130185–130213, 2024
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a1f52dc-98bd-4507-af31-9d066e995cb7 · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Automated Design of Agentic Systems
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c3c57b8-a7c9-4502-9dbb-dea711fc0f59 · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models AFlow: Automating Agentic Workflow Generation
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26de77b4-bcac-45a5-b411-d2f40aa55df2 · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Multi-agent Architecture Search via Agentic Supernet
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c10eb10e-014a-45aa-bf32-90065f8b8b7c · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Blood on the clocktower, 2024
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4c509bb-e103-454c-b7a1-5aa47847ca92 · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Improving fac- tuality and reasoning in language models through multiagent debate
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3087291-5e04-4242-b044-f3d893e28ac5 · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Self-refine: Iterative refinement with self-feedback.Advances in Neural Information Processing Systems, 36:46534–46594, 2023
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c90b2b0d-60c0-4ace-864f-e2afc45173c2 · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Dyflow: Dynamic workflow framework for agentic reasoning
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdf1ecc5-6a52-428f-b5b2-d3d115afe017 · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f5ebcac-d920-474e-bc80-35af9b5760a6 · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Direct preference optimization: Your language model is secretly a reward model
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3cf05bba-a277-4da6-91e2-2a1a66b57aa0 · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Glucose: Generalized and contextualized story explanations
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d30a6f0-1c1a-41be-95d2-18e1f0d21c23 · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Piqa: Reasoning about physical commonsense in natural language
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a74cdb97-f283-4414-8db2-991dbf3c0f13 · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models HellaSwag: Can a Machine Really Finish Your Sentence?
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75b2f6fc-b31c-4bdc-8509-7bafd5e238d5 · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Theory of mind may have spontaneously emerged in large language models
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10c98bb8-a181-4f77-b23d-126e62752d1b · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Breaking focus: Contextual distraction curse in large language models.arXiv preprint arXiv:2502.01609, 2025
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 376a9404-d10c-4302-84ff-918afe3a131b · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models FANToM: A Benchmark for Stress-testing Machine Theory of Mind in Interactions
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 834c773b-beb2-451b-93fe-ddf686541c21 · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Tomvalley: Evaluating the theory of mind reasoning of llms in realistic social context
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 024dd819-cad6-4501-8a0a-ebaa415fc6e6 · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models An Open Review of OpenReview: A Critical Analysis of the Machine Learning Conference Review Process
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 760e7e7f-9606-494e-b50d-4690a4b86426 · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models The Open Review-Based (ORB) dataset: Towards Automatic Assessment of Scientific Papers and Experiment Proposals in High-Energy Physics
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5f241ba2-4413-4f28-bf6e-d4d71f446b19 · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Superhuman ai for multiplayer poker.Science, 365(6456):885–890, 2019
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f76f0b62-ae4c-48e7-a4df-c9c64a52d911 · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Geolocation with real human gameplay data: A large-scale dataset and human-like reasoning framework.arXiv preprint arXiv:2502.13759, 2025
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49ecd989-94bc-4019-b6cf-4a2b8c9a7f91 · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Word form matters: Llms’ semantic reconstruction under typoglycemia, 2025
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d23c6e0-dd3b-4975-9538-84b8fd3bba74 · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models TextAttack: A Framework for Adversarial Attacks, Data Augmentation, and Adversarial Training in NLP
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d0967d5-4ab6-42ab-9711-1d6e63576fe6 · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models TrustLLM: Trustworthiness in Large Language Models
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0bf59a51-b52f-47b5-8970-cf99e84bf206 · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models GPT-4o System Card
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab8e9951-1361-4cbe-a916-4c539efab799 · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Gpt-4o mini: Advancing cost-efficient intelligence.https://openai.com/index/ gpt-4o-mini-advancing-\cost-efficient-intelligence/, 2024
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3fb4e4bc-8c02-4a0c-8ddf-2ea039558105 · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models o3-mini.https://docsbot.ai/models/o3-mini, January 2025
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06751ed8-7c43-42b4-80ae-1ac0f08b4b78 · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models OpenAI o1 System Card
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ec7c936-bbd4-42c2-b9a7-5b26e2a44b73 · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Gemini 2.5 pro experimental.https://ai.google.dev/gemini-api/ docs/models, March 2025
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation caaf9613-11be-4b6b-a4c8-cd19e9351c51 · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Phi-4 Technical Report
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 212e4679-bfb0-4684-89a4-2ab9909d1e2c · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Llama 3.1-8b.https://huggingface.co/meta-llama/Llama-3.1-8B, 2024
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c4a0d9d-9bee-408b-b4f2-5047eab5104b · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Llama 3.3-70b
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 933701dc-3182-4c36-9915-900439b7e4c9 · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Qwen2.5: A party of foundation models, September 2024
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66236904-c4d5-445f-8756-cb64ae00842c · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Qwq-32b: Embracing the power of reinforcement learning, March 2025
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea100ba9-af50-4419-b410-edca7141a3cd · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56e6ff9c-e1ef-4bbc-a08e-b14f0fb68dc8 · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Unresolved cited work
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f2ed3cf-d633-4a12-b608-799c816a1980 · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models 𝑢 is Criminal
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbdad187-d187-4f70-a823-edddd76df300 · outbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Player𝑣 says Player𝑢 is the criminal
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bfb07b62-2001-4fcf-bbaa-f2081bb94580 · inbound
SocialGrid: A Benchmark for Planning and Social Reasoning in Embodied Multi-Agent Systems SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation af7acf9e-6748-4e34-83d5-a80c1cd97551 · inbound
GroupToM-Bench: Benchmarking Group Theory of Mind and Nonlinear Social Emergence in MLLMs SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fde092a4-5309-4edc-9509-25b25cc7dd86 · inbound
OpenSkill: Open-World Self-Evolution for LLM Agents SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0785db62-f8a9-470a-88ea-c06e0644d9b4 · inbound
Getting Better at Working With You: Compiling User Corrections into Runtime Enforcement for Coding Agents SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 650c9ba9-2d06-45ee-84df-fb206940328e · inbound
PIVOTSBench: Evaluating Fine-Grained Interpersonal Relationship Reasoning in Multimodal Large Language Models SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 24a806cd-70cc-405e-9802-a554c05f7c9b · inbound
Inside the Unfair Judge: A Mechanistic Interpretability Account of LLM-as-Judge Bias SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models
Reference 218
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a631bd5-d10f-4067-a659-f09cd98f4aea · inbound
No One Wins in Nuclear War: A Social Simulation of Military Decision-making SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.