Pith. sign in

Paper Citation Record · LEDGER

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models

As of 23 August 2026, this Paper Citation Record lists 100 of 164 outbound references and 9 inbound Pith citation observations for arXiv:2505.23713.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.23713 v1

Coverage vector

measured 100 of 164 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:45:21.376172Z

measured 109 of 109 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:19:51.138159Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T10:39:44.817622Z

Reference resolution

100 of 164 outbound references displayed

  • verified exact3
  • verified fuzzy0
  • unresolved97
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4c19d23c-e099-4afe-8365-8597d00d0577 · outbound

This paper cites What can large language models do in chemistry? a comprehensive benchmark on eight tasks.Advances in Neural Information Processing Systems, 36:59662–59688, 2023.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models What can large language models do in chemistry? a comprehensive benchmark on eight tasks.Advances in Neural Information Processing Systems, 36:59662–59688, 2023

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:10.560836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:10.560836Z digest=sha256:ee2c41bb35509aeec4559ba817affea78a31c4aebe94c9f41a3cd1e561c60681

Observation c07aa0af-befa-41b6-a881-62f4cba94803 · outbound

This paper cites Unveiling the power of language models in chemical research question answering.Communications Chemistry, 8(1):4, 2025.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Unveiling the power of language models in chemical research question answering.Communications Chemistry, 8(1):4, 2025

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:10.641139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:10.641139Z digest=sha256:7af8b7e6ac09923fed126fe1e2e3ad5d2c7631ba6314bcf9462c1ded5ded9825

Observation c93a395a-dc3f-4c91-8ab2-24335af568d2 · outbound

This paper cites Medtrinity-25m: A large-scale multimodal dataset with multigranular annotations for medicine, 2025.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Medtrinity-25m: A large-scale multimodal dataset with multigranular annotations for medicine, 2025

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:10.681909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:10.681909Z digest=sha256:7ada1d187ad8cfbdf129b4f464ef0a1f74bbd6a25a97d69dd9b0a78afb3c9a0b

Observation 10d83bd2-6924-4e6f-a2db-8a76e44650ed · outbound

This paper cites Pre-trained multimodal large language model enhances dermatological diagnosis using skingpt-4.Nature Communications, 15(1):5649, 2024.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Pre-trained multimodal large language model enhances dermatological diagnosis using skingpt-4.Nature Communications, 15(1):5649, 2024

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:10.787564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:10.787564Z digest=sha256:d831cbd5b7f29fabbae0eee5b699318db225da9a2c1fadef35706b72757578bc

Observation ef91d5c4-f0e5-44ea-96bf-9ae94dfc3a25 · outbound

This paper cites Evalu- ating and mitigating bias in ai-based medical text generation.Nature Computational Science, pages 1–9, 2025.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Evalu- ating and mitigating bias in ai-based medical text generation.Nature Computational Science, pages 1–9, 2025

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:10.877179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:10.877179Z digest=sha256:9ecbbd6f90bcb5460d590987bd40f2196281fc92b541ba20c2a2c2e79372d0a6

Observation 73aa2aa5-0385-4117-b5ef-d1edfe8096e6 · outbound

This paper cites Llm-mod: Can large language models assist content moderation? InExtended Abstracts of the CHI Conference on Human Factors in Computing Systems, pages 1–8, 2024.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Llm-mod: Can large language models assist content moderation? InExtended Abstracts of the CHI Conference on Human Factors in Computing Systems, pages 1–8, 2024

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:10.952094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:10.952094Z digest=sha256:2d455846779b7a143fb27aac7d934a1d0128cedaa1476439d8f693d6bb74a376

Observation 0214642d-ba8c-4bd2-8a1e-d138d1e93ce2 · outbound

This paper cites ShieldGemma: Generative AI Content Moderation Based on Gemma.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models ShieldGemma: Generative AI Content Moderation Based on Gemma

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:11.036015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:11.036015Z digest=sha256:c2f642aea3cf905b420583c725416009fadfffbe785fb34413a3bb7975537f62

Observation 87e8edcc-6e5b-41f2-8c4f-92e4b614862b · outbound

This paper cites Scaling up llm reviews for google ads content moderation.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Scaling up llm reviews for google ads content moderation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:11.119173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:11.119173Z digest=sha256:265599f86c3a239b430f319d4f0bc6b379fbf846cfb6a0922dce560f3b4480f8

Observation c0b1784b-7f58-45fa-b2fa-a2a1da2ac8bd · outbound

This paper cites Hate Personified: Investigating the role of LLMs in content moderation.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Hate Personified: Investigating the role of LLMs in content moderation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:11.169211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:11.169211Z digest=sha256:40194d0ad6728f2d467363de0ae4987191dcf397bb957b0b77f5654c1abb0bd3

Observation 6ae55d95-54c4-4bc4-9ce5-a2aefbd9603e · outbound

This paper cites Autonomous agents for collaborative task under information asymmetry.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Autonomous agents for collaborative task under information asymmetry

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:11.239780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:11.239780Z digest=sha256:bbeb0181c575733350d87172de13c65e0673e64030bd6319b7c413e15ae262e8

Observation e7a85e38-eebe-4b4f-b1f6-b7b7a774a32a · outbound

This paper cites Halc: object hallucination reduction via adaptive focal-contrast decoding.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Halc: object hallucination reduction via adaptive focal-contrast decoding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:11.374316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:11.374316Z digest=sha256:1a4fae97fd6fc6b68452a3d9c01aa40f2b71e3bc8bbd75fc66424171234e81d9

Observation c8ecb7a4-1bdb-4242-8599-1271f87ff4f9 · outbound

This paper cites LLM-Assisted Content Analysis: Using Large Language Models to Support Deductive Coding.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models LLM-Assisted Content Analysis: Using Large Language Models to Support Deductive Coding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:11.486260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:11.486260Z digest=sha256:2aa5b87008d070b0f423e777c4ffaa440b23156426f4bfcafef4fdb9a3af377e

Observation 5b54ee77-7f94-4d69-9951-7ae33443b3b5 · outbound

This paper cites The Stepwise Deception: Simulating the Evolution from True News to Fake News with LLM Agents.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models The Stepwise Deception: Simulating the Evolution from True News to Fake News with LLM Agents

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:11.630040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:11.630040Z digest=sha256:909d6b88aac262af0b7aaa1482c31e62913f2368e1c34cc9db3b3f4bfed3fc85

Observation 64aad1e1-38c5-4a3f-b89d-1146e8187355 · outbound

This paper cites Decoding echo chambers: LLM- powered simulations revealing polarization in social networks.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Decoding echo chambers: LLM- powered simulations revealing polarization in social networks

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:11.698038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:11.698038Z digest=sha256:d40c5862f23cc914ea29423417fc95e556e4fcf308ae35e9a724959bca1ee5ae

Observation 8eebf24c-cb28-4630-a5e7-9ac7f70642e6 · outbound

This paper cites Safewatch: An efficient safety-policy following video guardrail model with transparent explanations.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Safewatch: An efficient safety-policy following video guardrail model with transparent explanations

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:11.850117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:11.850117Z digest=sha256:a00e44598b0d45646476a4f95f17e058b60b6067612376398cbbfd06ce9ca8bd

Observation 53d70317-cbd7-4854-a6df-03a57914247f · outbound

This paper cites Theory of Mind for Multi-Agent Collaboration via Large Language Models.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Theory of Mind for Multi-Agent Collaboration via Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:11.951679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:11.951679Z digest=sha256:672d25279d494cd28262f6f910514b0ff2e252933179c58867834cd5f583be26

Observation 9d09d512-727a-42b6-9f28-256481b93c14 · outbound

This paper cites Exploring Large Language Models for Word Games:Who is the Spy?.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Exploring Large Language Models for Word Games:Who is the Spy?

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:12.071230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:12.071230Z digest=sha256:6bdc67e141b1ceb3ed29b347f78d6cef1cf78bb85d0affa1a3fcfedeee25eaaf

Observation 2cfbf2c1-199c-4e48-a2b9-febad1800e59 · outbound

This paper cites Chatbot arena: An open platform for evaluating llms by human preference.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Chatbot arena: An open platform for evaluating llms by human preference

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:12.176186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:12.176186Z digest=sha256:be681814f3bf3bc2d837b98875096018f6aeb7ca14cbef4f4fe1419d15ea4314

Observation 5bba9b19-3104-4b81-8280-0cfa30e6eb40 · outbound

This paper cites On the Trustworthiness of Generative Foundation Models: Guideline, Assessment, and Perspective.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models On the Trustworthiness of Generative Foundation Models: Guideline, Assessment, and Perspective

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:12.276453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:12.276453Z digest=sha256:fe5bfff6944a823a2370806b37f2c11d92a49bfa5dd5fe426a693fe37ac5f5f2

Observation d9f38520-a65f-4128-a3f0-10133692af8c · outbound

This paper cites Trusteval: A dynamic evaluation toolkit on trustworthiness of generative foundation models.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Trusteval: A dynamic evaluation toolkit on trustworthiness of generative foundation models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:12.374411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:12.374411Z digest=sha256:9abd529f20211dde40727a26f085b365091846811a1f5bedf98f375b8b686908

Observation c47236ff-3e92-407a-b1a9-6bdb42ef4fcc · outbound

This paper cites Evaluating large language models in theory of mind tasks.Proceedings of the National Academy of Sciences, 121(45):e2405460121, 2024.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Evaluating large language models in theory of mind tasks.Proceedings of the National Academy of Sciences, 121(45):e2405460121, 2024

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:12.480039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:12.480039Z digest=sha256:9174d693f98ef59818f739930f10a9592e71331e62119bb3d2c065c773ca27ea

Observation 1861385d-1771-442f-b1f7-b6ffd1243484 · outbound

This paper cites MJ-Bench: Is Your Multimodal Reward Model Really a Good Judge for Text-to-Image Generation?.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models MJ-Bench: Is Your Multimodal Reward Model Really a Good Judge for Text-to-Image Generation?

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:12.627945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:12.627945Z digest=sha256:6a3372519d033d93ee73fa36475f68c82409fe8ef7c5e02686698d00aeb344c5

Observation ced0f817-bd1d-4b6e-8ce2-9bc40924d527 · outbound

This paper cites Cross-lingual pitfalls: Automatic probing cross-lingual weakness of multilingual large language models.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Cross-lingual pitfalls: Automatic probing cross-lingual weakness of multilingual large language models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:12.786418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:12.786418Z digest=sha256:f531afb7b38322b1d6806a2a2533aa757449f93e19dad60343e5ba491f027805

Observation 144b460c-365a-4c3b-9346-af7033c27931 · outbound

This paper cites Injecting Domain-Specific Knowledge into Large Language Models: A Comprehensive Survey.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Injecting Domain-Specific Knowledge into Large Language Models: A Comprehensive Survey

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:12.924587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:12.924587Z digest=sha256:615c68a2e8b3c369bfb35a4a9ccd18e2886f9a7897578ed287f6da9a4bc57fc7

Observation a0ce336a-c6c3-40bb-a732-555597637bcf · outbound

This paper cites Honestllm: Toward an honest and helpful large language model.Advances in Neural Information Processing Systems, 37:7213–7255, 2024.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Honestllm: Toward an honest and helpful large language model.Advances in Neural Information Processing Systems, 37:7213–7255, 2024

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:13.044373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:13.044373Z digest=sha256:e9436caa4aa36e8d2018a19d3a72c6e2a6ba72eab3242284905fa4190e2834f8

Observation b705d457-b9ad-4bf6-aedb-6512f0c9a1e6 · outbound

This paper cites Vuldetect- bench: Evaluating the deep capability of vulnerability detection with large language models, 2024.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Vuldetect- bench: Evaluating the deep capability of vulnerability detection with large language models, 2024

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:13.160590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:13.160590Z digest=sha256:39e1a7f041e0505e3a07e37d51a20d6bd50be6ea9fa50354df44f4cd26e7b1c0

Observation 53675f87-0002-40fc-ac33-6f073a12c82b · outbound

This paper cites Datagen: Unified synthetic dataset generation via large language models.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Datagen: Unified synthetic dataset generation via large language models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:13.306776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:13.306776Z digest=sha256:56710df851dfe834e4071e47c27c21793325ec04a72868a1e65b9aa9ab92edb1

Observation cb390568-9bb4-43d4-9110-5414a5a2fd84 · outbound

This paper cites Shaping the safety boundaries: Understanding and defending against jailbreaks in large language models, 2024.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Shaping the safety boundaries: Understanding and defending against jailbreaks in large language models, 2024

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:13.459963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:13.459963Z digest=sha256:3f10ad4e6130ca8aff65bf33ad878a23a7ebff96c78aabc4a3851cd98954356a

Observation 6c5fecf5-1d5a-4114-beb0-125e884bc2b3 · outbound

This paper cites Nqe: N-ary query embedding for complex query answering over hyper-relational knowledge graphs.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Nqe: N-ary query embedding for complex query answering over hyper-relational knowledge graphs

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:13.575817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:13.575817Z digest=sha256:42b247c85fe7341bba5e2b7812bc1aac2d2ddf78cb2c5ca3765abfb3db73465e

Observation 4c05919c-aa1e-4ef6-8b37-46ce0ca263c2 · outbound

This paper cites Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:13.720220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:13.720220Z digest=sha256:6e80f77a84c5e516df80f9ea9c8248007f0168cfc6275473eafdccdc075ed194

Observation 68b15ff9-d70c-4251-b8c0-e09a4dad781e · outbound

This paper cites Gta: Graph theory agent and benchmark for algorithmic graph reasoning with llms, 2025.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Gta: Graph theory agent and benchmark for algorithmic graph reasoning with llms, 2025

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:13.834447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:13.834447Z digest=sha256:2bf79ba2ad8503340df7b6f19f0da457f16bd1fb49dca94b2c7a0f43d000f8f6

Observation 96fdb768-209a-4ab6-be35-47b046e434cc · outbound

This paper cites SocialIQA: Commonsense Reasoning about Social Interactions.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models SocialIQA: Commonsense Reasoning about Social Interactions

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:13.983646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:13.983646Z digest=sha256:be5e0ecd781a69f0e2fdffa663f6c9956a96b061288376dc0fb29ee63ebf83cf

Observation 2caac88c-83d4-43a6-b5e8-d34fda8f6c67 · outbound

This paper cites Atomic: Anatlasofmachinecommonsense for if-then reasoning.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Atomic: Anatlasofmachinecommonsense for if-then reasoning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:14.142371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:14.142371Z digest=sha256:301ead42c352e7d25fdf36264c357edb6684964229c83e3569364f2c5c7acde8

Observation 562d762f-b506-4b2a-93db-7968ba8248e3 · outbound

This paper cites CREAK: A Dataset for Commonsense Reasoning over Entity Knowledge.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models CREAK: A Dataset for Commonsense Reasoning over Entity Knowledge

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:14.262662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:14.262662Z digest=sha256:dafd87b895148c705e0a2d5f40cc23859eefb9e642efd4e3924fe6b104b8c180

Observation d4eb93cc-e80b-4bc4-b8cb-90455851fc98 · outbound

This paper cites GoEmotions: A Dataset of Fine-Grained Emotions.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models GoEmotions: A Dataset of Fine-Grained Emotions

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:14.387941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:14.387941Z digest=sha256:49247318035221358fd5c633791735a625e5af127909788a0ff02e09eb4783d4

Observation 892c49eb-0b48-4ca7-ba06-f328200e6c82 · outbound

This paper cites CommonGen: A constrained text generation challenge for generative commonsense reasoning.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models CommonGen: A constrained text generation challenge for generative commonsense reasoning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:14.537634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:14.537634Z digest=sha256:a905cf53f41358fc75aaca2182ec7df5372533226c87e473d7613ea31edb3881

Observation fae1bdab-2a69-466e-af42-439d37bc5b3a · outbound

This paper cites Clever Hans or Neural Theory of Mind? Stress Testing Social Reasoning in Large Language Models.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Clever Hans or Neural Theory of Mind? Stress Testing Social Reasoning in Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:14.673881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:14.673881Z digest=sha256:526aa28e14ad1024e1fdd5de72e28851ead7326c53b69ed8d6cb77524b2b9d83

Observation 3f772664-b5eb-427f-8021-856e067cfbc8 · outbound

This paper cites Evaluating Theory of Mind in Question Answering.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Evaluating Theory of Mind in Question Answering

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:14.837395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:14.837395Z digest=sha256:d9cbee6db0f25c0f269e26f908d48d08891a1e82fba9eec0d82a5e7522616621

Observation d50980b8-64d9-48ee-b6e8-489b317aa38a · outbound

This paper cites Social Chemistry 101: Learning to Reason about Social and Moral Norms.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Social Chemistry 101: Learning to Reason about Social and Moral Norms

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:15.026411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:15.026411Z digest=sha256:53e618cc890014c375d5ac056159d2677fa25746a9be99f698e73c53c9ca2afc

Observation b9e34c8e-6b11-4a98-bdea-0ba7aeae428c · outbound

This paper cites Aligning ai with shared human values.Proceedings of the International Conference on Learning Representations (ICLR), 2021.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Aligning ai with shared human values.Proceedings of the International Conference on Learning Representations (ICLR), 2021

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:15.176837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:15.176837Z digest=sha256:d52c483ca16e6125758507b0d573d7d39af13f7fa22841d92f8849f2bf076702

Observation 5096160d-dc35-4a12-bd7d-d52ca33d0792 · outbound

This paper cites DebateBench: A Challenging Long Context Reasoning Benchmark For Large Language Models.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models DebateBench: A Challenging Long Context Reasoning Benchmark For Large Language Models

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:45:27.847772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T12:45:15.313592Z digest=sha256:3732291d1811dc19143b45a0e556395f2b205aeccf3b749e7ef232b93bd3a5c3

Observation ed687835-2afc-4ef2-b5bd-8032534ebd3b · outbound

This paper cites Understanding social reasoning in language models with language models.Advances in Neural Information Processing Systems, 36:13518–13529, 2023.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Understanding social reasoning in language models with language models.Advances in Neural Information Processing Systems, 36:13518–13529, 2023

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:15.443725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:15.443725Z digest=sha256:88270371034b02651fa83a13fcb8ecc1a6a2539e518d3e8d4d8ef67a14c0ed24

Observation bed73b27-ddcb-4714-b0e6-b02bd592d5fa · outbound

This paper cites Mastering the Game of No-Press Diplomacy via Human-Regularized Reinforcement Learning and Planning.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Mastering the Game of No-Press Diplomacy via Human-Regularized Reinforcement Learning and Planning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:15.575617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:15.575617Z digest=sha256:7cb5f26f11a3ef80947ff67c721db8206c9c75d6311dbb353ed7aa196ca6bd4b

Observation 160bc2c0-759b-4f4d-9147-24b6198c4e06 · outbound

This paper cites AvalonBench: Evaluating LLMs Playing the Game of Avalon.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models AvalonBench: Evaluating LLMs Playing the Game of Avalon

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:15.723111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:15.723111Z digest=sha256:f8aae9636acc5891e7235c8daacb3db519d293a5c94f6d087e005ec3b9c6ef63

Observation 6df2bdba-089d-4aee-bafa-a3f3e4da6171 · outbound

This paper cites Werewolf Arena: A Case Study in LLM Evaluation via Social Deduction.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Werewolf Arena: A Case Study in LLM Evaluation via Social Deduction

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:15.845155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:15.845155Z digest=sha256:08be3e844da6bbf963c996ce26f38085eedfd02c327cea60bfe8b7b5950465a8

Observation 8346a62d-881e-42ce-8837-f77d96954f15 · outbound

This paper cites Exploring Large Language Models for Communication Games: An Empirical Study on Werewolf.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Exploring Large Language Models for Communication Games: An Empirical Study on Werewolf

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:15.951141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:15.951141Z digest=sha256:f89657b916afd8c239f417caa62683137d9601eec235b358a88e5b1391c2426c

Observation 7afb8966-e0ae-4455-b482-01614fbc96dd · outbound

This paper cites Does the chimpanzee have a theory of mind?Behavioral and brain sciences, 1(4):515–526, 1978.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Does the chimpanzee have a theory of mind?Behavioral and brain sciences, 1(4):515–526, 1978

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:16.080047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:16.080047Z digest=sha256:b1482ef0f71c2e931fcd76593f6b681775e55767389650fc17d803afbaa59783

Observation 2771bbea-358a-45a0-8a3c-0386c43f734a · outbound

This paper cites Oxford University Press, 2014.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Oxford University Press, 2014

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:16.198939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:16.198939Z digest=sha256:045c331ca11c605af6bba0ce413b88cbdb4e57f1d4ca611481ec47342d96e21d

Observation caa6c6f4-d336-4015-bb62-7ccfabd777d7 · outbound

This paper cites MIT press, 1999.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models MIT press, 1999

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:16.279846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:16.279846Z digest=sha256:b109257f12772f925bb5904caa3c7d5bfe784af441fd7dcb8f7e3cbb38406586

Observation 4e9c919f-6e1e-4fc7-ba27-efe5618d6d4e · outbound

This paper cites Social cognition in humans.Current biology, 17(16):R724–R732, 2007.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Social cognition in humans.Current biology, 17(16):R724–R732, 2007

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:16.346129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:16.346129Z digest=sha256:30e89a199160921dee9189780f0769529f3c5410ea3d00f29c27d526f1e7cb3c

Observation e27229f5-1234-43d0-bed7-c944e88d070f · outbound

This paper cites Counterfactual thinking.Psychological bulletin, 121(1):133, 1997.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Counterfactual thinking.Psychological bulletin, 121(1):133, 1997

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:16.404998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:16.404998Z digest=sha256:0a3eb1c17ea5c9c2e8fc470afcb2a860ee8f3e4963d8b934a056d6380e0862a0

Observation 42f63aa8-154a-4dfd-a993-13a2ecee78e4 · outbound

This paper cites CLOMO: Counterfactual Logical Modification with Large Language Models.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models CLOMO: Counterfactual Logical Modification with Large Language Models

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:45:27.628490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T12:45:16.459659Z digest=sha256:005f6ac94caae9ab86f2bc134baa961c94caa55256fe403a4b6624b600d3058a

Observation 06583048-38d9-4073-8ab5-b8fd7f349246 · outbound

This paper cites LLM as a Mastermind: A Survey of Strategic Reasoning with Large Language Models.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models LLM as a Mastermind: A Survey of Strategic Reasoning with Large Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:16.514813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:16.514813Z digest=sha256:7cd912eb4dd7ede08755a31572ef61a23a8d097589201d248462f03bf214107c

Observation 57e784be-1ab8-44fe-a2cb-722ecdff9b39 · outbound

This paper cites What is agency?American journal of sociology, 103(4):962– 1023, 1998.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models What is agency?American journal of sociology, 103(4):962– 1023, 1998

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:16.598584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:16.598584Z digest=sha256:e6cec79d8a5e3a1cbdc6033957b2b23052cd031b3c53a057c045f9e2ba82e943

Observation a5ac645c-90ae-47fd-a850-83a416fdf9b5 · outbound

This paper cites Penguin, 2014.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Penguin, 2014

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:16.670659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:16.670659Z digest=sha256:d8358314476d85f687ac9dd106892bf6d0ea54331fcc8e649f51a3142fdc65a2

Observation 531f0aae-4831-4879-8414-49c086fa0760 · outbound

This paper cites Cambridge University Press, 2008.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Cambridge University Press, 2008

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:16.710690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:16.710690Z digest=sha256:45b657cebdd08910f4c2460284c74ca5e4ea5cd1d90496de6e581c957ca104cc

Observation a7d3c459-8d16-4d8c-9bac-7aa76bc442e6 · outbound

This paper cites Generative agents: Interactive simulacra of human behavior.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Generative agents: Interactive simulacra of human behavior

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:16.890152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:16.890152Z digest=sha256:c78c39c5330add662b0262f380eefee70e2e90512417e57321be233a802754f4

Observation 34db5b5e-464e-4998-a1af-daab964ecc81 · outbound

This paper cites Shieldagent: Shielding agents via verifiable safety policy reasoning.arXiv preprint arXiv:2503.22738, 2025.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Shieldagent: Shielding agents via verifiable safety policy reasoning.arXiv preprint arXiv:2503.22738, 2025

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:16.957072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:16.957072Z digest=sha256:6d6ec3508ab80e67332d5356187bfda877e1016c04b2094472ec5c7dd3e8c342

Observation a3664a79-20f0-44d4-b9a6-75cf166c320a · outbound

This paper cites Fusing heterogeneous data: A case for remote sensing and social media.IEEE Transactions on Geoscience and Remote Sensing, 56(12):6956–6968, 2018.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Fusing heterogeneous data: A case for remote sensing and social media.IEEE Transactions on Geoscience and Remote Sensing, 56(12):6956–6968, 2018

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:17.044966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:17.044966Z digest=sha256:8894c0bef8cd5bd1c517978e323996a2fff5baa266bae70d870147aaa99f8f88

Observation 5da8a377-711a-4360-b89c-e46dd09446d9 · outbound

This paper cites How noisy social media text, how diffrnt social media sources? InProceedings of the sixth international joint conference on natural language processing, pages 356–364, 2013.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models How noisy social media text, how diffrnt social media sources? InProceedings of the sixth international joint conference on natural language processing, pages 356–364, 2013

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:17.136964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:17.136964Z digest=sha256:6c10d6acf20139a1affb05847a9842c7cacf4a1aabe399d527813e39800e20bb

Observation c6d3cc52-1d5d-4440-a460-5bdb623a1096 · outbound

This paper cites Social-ecological systems as complex adaptive systems.Ecology and Society, 23(4), 2018.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Social-ecological systems as complex adaptive systems.Ecology and Society, 23(4), 2018

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:17.228510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:17.228510Z digest=sha256:789e05a8573d721a1fdfe047d7b8a272f463833cb27dbb699ec5183af3adbb08

Observation 70afb51a-c228-4a3c-900a-458fff4c341f · outbound

This paper cites The science of fake news.Science, 359(6380):1094–1096, 2018.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models The science of fake news.Science, 359(6380):1094–1096, 2018

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:17.336305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:17.336305Z digest=sha256:381e08b3879548442a51634cc2170527de5e15b40fa761a6b7bcb78c724c5756

Observation f9039d21-c660-4f93-bfb5-837990d164fc · outbound

This paper cites Agentpoison: Red-teaming llm agents via poisoning memory or knowledge bases.Advances in Neural Information Processing Systems, 37:130185–130213, 2024.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Agentpoison: Red-teaming llm agents via poisoning memory or knowledge bases.Advances in Neural Information Processing Systems, 37:130185–130213, 2024

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:17.421245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:17.421245Z digest=sha256:6ada5b48f38fdc4bab065c9658efb182eab2b963159abdbef75217033e970322

Observation 7a1f52dc-98bd-4507-af31-9d066e995cb7 · outbound

This paper cites Automated Design of Agentic Systems.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Automated Design of Agentic Systems

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:17.499384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:17.499384Z digest=sha256:6d1b70c8c2bb5de24ab7bb84226f58de9d3ef0ef7197349ac05ec896b71415ee

Observation 5c3c57b8-a7c9-4502-9dbb-dea711fc0f59 · outbound

This paper cites AFlow: Automating Agentic Workflow Generation.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models AFlow: Automating Agentic Workflow Generation

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:17.603078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:17.603078Z digest=sha256:3ff49d52325f7ef72dbb35719e862d94fcb630ce536f80de8b14b1d5922c1404

Observation 26de77b4-bcac-45a5-b411-d2f40aa55df2 · outbound

This paper cites Multi-agent Architecture Search via Agentic Supernet.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Multi-agent Architecture Search via Agentic Supernet

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:17.691822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:17.691822Z digest=sha256:ffd1d3fd657d03693413bd34dcedbd4e71efbf4f326945a4aeec87907914f2dd

Observation c10eb10e-014a-45aa-bf32-90065f8b8b7c · outbound

This paper cites Blood on the clocktower, 2024.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Blood on the clocktower, 2024

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:17.795238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:17.795238Z digest=sha256:556e3614b63ba0c7e0ebfcab36ec01eeb88a46e11055d1cb4158ca61d0829397

Observation b4c509bb-e103-454c-b7a1-5aa47847ca92 · outbound

This paper cites Improving fac- tuality and reasoning in language models through multiagent debate.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Improving fac- tuality and reasoning in language models through multiagent debate

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:17.883268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:17.883268Z digest=sha256:48e9bed4529d5dc52bce1b5c96c3230a5d3969d07b6feb041cb4bd1f00e31564

Observation a3087291-5e04-4242-b044-f3d893e28ac5 · outbound

This paper cites Self-refine: Iterative refinement with self-feedback.Advances in Neural Information Processing Systems, 36:46534–46594, 2023.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Self-refine: Iterative refinement with self-feedback.Advances in Neural Information Processing Systems, 36:46534–46594, 2023

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:17.974595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:17.974595Z digest=sha256:9f36e2d0f8a4c2ac26122a9a09c9bcbb24ac67582dc04f163b9649918626fdfa

Observation c90b2b0d-60c0-4ace-864f-e2afc45173c2 · outbound

This paper cites Dyflow: Dynamic workflow framework for agentic reasoning.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Dyflow: Dynamic workflow framework for agentic reasoning

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:18.050316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:18.050316Z digest=sha256:8ac8539c7c34486097e2971a6fe61c94d1b2bfdae0fb18af180710191b8f876e

Observation fdf1ecc5-6a52-428f-b5b2-d3d115afe017 · outbound

This paper cites Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:18.133257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:18.133257Z digest=sha256:86984267416c95d7d54f3a503c617d32c22fdbea308ad96b7d82e69592295004

Observation 4f5ebcac-d920-474e-bc80-35af9b5760a6 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Direct preference optimization: Your language model is secretly a reward model

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:18.247145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:18.247145Z digest=sha256:87f010a00b80ea4c1e0cb70c1b4dd3a7e6ebe72f92d15e77ae131d064c01f913

Observation 3cf05bba-a277-4da6-91e2-2a1a66b57aa0 · outbound

This paper cites Glucose: Generalized and contextualized story explanations.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Glucose: Generalized and contextualized story explanations

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:18.315743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:18.315743Z digest=sha256:dca0feb284e594fa47454456eade946aa4c902abdc18b8e032b4fa42c7945ebd

Observation 9d30a6f0-1c1a-41be-95d2-18e1f0d21c23 · outbound

This paper cites Piqa: Reasoning about physical commonsense in natural language.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Piqa: Reasoning about physical commonsense in natural language

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:18.418665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:18.418665Z digest=sha256:83de705f63323ff677b469b7131e9c51175a6c8019dea7651700af14da6f23d8

Observation a74cdb97-f283-4414-8db2-991dbf3c0f13 · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:18.491278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:18.491278Z digest=sha256:07fa30d90066fc73ab2bb424fd36008d389f0161c8048df8e29f2e6db60c485a

Observation 75b2f6fc-b31c-4bdc-8509-7bafd5e238d5 · outbound

This paper cites Theory of mind may have spontaneously emerged in large language models.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Theory of mind may have spontaneously emerged in large language models

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:18.564499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:18.564499Z digest=sha256:f8545d2dfe35a2f7973635580615cc3eea990b47ee28512abebf89599a3d2415

Observation 10c98bb8-a181-4f77-b23d-126e62752d1b · outbound

This paper cites Breaking focus: Contextual distraction curse in large language models.arXiv preprint arXiv:2502.01609, 2025.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Breaking focus: Contextual distraction curse in large language models.arXiv preprint arXiv:2502.01609, 2025

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:18.658963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:18.658963Z digest=sha256:f8ddba98cda6eb4f939a2494efef9368c37ee3def8e0aa775cd7f9d347b69282

Observation 376a9404-d10c-4302-84ff-918afe3a131b · outbound

This paper cites FANToM: A Benchmark for Stress-testing Machine Theory of Mind in Interactions.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models FANToM: A Benchmark for Stress-testing Machine Theory of Mind in Interactions

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:18.777761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:18.777761Z digest=sha256:840c2ba0e0ac906f0eb9aaf746a02b4a9de7c7de1d5804d492debe0a693540de

Observation 834c773b-beb2-451b-93fe-ddf686541c21 · outbound

This paper cites Tomvalley: Evaluating the theory of mind reasoning of llms in realistic social context.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Tomvalley: Evaluating the theory of mind reasoning of llms in realistic social context

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:18.867408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:18.867408Z digest=sha256:018c0264e2a3fd7eedb532df5334bea8f345eadf1b8ef38bbf72abd7cf01cc93

Observation 024dd819-cad6-4501-8a0a-ebaa415fc6e6 · outbound

This paper cites An Open Review of OpenReview: A Critical Analysis of the Machine Learning Conference Review Process.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models An Open Review of OpenReview: A Critical Analysis of the Machine Learning Conference Review Process

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:18.940764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:18.940764Z digest=sha256:a5ee3a8b572d6728c809d52fd5200d20fffddecd0aaf2c2c578f6a80d526ba73

Observation 760e7e7f-9606-494e-b50d-4690a4b86426 · outbound

This paper cites The Open Review-Based (ORB) dataset: Towards Automatic Assessment of Scientific Papers and Experiment Proposals in High-Energy Physics.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models The Open Review-Based (ORB) dataset: Towards Automatic Assessment of Scientific Papers and Experiment Proposals in High-Energy Physics

Reference 81

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:45:27.218520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T12:45:19.052441Z digest=sha256:b8d95d76845948c6aedaa8627fb2c52a50773c0fbdc0717f7aac95629fac6eaf

Observation 5f241ba2-4413-4f28-bf6e-d4d71f446b19 · outbound

This paper cites Superhuman ai for multiplayer poker.Science, 365(6456):885–890, 2019.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Superhuman ai for multiplayer poker.Science, 365(6456):885–890, 2019

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:19.122093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:19.122093Z digest=sha256:39f28d554e52c5f48537f402ebc6c632ed67dc39214421e39f31961a5716e266

Observation f76f0b62-ae4c-48e7-a4df-c9c64a52d911 · outbound

This paper cites Geolocation with real human gameplay data: A large-scale dataset and human-like reasoning framework.arXiv preprint arXiv:2502.13759, 2025.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Geolocation with real human gameplay data: A large-scale dataset and human-like reasoning framework.arXiv preprint arXiv:2502.13759, 2025

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:19.216652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:19.216652Z digest=sha256:15d2515e2aee738d1faafae31041efd541977a7669a9a5bf1bcc80fa6b4fbb3a

Observation 49ecd989-94bc-4019-b6cf-4a2b8c9a7f91 · outbound

This paper cites Word form matters: Llms’ semantic reconstruction under typoglycemia, 2025.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Word form matters: Llms’ semantic reconstruction under typoglycemia, 2025

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:19.323926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:19.323926Z digest=sha256:472b644241aae6e1cf26178e90697be5fb697fc3e56e08c5d096cd29de0d507f

Observation 3d23c6e0-dd3b-4975-9538-84b8fd3bba74 · outbound

This paper cites TextAttack: A Framework for Adversarial Attacks, Data Augmentation, and Adversarial Training in NLP.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models TextAttack: A Framework for Adversarial Attacks, Data Augmentation, and Adversarial Training in NLP

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:19.444402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:19.444402Z digest=sha256:39ba6ea24d28984000fa5b8aac1ce56a9dba012a3f72660d17475fbb8b139191

Observation 6d0967d5-4ab6-42ab-9711-1d6e63576fe6 · outbound

This paper cites TrustLLM: Trustworthiness in Large Language Models.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models TrustLLM: Trustworthiness in Large Language Models

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:19.569117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:19.569117Z digest=sha256:639cd3e78c0a8056897f2cf90ee98845980739d3268680ab61e0c18fb6312b77

Observation 0bf59a51-b52f-47b5-8970-cf99e84bf206 · outbound

This paper cites GPT-4o System Card.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models GPT-4o System Card

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:19.686427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:19.686427Z digest=sha256:ec1125b3eb937e13f48d9267021d405c32915ddc66a4f77eae804c0eca203e41

Observation ab8e9951-1361-4cbe-a916-4c539efab799 · outbound

This paper cites Gpt-4o mini: Advancing cost-efficient intelligence.https://openai.com/index/ gpt-4o-mini-advancing-\cost-efficient-intelligence/, 2024.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Gpt-4o mini: Advancing cost-efficient intelligence.https://openai.com/index/ gpt-4o-mini-advancing-\cost-efficient-intelligence/, 2024

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:19.775775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:19.775775Z digest=sha256:dcda1c0b2e40b63c5704c0e75180def2f52a127e12451164ec0b50ffb738d241

Observation 3fb4e4bc-8c02-4a0c-8ddf-2ea039558105 · outbound

This paper cites o3-mini.https://docsbot.ai/models/o3-mini, January 2025.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models o3-mini.https://docsbot.ai/models/o3-mini, January 2025

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:19.864247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:19.864247Z digest=sha256:20e1c5ebe4424a4a09a912049f8d0e4318cd2238ed6caed0e646cb13bcff95b8

Observation 06751ed8-7c43-42b4-80ae-1ac0f08b4b78 · outbound

This paper cites OpenAI o1 System Card.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models OpenAI o1 System Card

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:20.032059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:20.032059Z digest=sha256:f6e83fdcebe2ef916e2646a0a82443878de71e841dc1aa03bc3ef8971ef552e2

Observation 3ec7c936-bbd4-42c2-b9a7-5b26e2a44b73 · outbound

This paper cites Gemini 2.5 pro experimental.https://ai.google.dev/gemini-api/ docs/models, March 2025.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Gemini 2.5 pro experimental.https://ai.google.dev/gemini-api/ docs/models, March 2025

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:20.213072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:20.213072Z digest=sha256:bb57b7f0d9d8fe08e7a86c24de01ca84d4606e8fe1dfd8d163fac0b159af87ae

Observation caaf9613-11be-4b6b-a4c8-cd19e9351c51 · outbound

This paper cites Phi-4 Technical Report.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Phi-4 Technical Report

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:20.381718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:20.381718Z digest=sha256:96281c5ca83f943d3c49bf89f6eeab5a55b5efe6ddfc8b5207673942397ea09e

Observation 212e4679-bfb0-4684-89a4-2ab9909d1e2c · outbound

This paper cites Llama 3.1-8b.https://huggingface.co/meta-llama/Llama-3.1-8B, 2024.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Llama 3.1-8b.https://huggingface.co/meta-llama/Llama-3.1-8B, 2024

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:20.477779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:20.477779Z digest=sha256:926e5642e270d1e326d806c03332d9088e02d9baf3d4b9b592b40bb588db619a

Observation 6c4a0d9d-9bee-408b-b4f2-5047eab5104b · outbound

This paper cites Llama 3.3-70b.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Llama 3.3-70b

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:20.590936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:20.590936Z digest=sha256:e43d70e926cef7ec0a9ce6dc40e2adedeccf811ebdc79ac668abc0de4cb7b07a

Observation 933701dc-3182-4c36-9915-900439b7e4c9 · outbound

This paper cites Qwen2.5: A party of foundation models, September 2024.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Qwen2.5: A party of foundation models, September 2024

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:20.676823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:20.676823Z digest=sha256:c4d84c8feef240016c1db07ee3ddcb83e93affe01f510e1f1a84d8a84fba17c6

Observation 66236904-c4d5-445f-8756-cb64ae00842c · outbound

This paper cites Qwq-32b: Embracing the power of reinforcement learning, March 2025.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Qwq-32b: Embracing the power of reinforcement learning, March 2025

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:20.773301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:20.773301Z digest=sha256:7ce762046e3ca1772d178efcd193cb0e1b89e2bf61710122e6ee2a4ec1c9064b

Observation ea100ba9-af50-4419-b410-edca7141a3cd · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:20.998512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:20.998512Z digest=sha256:815d45a055a85f668590c01d0ab8e3e35b3436ba9f700b8eeb7b2897d4657cbe

Observation 56e6ff9c-e1ef-4bbc-a08e-b14f0fb68dc8 · outbound

This paper cites an unresolved cited work.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Unresolved cited work

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:21.224254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:21.224254Z digest=sha256:9d5bd3a4893f2952cae7ec17237ba87c07125256626a713edcdd440bfc636e5f

Observation 3f2ed3cf-d633-4a12-b608-799c816a1980 · outbound

This paper cites 𝑢 is Criminal.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models 𝑢 is Criminal

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:21.292902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:21.292902Z digest=sha256:327a2c56d308fe1d8c3ce878a9ccca2563ad1d85f815994e0a9e759b69774555

Observation bbdad187-d187-4f70-a823-edddd76df300 · outbound

This paper cites Player𝑣 says Player𝑢 is the criminal.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Player𝑣 says Player𝑢 is the criminal

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:21.376172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:21.376172Z digest=sha256:268c61225818e6859726c1a55f4d0ca1b8af60326b2adfe7656f20099ba6d6b3

Pith citing papers

Observation 2f9f7f8e-aeee-4438-bf4a-45904ddd17b0 · inbound

AI Awareness cites this paper.

AI Awareness SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models

Reference 161

Resolution
unresolved
no resolver link, observed 2026-08-16T10:19:51.138159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:19:51.138159Z digest=sha256:bef7ae899fa0c686e82a3de20069cbf32c6ec411ff5d8c9685d1b729e420b4e5

Observation 7646a186-5934-4176-8c13-bb112357b79c · inbound

Temporal-IRL: Modeling Port Congestion and Berth Scheduling with Inverse Reinforcement Learning cites this paper.

Temporal-IRL: Modeling Port Congestion and Berth Scheduling with Inverse Reinforcement Learning SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T18:29:57.188033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T18:29:57.188033Z digest=sha256:0a5479030c3f483970c59917a32e820517a5e78e98150d4302e6b3d150ff2c27

Observation bfb07b62-2001-4fcf-bbaa-f2081bb94580 · inbound

SocialGrid: A Benchmark for Planning and Social Reasoning in Embodied Multi-Agent Systems cites this paper.

SocialGrid: A Benchmark for Planning and Social Reasoning in Embodied Multi-Agent Systems SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:48:01.696600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-10T08:45:54.303143Z digest=sha256:a7bc8f21c0f8dd04379a14dbb126c375a03136c5fcaa4e2dd803202a0da5fdda

Observation af7acf9e-6748-4e34-83d5-a80c1cd97551 · inbound

GroupToM-Bench: Benchmarking Group Theory of Mind and Nonlinear Social Emergence in MLLMs cites this paper.

GroupToM-Bench: Benchmarking Group Theory of Mind and Nonlinear Social Emergence in MLLMs SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models

Reference 54

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T02:46:29.084004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-28T10:37:45.062718Z digest=sha256:f1fd7579502a6c30e99ecdc5314cfbf166d7ec045f2ff6dc8d0945d25ef17598

Observation fde092a4-5309-4edc-9509-25b25cc7dd86 · inbound

OpenSkill: Open-World Self-Evolution for LLM Agents cites this paper.

OpenSkill: Open-World Self-Evolution for LLM Agents SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T13:56:59.840420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-28T00:47:38.201358Z digest=sha256:3622793177d895762b16eb5f970c891e3afa3444f447e5fc221cbc573c41fe69

Observation 0785db62-f8a9-470a-88ea-c06e0644d9b4 · inbound

Getting Better at Working With You: Compiling User Corrections into Runtime Enforcement for Coding Agents cites this paper.

Getting Better at Working With You: Compiling User Corrections into Runtime Enforcement for Coding Agents SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:28:31.390518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-27T07:04:51.970049Z digest=sha256:e6de792005d11b923988cb2f6b0bd64ec4e06b25fedf976c544078259d095748

Observation 650c9ba9-2d06-45ee-84df-fb206940328e · inbound

PIVOTSBench: Evaluating Fine-Grained Interpersonal Relationship Reasoning in Multimodal Large Language Models cites this paper.

PIVOTSBench: Evaluating Fine-Grained Interpersonal Relationship Reasoning in Multimodal Large Language Models SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models

Reference 1

Resolution
malformed identifier
arxiv_id, observed 2026-07-04T10:39:44.819288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-26T08:42:26.713614Z digest=sha256:8a7e615714affcaa4cb4cf278a6313f18826e0b15b0fb97072cca10a0818bcc2

Observation 24a806cd-70cc-405e-9802-a554c05f7c9b · inbound

Inside the Unfair Judge: A Mechanistic Interpretability Account of LLM-as-Judge Bias cites this paper.

Inside the Unfair Judge: A Mechanistic Interpretability Account of LLM-as-Judge Bias SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models

Reference 218

Resolution
unresolved
no resolver link, observed 2026-07-14T02:33:34.084111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T02:33:34.084111Z digest=sha256:826d343e2096851dfa016337880c9c419fe2ce0b9772e5025b3e1d2b14bd8313

Observation 7a631bd5-d10f-4067-a659-f09cd98f4aea · inbound

No One Wins in Nuclear War: A Social Simulation of Military Decision-making cites this paper.

No One Wins in Nuclear War: A Social Simulation of Military Decision-making SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T19:16:48.882711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T19:16:48.882711Z digest=sha256:f8ad54729a5074b49a59c4c23dfdd7fad8db049e1cd8139bf0b4c4638f0f9806