Pith. sign in

Paper Citation Record · LEDGER

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments

As of 14 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 0 inbound Pith citation observations for arXiv:2608.09128.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.09128 v1

Coverage vector

measured 68 of 68 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T23:05:34.799690Z

measured 68 of 68 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

68 of 68 outbound references displayed

  • verified exact0
  • verified fuzzy31
  • unresolved37
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 66099bef-a43e-4607-8aed-e7f68deef70a · outbound

This paper cites Acta Universitatis Sapientiae, Informatica , volume=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Acta Universitatis Sapientiae, Informatica , volume=

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.528855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.540529Z digest=sha256:fe436cdc84d550301d29713eff19fba51bfcd2e7f929a5776266e43b9c6a786b

Observation 3daf168c-8bbb-42d9-aeae-cfefd8631438 · outbound

This paper cites NAACL , year=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments NAACL , year=

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.516410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.546631Z digest=sha256:acf15e144e2cdbbc8799b240a40f06542f67799eb0f42272e814f42988229f5d

Observation ce18d556-0ae4-4e64-9da2-f90757690548 · outbound

This paper cites 2026 , booktitle=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments 2026 , booktitle=

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.502952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.550779Z digest=sha256:6dddb30afc67eef9e9402e5bcb5da4b46398563e3db7d4569dd7f41f9985e663

Observation 29ab7034-16d0-42d2-8406-7e74c1caee57 · outbound

This paper cites Findings of ACL , year=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Findings of ACL , year=

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.489237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.554824Z digest=sha256:1807f487fc71257d07b1bf2ec4910a4abcdf3eba9fc4820ede72cedfe36a0b1c

Observation 594e5ffc-4f00-47f3-bef2-f6cf10b0011c · outbound

This paper cites 2024 , booktitle=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments 2024 , booktitle=

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.477232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.559522Z digest=sha256:85f231b0b7e3ec6442c627b34c6a670dc38b0f50d254957ac327ec138f58caad

Observation 14d77b26-a516-41de-a5a5-469d56fcc0f1 · outbound

This paper cites 2025 , eprint=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments 2025 , eprint=

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.465794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.563464Z digest=sha256:64adc44dafa9e75b9d56db5a4f2f134018fe65d080fa581d2f9b78a7f3f9bd2e

Observation d74d4f7d-84bb-45ba-b997-44c7be788b18 · outbound

This paper cites Advances in neural information processing systems , volume=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Advances in neural information processing systems , volume=

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.567592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.567592Z digest=sha256:bf8273f48d329cc543eb0a907c57797fdac5d96df5bcac2d1f0b3f21b27e103a

Observation 6e832812-f636-4569-b8b4-254b609b2c2f · outbound

This paper cites 2025 , month = may, note =.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments 2025 , month = may, note =

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.445925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.571679Z digest=sha256:85d9b64d0b9f17fb6935f08a05f714af4b0f4871eb74dca433d59083318efa71

Observation 89396d3a-c612-4fdf-a82c-907dfb2ad00e · outbound

This paper cites Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages=

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.575440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.575440Z digest=sha256:97bb25b68ce22c8f1cb5a6aa40fb3efbda32764c7b7d80c49ebefdfdffb721f9

Observation 52c6cd20-333f-404f-b795-6f83819e7cf6 · outbound

This paper cites Societal Alignment Frameworks Can Improve LLM Alignment.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Societal Alignment Frameworks Can Improve LLM Alignment

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.579870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.579870Z digest=sha256:e49539e0543eab729c50b73ca680777a7ece7d7e7cf6875613999e1dc201a9df

Observation 3675d94c-a18b-443c-9102-a62a273b4660 · outbound

This paper cites 2005 , publisher=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments 2005 , publisher=

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.584416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.584416Z digest=sha256:dbf220676afd006c44d69aabdb83604558e3de6bef32e2b0e0345d33f649b973

Observation 8b79025d-e39c-446b-94b7-b8c4220a931c · outbound

This paper cites Concrete Problems in AI Safety.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Concrete Problems in AI Safety

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.588577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.588577Z digest=sha256:8fe46894b836d5e68e1d8cb2ea68576997c7d94066455feac8b8cca40ff5328f

Observation d894df8e-ee2f-42df-9b5e-1ddcdea9fabb · outbound

This paper cites The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.592692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.592692Z digest=sha256:6b06c135722dcf3f61610ccb19968c74a54f2e4aadf6d5dcdde1acbbb2e54f44

Observation e58cea66-1fc2-4454-862b-38e8fe86ffcc · outbound

This paper cites Artificial intelligence , volume=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Artificial intelligence , volume=

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.597498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.597498Z digest=sha256:82890cb4389774033b16239d1b09c3effce53ed5198561bb8bd0cba459a7d66b

Observation 45bf8575-057e-4261-a378-331508fe4b8b · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.601259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.601259Z digest=sha256:e2d215e2c0e9c8589eb348926e445f4132f59c78dce56f0ed9b6ae0b75f425ef

Observation da6f06b8-be26-4c03-ac42-45b7dc6c5ee9 · outbound

This paper cites The Llama 3 Herd of Models.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments The Llama 3 Herd of Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.605581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.605581Z digest=sha256:8f927079e936e613c072f276d7be5a0fbf283a81519d22d93518b1ac539bf50a

Observation e9625301-19d0-46f4-a001-fee5262eba8c · outbound

This paper cites 2024 , eprint=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments 2024 , eprint=

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.609259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.609259Z digest=sha256:eb2789277e99eeb2ce6200d9367c63b077d2d9f0beb235ab13ab9400352fe583

Observation b0f12945-cad7-4cf5-8ef4-1249fd64fc6a · outbound

This paper cites 2025 , eprint=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments 2025 , eprint=

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.398675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.613177Z digest=sha256:4ae35c4b210f99b5710eb5164a855dc3a668bfa57523a63bbe9885ef314b7477

Observation f7dac24d-6f34-45eb-a0b2-19621337f3a0 · outbound

This paper cites 2025 , eprint=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments 2025 , eprint=

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.385597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.616839Z digest=sha256:b3ffa3012834835f1247eeef49640926e45f5d94b08a72a9610fa58474231542

Observation 954e4e6b-ef1d-49a7-b228-086e3cca04e2 · outbound

This paper cites 2026 , eprint=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments 2026 , eprint=

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.371792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.620698Z digest=sha256:e6206b8c89052e43aeadc038403c6b2749ec61cbf658170827d131b70edb7fe9

Observation 3d011c37-78ee-4b60-bb73-3d9b00f803b9 · outbound

This paper cites 2025 , eprint=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments 2025 , eprint=

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.624472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.624472Z digest=sha256:4c99badd3205c99dce6d865e700cfeaebe6daa75e28372e194f09bea4b5e5695

Observation 84d06b87-c67f-491f-b547-53ee684b4f78 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Advances in Neural Information Processing Systems , volume=

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.628227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.628227Z digest=sha256:669691608b4f04980dc53fbe2d1dcac7b6750b35840717bd349724fb78b9f4a7

Observation 773ede4d-d5d0-4f6a-ae28-e37296867b06 · outbound

This paper cites Proceedings of the 36th annual acm symposium on user interface software and technology , pages=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Proceedings of the 36th annual acm symposium on user interface software and technology , pages=

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.631953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.631953Z digest=sha256:82b01b348fcd63e166d53fdafb461f568735fba236e597e1e9220247eff58eff

Observation c8645620-ea0e-4777-b1cc-ccc40dd6a6ee · outbound

This paper cites Proceedings of the National Academy of Sciences , volume=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Proceedings of the National Academy of Sciences , volume=

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.635924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.635924Z digest=sha256:15ea2599158948421f947cd70d238195a001132e1699a276ad62b11ca01c6dcb

Observation 2894105d-3f0e-4a29-aa5b-eddbeedf496f · outbound

This paper cites Science , volume=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Science , volume=

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.329977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.639702Z digest=sha256:9b1d7dead62042d33f2292fd9adca63258d6631496b59c5d8bc4e7e691f4e6dd

Observation bbe5a7ab-7a84-44d4-a768-0abee4856344 · outbound

This paper cites Exploring Large Language Models for Communication Games: An Empirical Study on Werewolf.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Exploring Large Language Models for Communication Games: An Empirical Study on Werewolf

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.643711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.643711Z digest=sha256:b51ee7707e9cd9c4009d9a06bd9e8e5c03c9a6d80b408c5fd4ef4ceb8cb78b23

Observation 975706fe-9bd4-4fc4-81a7-b86d4ce42419 · outbound

This paper cites Nature Human Behaviour , volume=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Nature Human Behaviour , volume=

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.647924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.647924Z digest=sha256:144b4c53641c574882ae31392fdc3b6dd4cbd80948a958dfba12d2712736fcf7

Observation 4454e18a-3ab3-4ea1-98bf-c82f13f29c35 · outbound

This paper cites Advances in neural information processing systems , volume=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Advances in neural information processing systems , volume=

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.651661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.651661Z digest=sha256:e0972b90761d363dc5a2f887cf93ec4528e585683627d83f66360ccd183203f9

Observation bdf964a1-19f4-4463-a7e5-9bb85bd572c8 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Advances in Neural Information Processing Systems , volume=

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.655444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.655444Z digest=sha256:7b4c742a7bd934fea911ed4df3f2ea724881202f8a9edfb562c9b074003f8b2f

Observation 77601210-e2c9-4ae2-a34a-c04fb698737a · outbound

This paper cites Advances in neural information processing systems , volume=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Advances in neural information processing systems , volume=

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.659162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.659162Z digest=sha256:5ad7503cb56a94ad45d4736653f5c57f318166efd78e579f04319177b713a21f

Observation 17b30183-258b-43a1-985e-8288a569860e · outbound

This paper cites LLM-Based Agent Society Investigation: Collaboration and Confrontation in Avalon Gameplay.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments LLM-Based Agent Society Investigation: Collaboration and Confrontation in Avalon Gameplay

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.662836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.662836Z digest=sha256:18725112633446c89c60bcb4db43a4492af4599e8fd82d2efd1c0cdedcebe296

Observation 127bc410-80ff-4c93-a5df-4d0663cca398 · outbound

This paper cites Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.666916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.666916Z digest=sha256:48b564ed6a348c413a96d7fc6771f9208d808639c45a84107deebc3079f7937e

Observation 496b2712-ea19-4b53-a284-9e8520e591f6 · outbound

This paper cites 1978 , publisher=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments 1978 , publisher=

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.281839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.670736Z digest=sha256:66767375ac8da8c2b964facd436fc849e8c3316ab674c35edb2442ee4c1fbe0c

Observation a19a1aeb-11ce-43c4-a3f4-6b1ea2129004 · outbound

This paper cites International Conference on Machine Learning , pages=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments International Conference on Machine Learning , pages=

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.270011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.674770Z digest=sha256:6f556a3eb7464306994569f9405daffd2a9e11c96402696ee3bed9f2f597f1cf

Observation 8dbaa87e-8e9a-4f7d-8174-a8fb22c37e6e · outbound

This paper cites the method of paired comparisons , author=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments the method of paired comparisons , author=

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.678632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.678632Z digest=sha256:2e3fcb24581b352b29c005227e70115b161d0c418548afe7582fad5130f8c371

Observation bfa86fc2-5cce-4729-a821-5265ed3d9a80 · outbound

This paper cites Harper's Magazine , volume=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Harper's Magazine , volume=

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.251302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.682484Z digest=sha256:7c1f71455924851c3cb11718113bd81e9a5bc83d928c8a586c07d8109ebcbcee

Observation 5e5c370e-7aab-42ea-b054-d79e08c65fed · outbound

This paper cites 2011 , publisher=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments 2011 , publisher=

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.686119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.686119Z digest=sha256:a079f2265405824869a66c892a8331aae2be0ec3f23e90093cbf009940575ab7

Observation e78b11c9-eff8-4e9c-a444-ab8939b2b668 · outbound

This paper cites science , volume=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments science , volume=

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.231837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.689919Z digest=sha256:5b76e60145aa23d71b12bf77da69dc0a5509c90bb6e53e38fa5726c6001255e0

Observation 70a90aee-3bb6-4f08-9859-8440827bb1d2 · outbound

This paper cites 1980 , publisher=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments 1980 , publisher=

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.219880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.693614Z digest=sha256:db3d31caab35e2fccaf9284ebc6996d193b92b9fa7660357968d52de6d8bae6e

Observation 9923f44d-4655-4f10-92c0-bff98c3bd522 · outbound

This paper cites 1990 , publisher=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments 1990 , publisher=

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.697208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.697208Z digest=sha256:9b2693d0c37f10b0f53ae04f38b5b059946d5d3140531c1facc03d975522b082

Observation 72eece5e-e793-4ef5-ba53-1c448cd309d5 · outbound

This paper cites American Economic Review , volume=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments American Economic Review , volume=

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.199317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.700971Z digest=sha256:83049aa75c7c189ee5ea78f5cd23b3cd72018583974520e2c80051195d8b342e

Observation 5e426b3f-8ced-4adf-9d53-be49d5a8976c · outbound

This paper cites Journal of Economic theory , volume=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Journal of Economic theory , volume=

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.186977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.704798Z digest=sha256:adb5d3853c0d31c1d3dd24e374569ca1f6e63df64b454713b0ca35c0b9a40f88

Observation d61c3931-0e43-4924-b939-4beba99641db · outbound

This paper cites Econometrica: Journal of the Econometric Society , pages=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Econometrica: Journal of the Econometric Society , pages=

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.174422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.708341Z digest=sha256:b9d77aadc2f5895b16e6e39c8657f800aa8603360e7fb00cd0f9de9e81a4270c

Observation 3b7977d3-78c1-4106-970a-0b4540b5b125 · outbound

This paper cites , author=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments , author=

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.160685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.711994Z digest=sha256:b5d71d73b80ed79bd9b42ef5f3e947fd2f55b24990ea056cff1699501d2f8dd0

Observation 8fe176d4-dc5e-4082-8cc1-b7d4313b7eb5 · outbound

This paper cites John thinks that Mary thinks that….

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments John thinks that Mary thinks that…

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.715483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.715483Z digest=sha256:f38f42aa2f9991b83fa0a98d99b642c60d5f2afcc3141125f6a6280d4b5dbe84

Observation 156be659-4703-4394-bce9-251668320b51 · outbound

This paper cites Evolutionary Anthropology: Issues, News, and Reviews , volume=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Evolutionary Anthropology: Issues, News, and Reviews , volume=

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.147997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.721296Z digest=sha256:17ef8f1cf1d63539e1624280940abda5e5eb66527d18b3f4271b6277b2415656

Observation 191024f9-c656-491a-a870-4531d69f5959 · outbound

This paper cites 1984 , publisher=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments 1984 , publisher=

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.136750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.725068Z digest=sha256:77e05ff7be9ea7243813b3a9f05fd185203aa2ab56c2a9e75152150a35218e16

Observation c96f0ba7-6dc2-493b-beb6-1dc1959b545d · outbound

This paper cites Cognition , volume=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Cognition , volume=

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.125061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.728658Z digest=sha256:9874cc607119856334184348db1c20660ad7eff09b09deeb849dc82a545a8098

Observation dbe2ea54-809a-4c1a-919d-03136eda64ba · outbound

This paper cites Cognition , volume=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Cognition , volume=

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.111809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.731838Z digest=sha256:65f3e0d7f0a48aa0e0ce24c25ff4e9ed67bc95d3b7252d678b6d2f408d72bc79

Observation 519e20e8-bb82-4bca-b07b-8c09befafff7 · outbound

This paper cites 1988 , publisher=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments 1988 , publisher=

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.099667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.735336Z digest=sha256:59e244fad3de208fac29e50d1e35a398dcc97aa9e7202c640f3b6899146fbd85

Observation 2a5a11b2-97c1-415a-9ec8-739f58cfd8bf · outbound

This paper cites an unresolved cited work.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Unresolved cited work

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.738866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.738866Z digest=sha256:2fe1de44d644bc0eff8c56fb8c3cee70d1b3f05943029b73ef4a48c2882da52d

Observation 8efb1f2c-664e-4b59-b3d8-904f78ce0a79 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Advances in Neural Information Processing Systems , volume=

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.742161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.742161Z digest=sha256:89b9a3056a16513d8025a74bc6ef91b6c995de34cc85a13d4ea0d929540290ce

Observation 9b046e81-6669-4544-82ea-8f22a48688d7 · outbound

This paper cites Proceedings of the 2022 conference on empirical methods in natural language processing , pages=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Proceedings of the 2022 conference on empirical methods in natural language processing , pages=

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.073627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.745589Z digest=sha256:9d4dc26b240235374335301d3df0e80c0d426868b8a05fe8e8abde73929d1137

Observation 0bdd7741-8875-4777-b65a-8942b19cddd3 · outbound

This paper cites Evaluating Large Language Models in Theory of Mind Tasks.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Evaluating Large Language Models in Theory of Mind Tasks

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.748777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.748777Z digest=sha256:f4de53bf261cd89dded15ccabc24f05e9ebb0b9b74508a1cd502cf54ebaf7be4

Observation bfd5295b-c9a8-40a6-878a-bc1bde0975ae · outbound

This paper cites Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.752241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.752241Z digest=sha256:353e36dbfa3d5ef04325264a2833fddd94d247eed2b5baf970c0c168f27df10b

Observation f800c627-c57e-4d71-9de6-51a68e340350 · outbound

This paper cites AvalonBench: Evaluating LLMs Playing the Game of Avalon.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments AvalonBench: Evaluating LLMs Playing the Game of Avalon

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.755550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.755550Z digest=sha256:f3c0a50a65406c11c8f9fa68830c6a64b48cb05a5ae9ae7c29f83ff9dddc8ed7

Observation 6f7e96a1-d04a-46b3-8b66-362b221ffcdd · outbound

This paper cites GTBench: Uncovering the Strategic Reasoning Limitations of LLMs via Game-Theoretic Evaluations.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments GTBench: Uncovering the Strategic Reasoning Limitations of LLMs via Game-Theoretic Evaluations

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.759492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.759492Z digest=sha256:281a4d2bc07dd4302e8f8ea2749c0ecdcea447d8277a15553a0f12d62360cc14

Observation a9219dcf-52df-49df-b870-ed3972b466cb · outbound

This paper cites International Conference on Learning Representations , volume=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments International Conference on Learning Representations , volume=

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.061007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.764350Z digest=sha256:8ccdcc7d36c8abf906149e777758970ca40a72ba4f0292f1a7ec8442e08d7bef

Observation ffbaccbe-f001-4bae-aca2-42e1d6af22eb · outbound

This paper cites Voyager: An Open-Ended Embodied Agent with Large Language Models.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Voyager: An Open-Ended Embodied Agent with Large Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.767877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.767877Z digest=sha256:3579a745d1652e4841830dd7d582f85ba177b8c07464e3aae16c8bddf1ab1ce2

Observation 0d998186-d093-463f-b125-9c66efc035c4 · outbound

This paper cites Proceedings of the AAAI Conference on Artificial Intelligence , volume=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Proceedings of the AAAI Conference on Artificial Intelligence , volume=

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.771496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.771496Z digest=sha256:9006d97525ca12c026a1b3d2a6f5448f67a4ec6b6e0815ce3078a4b334b48ab9

Observation 3fda412d-e11f-4999-9d0e-50d33c78b3ed · outbound

This paper cites Handbook of Intelligence , editor=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Handbook of Intelligence , editor=

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.039264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.774899Z digest=sha256:1bd14b3ec772614e0fbbe78d68574200159389b88fb64ff2a3622e8254ec173f

Observation 81347811-66ea-4fea-af28-c0484b9fe772 · outbound

This paper cites 2025 , eprint=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments 2025 , eprint=

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.778497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.778497Z digest=sha256:1d47eaed935fd140988265cbb2792303fb25e0c4b9cb42c81e665644e8ef8a4b

Observation e596d2f7-e9ad-4f6c-ac81-4781c8d7da5d · outbound

This paper cites 2023 , eprint=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments 2023 , eprint=

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.781939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.781939Z digest=sha256:3777fffc32674ffd809f7f3eceae55c3301265d117a0a380399007845f57ce61

Observation 8898c9df-a433-4473-89cc-42766f86ecae · outbound

This paper cites Proceedings of the 41st International Conference on Machine Learning , pages=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Proceedings of the 41st International Conference on Machine Learning , pages=

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.011999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.785370Z digest=sha256:b238bd94c00bd7029cf5fc699b9ffd65ca2e77e09db3945700e843a6e2d04f44

Observation 5c155570-937a-4238-9a6d-c70946b3ad1f · outbound

This paper cites Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages=

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.788809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.788809Z digest=sha256:482872f0fe89b257be5741658b0d08a6435173e40a6d7350e39ceff1eeb8c6d1

Observation 2cbbefed-68ef-4c45-9eea-87278b528cb0 · outbound

This paper cites an unresolved cited work.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Unresolved cited work

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.792530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.792530Z digest=sha256:531dc64aafed590f32e1c382a5570f687346184e867f702cedf835cd70ee3cf3

Observation 77bc36d6-7d6d-482c-a4b2-f9903d3755ec · outbound

This paper cites 2024 , publisher=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments 2024 , publisher=

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:34.991239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.796172Z digest=sha256:75e2ca21f7c92a63aef28e0974c1ebf7dc79d7a4c756c572593b9b156bf5a4d3

Observation 89df69cd-4405-40d8-a20b-7102f040f3d0 · outbound

This paper cites Journal of Consumer Research , volume=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Journal of Consumer Research , volume=

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:34.979116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.799690Z digest=sha256:ad31a5023a76007fa75a5b681b1da0ed885dbadc143d8353a05ea24e5ff7fd9f

Pith citing papers

No inbound Pith citation observations are available.