Pith. sign in

Paper Citation Record · LEDGER

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments

As of 22 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 0 inbound Pith citation observations for arXiv:2608.09128.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.09128 v1

Coverage vector

measured 68 of 68 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T23:05:34.799690Z

measured 68 of 68 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

68 of 68 outbound references displayed

  • verified exact0
  • verified fuzzy31
  • unresolved37
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 66099bef-a43e-4607-8aed-e7f68deef70a · outbound

This paper cites Acta Universitatis Sapientiae, Informatica , volume=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Acta Universitatis Sapientiae, Informatica , volume=

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.528855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.540529Z digest=sha256:42286bae2b0a92a03c65d48ba1931ed9b02ac245d37e377d36feb816235b0ec5

Observation 3daf168c-8bbb-42d9-aeae-cfefd8631438 · outbound

This paper cites NAACL , year=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments NAACL , year=

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.516410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.546631Z digest=sha256:1b8d602e34a351c5df757e4d8dc779c966176fc33e79a8b25c3bece2996fe88d

Observation ce18d556-0ae4-4e64-9da2-f90757690548 · outbound

This paper cites 2026 , booktitle=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments 2026 , booktitle=

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.502952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.550779Z digest=sha256:43158133d73dde3a85c137b964c5db62ae7035eb1baf868099e8c18721d998b0

Observation 29ab7034-16d0-42d2-8406-7e74c1caee57 · outbound

This paper cites Findings of ACL , year=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Findings of ACL , year=

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.489237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.554824Z digest=sha256:5c8da281e5bf77ad41e793de80ceb9cd67c0d55b4d4bbfd8655ec6273caf32f3

Observation 594e5ffc-4f00-47f3-bef2-f6cf10b0011c · outbound

This paper cites 2024 , booktitle=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments 2024 , booktitle=

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.477232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.559522Z digest=sha256:1d81297b00a5143ba9bf70359d2a697e195a87e875eee0d563e148bb0f4769f9

Observation 14d77b26-a516-41de-a5a5-469d56fcc0f1 · outbound

This paper cites 2025 , eprint=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments 2025 , eprint=

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.465794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.563464Z digest=sha256:2d0fa9aa2bf0a50b3ba75553cb86ee8c55711167ce3dbf4ce6a879cd05ee8c1a

Observation d74d4f7d-84bb-45ba-b997-44c7be788b18 · outbound

This paper cites Advances in neural information processing systems , volume=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Advances in neural information processing systems , volume=

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.567592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.567592Z digest=sha256:709dc6198f362d85efabd609a2e029e547f01707f483ca04da7e2aa1cfd17d7e

Observation 6e832812-f636-4569-b8b4-254b609b2c2f · outbound

This paper cites 2025 , month = may, note =.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments 2025 , month = may, note =

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.445925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.571679Z digest=sha256:472d4e0860d7a27cb6ff9d7ed1dcc3bc0da0fa9cd252f5ee75362a888dfa9dda

Observation 89396d3a-c612-4fdf-a82c-907dfb2ad00e · outbound

This paper cites Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages=

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.575440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.575440Z digest=sha256:673d5733f8603f25ea73d9ecf062830041685a4a4769910842474973aa60522d

Observation 52c6cd20-333f-404f-b795-6f83819e7cf6 · outbound

This paper cites Societal Alignment Frameworks Can Improve LLM Alignment.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Societal Alignment Frameworks Can Improve LLM Alignment

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.579870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.579870Z digest=sha256:ae8aeeabaf9baf49f177312f5d57f17465b6e5b6bc32dfb9b7810a5b094ed51e

Observation 3675d94c-a18b-443c-9102-a62a273b4660 · outbound

This paper cites 2005 , publisher=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments 2005 , publisher=

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.584416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.584416Z digest=sha256:0af2a08baf136d67721a56bee0ad0a59c52351f9fbeb300a87ba401b4d66d2a2

Observation 8b79025d-e39c-446b-94b7-b8c4220a931c · outbound

This paper cites Concrete Problems in AI Safety.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Concrete Problems in AI Safety

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.588577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.588577Z digest=sha256:1b643f96f70b2875918c15276c3a636d5ee88c8d3f4d62ce4ab91d738a722925

Observation d894df8e-ee2f-42df-9b5e-1ddcdea9fabb · outbound

This paper cites The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.592692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.592692Z digest=sha256:1bff7881912fc3380d3e32c84f32d1f351314af5bda6245706ff184547e7db98

Observation e58cea66-1fc2-4454-862b-38e8fe86ffcc · outbound

This paper cites Artificial intelligence , volume=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Artificial intelligence , volume=

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.597498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.597498Z digest=sha256:9fc5f57a00d913d4178be52f96f00afc52bfb75b4f9711822b019cdaa5edafff

Observation 45bf8575-057e-4261-a378-331508fe4b8b · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.601259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.601259Z digest=sha256:45b65a826235930b88aaadf289e7dca558becc1729be42f3d993c9e10da65096

Observation da6f06b8-be26-4c03-ac42-45b7dc6c5ee9 · outbound

This paper cites The Llama 3 Herd of Models.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments The Llama 3 Herd of Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.605581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.605581Z digest=sha256:5718872602ebd0538982ace5655dca59ee04cc577bf132956629078314ca485c

Observation e9625301-19d0-46f4-a001-fee5262eba8c · outbound

This paper cites 2024 , eprint=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments 2024 , eprint=

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.609259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.609259Z digest=sha256:bba2346e6139ebd4ee455da339ce2a12adc9a590c165b54de75078afbbb292f9

Observation b0f12945-cad7-4cf5-8ef4-1249fd64fc6a · outbound

This paper cites 2025 , eprint=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments 2025 , eprint=

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.398675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.613177Z digest=sha256:1a91affb6f6d85bb40e8f41b38e505c0cf16bcb71d2ee5a4e2f8e90bce266484

Observation f7dac24d-6f34-45eb-a0b2-19621337f3a0 · outbound

This paper cites 2025 , eprint=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments 2025 , eprint=

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.385597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.616839Z digest=sha256:b9b40b0f0ead892e4c49d954ce141643f7019055cd72e31aa5f6e14a82862794

Observation 954e4e6b-ef1d-49a7-b228-086e3cca04e2 · outbound

This paper cites 2026 , eprint=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments 2026 , eprint=

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.371792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.620698Z digest=sha256:63d2880b8e5a56a7249ded011ded4a6b76e722e872fa7457605ba96ae9d58d58

Observation 3d011c37-78ee-4b60-bb73-3d9b00f803b9 · outbound

This paper cites 2025 , eprint=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments 2025 , eprint=

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.624472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.624472Z digest=sha256:030b22422f7ca8343f063d3976d248ce6bf2cb2dac56ba41a723431f2c82512c

Observation 84d06b87-c67f-491f-b547-53ee684b4f78 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Advances in Neural Information Processing Systems , volume=

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.628227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.628227Z digest=sha256:5a25a8bcd7cec7c02dedab54609df97e431e2a13eed9c2137eb3997005c5fb16

Observation 773ede4d-d5d0-4f6a-ae28-e37296867b06 · outbound

This paper cites Proceedings of the 36th annual acm symposium on user interface software and technology , pages=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Proceedings of the 36th annual acm symposium on user interface software and technology , pages=

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.631953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.631953Z digest=sha256:92e36224e9492f7d3afea3db55ca46967448e75438ddec412d26884eb8452e09

Observation c8645620-ea0e-4777-b1cc-ccc40dd6a6ee · outbound

This paper cites Proceedings of the National Academy of Sciences , volume=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Proceedings of the National Academy of Sciences , volume=

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.635924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.635924Z digest=sha256:d939b9f296d2af37e0d8cb33f1361a87571ca26f1ca27faf0043cc8c7dce4033

Observation 2894105d-3f0e-4a29-aa5b-eddbeedf496f · outbound

This paper cites Science , volume=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Science , volume=

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.329977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.639702Z digest=sha256:2fb3f6b11b1ca95a8e38da76d9c57d38418db1c2166cf10ca71555befba6fd58

Observation bbe5a7ab-7a84-44d4-a768-0abee4856344 · outbound

This paper cites Exploring Large Language Models for Communication Games: An Empirical Study on Werewolf.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Exploring Large Language Models for Communication Games: An Empirical Study on Werewolf

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.643711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.643711Z digest=sha256:9cef750a6fd9df019a00584ff591b4097789612086a0be961ea94049dbc9e833

Observation 975706fe-9bd4-4fc4-81a7-b86d4ce42419 · outbound

This paper cites Nature Human Behaviour , volume=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Nature Human Behaviour , volume=

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.647924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.647924Z digest=sha256:3bc2e9a216151842558868a7855d506d65e35fcd4861c504e9ec1e31da29080a

Observation 4454e18a-3ab3-4ea1-98bf-c82f13f29c35 · outbound

This paper cites Advances in neural information processing systems , volume=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Advances in neural information processing systems , volume=

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.651661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.651661Z digest=sha256:d7f0b7806651cd3a7a9a9f743efc3d0daded39152eedf6831b78e601b232f202

Observation bdf964a1-19f4-4463-a7e5-9bb85bd572c8 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Advances in Neural Information Processing Systems , volume=

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.655444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.655444Z digest=sha256:5ab579589248cfd3c1b2f000c13258e94d923d6da7198130132b03afd45a12d8

Observation 77601210-e2c9-4ae2-a34a-c04fb698737a · outbound

This paper cites Advances in neural information processing systems , volume=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Advances in neural information processing systems , volume=

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.659162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.659162Z digest=sha256:52a8f2f2d245843707a4c7d772f6172dcb0d7f5faa3dbfa43ab5fc39112d537e

Observation 17b30183-258b-43a1-985e-8288a569860e · outbound

This paper cites LLM-Based Agent Society Investigation: Collaboration and Confrontation in Avalon Gameplay.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments LLM-Based Agent Society Investigation: Collaboration and Confrontation in Avalon Gameplay

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.662836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.662836Z digest=sha256:a6d2c7a6e08b92be60193f78985a28c90729c749227c3bd85c97f4f422d7889e

Observation 127bc410-80ff-4c93-a5df-4d0663cca398 · outbound

This paper cites Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.666916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.666916Z digest=sha256:b53ab217379a4f5404e13d5a5498ee8b6c4203110fa69d195865e268a993c453

Observation 496b2712-ea19-4b53-a284-9e8520e591f6 · outbound

This paper cites 1978 , publisher=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments 1978 , publisher=

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.281839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.670736Z digest=sha256:d8e38161d74bb10e926dd69e3f93b6f6f533bec0c88d5ba4bfcea40bdd3ce356

Observation a19a1aeb-11ce-43c4-a3f4-6b1ea2129004 · outbound

This paper cites International Conference on Machine Learning , pages=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments International Conference on Machine Learning , pages=

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.270011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.674770Z digest=sha256:bebb88a0bd64ddaec3b3767fd246daa32636132c63da6a5d18b3016125f99953

Observation 8dbaa87e-8e9a-4f7d-8174-a8fb22c37e6e · outbound

This paper cites the method of paired comparisons , author=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments the method of paired comparisons , author=

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.678632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.678632Z digest=sha256:64444d4013f7829b097261e4b8989d4e47440bc5b6fb5145b57538ee3a1244c1

Observation bfa86fc2-5cce-4729-a821-5265ed3d9a80 · outbound

This paper cites Harper's Magazine , volume=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Harper's Magazine , volume=

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.251302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.682484Z digest=sha256:c272c9ce36b491d45a3cf0e5d080211bd56c5310bca57acf56a2a24aa4e6a047

Observation 5e5c370e-7aab-42ea-b054-d79e08c65fed · outbound

This paper cites 2011 , publisher=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments 2011 , publisher=

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.686119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.686119Z digest=sha256:bb511408cd5fff0cff0810db3b6307b44f8f6c59844e841f3526e54633fff963

Observation e78b11c9-eff8-4e9c-a444-ab8939b2b668 · outbound

This paper cites science , volume=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments science , volume=

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.231837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.689919Z digest=sha256:fc630528898f0e9efdbd3ea79b7a5c56ac92de0b37eb75a58904232f5c104ff5

Observation 70a90aee-3bb6-4f08-9859-8440827bb1d2 · outbound

This paper cites 1980 , publisher=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments 1980 , publisher=

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.219880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.693614Z digest=sha256:301bdd8e8e7e82712d4cb6ac2a56580124e571a5e1be925b445c1cce1ac24711

Observation 9923f44d-4655-4f10-92c0-bff98c3bd522 · outbound

This paper cites 1990 , publisher=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments 1990 , publisher=

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.697208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.697208Z digest=sha256:9f863c532d5f6a38c67dd68185eccaf9ff233aaa3e30175cc4eb6acfee93d473

Observation 72eece5e-e793-4ef5-ba53-1c448cd309d5 · outbound

This paper cites American Economic Review , volume=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments American Economic Review , volume=

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.199317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.700971Z digest=sha256:347c557ee871eb971f1565e7a9bfcf7c4f66b255da8481cf3376083696c85d7e

Observation 5e426b3f-8ced-4adf-9d53-be49d5a8976c · outbound

This paper cites Journal of Economic theory , volume=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Journal of Economic theory , volume=

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.186977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.704798Z digest=sha256:63d9be2e44bad85784cef7bf1407b1b2a48a464edab18a906d5a1bd41e786ed6

Observation d61c3931-0e43-4924-b939-4beba99641db · outbound

This paper cites Econometrica: Journal of the Econometric Society , pages=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Econometrica: Journal of the Econometric Society , pages=

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.174422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.708341Z digest=sha256:51982f6bce980d52f3ce8b65c0d54f8fe43e3574af0a7fd4b30de4dadab347c5

Observation 3b7977d3-78c1-4106-970a-0b4540b5b125 · outbound

This paper cites , author=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments , author=

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.160685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.711994Z digest=sha256:6429e0050cce6df3b46127d6af494df840e166e6ffd1b05e5c666cddc284ed25

Observation 8fe176d4-dc5e-4082-8cc1-b7d4313b7eb5 · outbound

This paper cites John thinks that Mary thinks that….

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments John thinks that Mary thinks that…

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.715483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.715483Z digest=sha256:fc315a4efe9ef7d6a4f4ea8bd09248444ed76dc30d08a4e2ba35946d707cb135

Observation 156be659-4703-4394-bce9-251668320b51 · outbound

This paper cites Evolutionary Anthropology: Issues, News, and Reviews , volume=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Evolutionary Anthropology: Issues, News, and Reviews , volume=

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.147997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.721296Z digest=sha256:d2c0601370b5d748199348febadd2121cef9fdd9c791707e0917d7f61e26b094

Observation 191024f9-c656-491a-a870-4531d69f5959 · outbound

This paper cites 1984 , publisher=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments 1984 , publisher=

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.136750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.725068Z digest=sha256:22db6f6109c18bc57403ed7310dd51e40419b56c37cdac1d564178443d360fa3

Observation c96f0ba7-6dc2-493b-beb6-1dc1959b545d · outbound

This paper cites Cognition , volume=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Cognition , volume=

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.125061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.728658Z digest=sha256:8f3fcb6c5d9beba60dbab421eeef6868468bb78c27697a23bf5c62db567683ff

Observation dbe2ea54-809a-4c1a-919d-03136eda64ba · outbound

This paper cites Cognition , volume=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Cognition , volume=

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.111809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.731838Z digest=sha256:6c14a0e15d55d199febf0fe8e027fb7c24b182909eef560cc19eb8191606b8ad

Observation 519e20e8-bb82-4bca-b07b-8c09befafff7 · outbound

This paper cites 1988 , publisher=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments 1988 , publisher=

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.099667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.735336Z digest=sha256:f39d46a8c48520653df28f875bfd35f82b1b73ea5a488eec5b2b6143354657b2

Observation 2a5a11b2-97c1-415a-9ec8-739f58cfd8bf · outbound

This paper cites an unresolved cited work.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Unresolved cited work

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.738866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.738866Z digest=sha256:3d78108073eee41f826c885394ee7c19a86d410c1d540c0323682c70a44fc8ea

Observation 8efb1f2c-664e-4b59-b3d8-904f78ce0a79 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Advances in Neural Information Processing Systems , volume=

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.742161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.742161Z digest=sha256:e0477f790244ebe0d1db1634534a881bfc59a716e3a86e9a6170f194667a7919

Observation 9b046e81-6669-4544-82ea-8f22a48688d7 · outbound

This paper cites Proceedings of the 2022 conference on empirical methods in natural language processing , pages=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Proceedings of the 2022 conference on empirical methods in natural language processing , pages=

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.073627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.745589Z digest=sha256:d806decf75a999b5c8d0f3df59b5e811a8a1cb671de7cd6087a711e4d16ceedb

Observation 0bdd7741-8875-4777-b65a-8942b19cddd3 · outbound

This paper cites Evaluating Large Language Models in Theory of Mind Tasks.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Evaluating Large Language Models in Theory of Mind Tasks

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.748777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.748777Z digest=sha256:7c360c10cbf8fd101b5fa960d8af515dcc2fcc61d0cdeaf428e7b5fab259bbbb

Observation bfd5295b-c9a8-40a6-878a-bc1bde0975ae · outbound

This paper cites Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.752241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.752241Z digest=sha256:133767f33828bd8adde3541e446488769f060c192a872d6ae00404d7485e3058

Observation f800c627-c57e-4d71-9de6-51a68e340350 · outbound

This paper cites AvalonBench: Evaluating LLMs Playing the Game of Avalon.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments AvalonBench: Evaluating LLMs Playing the Game of Avalon

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.755550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.755550Z digest=sha256:62fc6fa4e5138346f7c182002d65860a669739f32591fefffca740b65175fdcc

Observation 6f7e96a1-d04a-46b3-8b66-362b221ffcdd · outbound

This paper cites GTBench: Uncovering the Strategic Reasoning Limitations of LLMs via Game-Theoretic Evaluations.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments GTBench: Uncovering the Strategic Reasoning Limitations of LLMs via Game-Theoretic Evaluations

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.759492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.759492Z digest=sha256:37a130f52fc66f6b215472f6f7f9d20227025aac315259599cc7aed93406f612

Observation a9219dcf-52df-49df-b870-ed3972b466cb · outbound

This paper cites International Conference on Learning Representations , volume=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments International Conference on Learning Representations , volume=

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.061007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.764350Z digest=sha256:84e44805c9131829b44759034fb96f1a625f3c1959edb75f51a01779fc841a8f

Observation ffbaccbe-f001-4bae-aca2-42e1d6af22eb · outbound

This paper cites Voyager: An Open-Ended Embodied Agent with Large Language Models.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Voyager: An Open-Ended Embodied Agent with Large Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.767877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.767877Z digest=sha256:179abe42e7f5f719155a3265c91f6e0c02ae870a41734aee52b991c8bfad3192

Observation 0d998186-d093-463f-b125-9c66efc035c4 · outbound

This paper cites Proceedings of the AAAI Conference on Artificial Intelligence , volume=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Proceedings of the AAAI Conference on Artificial Intelligence , volume=

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.771496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.771496Z digest=sha256:134391567cf439ddddf55f341f8700bd9d7dbcd2433adc296afe0c959254a752

Observation 3fda412d-e11f-4999-9d0e-50d33c78b3ed · outbound

This paper cites Handbook of Intelligence , editor=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Handbook of Intelligence , editor=

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.039264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.774899Z digest=sha256:a995d7bc8d171dba9263354f1d6356fa0d22be156689e48d86417a597e4be15f

Observation 81347811-66ea-4fea-af28-c0484b9fe772 · outbound

This paper cites 2025 , eprint=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments 2025 , eprint=

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.778497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.778497Z digest=sha256:da6cf4bb0f86a4d675ffe8fb50e98c0d926dc65a95ae98b7af06b518f29e7e68

Observation e596d2f7-e9ad-4f6c-ac81-4781c8d7da5d · outbound

This paper cites 2023 , eprint=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments 2023 , eprint=

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.781939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.781939Z digest=sha256:a66f11a26643710c33c79cabfc6c093f8592f882abd50764066e319d8cb9bb92

Observation 8898c9df-a433-4473-89cc-42766f86ecae · outbound

This paper cites Proceedings of the 41st International Conference on Machine Learning , pages=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Proceedings of the 41st International Conference on Machine Learning , pages=

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.011999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.785370Z digest=sha256:4fc829bba865809e9c0041b3f1c99fc22c0e0fd4a51333ad9be201c69156c664

Observation 5c155570-937a-4238-9a6d-c70946b3ad1f · outbound

This paper cites Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages=

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.788809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.788809Z digest=sha256:daf9f5ca6dec4f23e048a33db446c85b6f5dbc9c1237c15a4afd1e939d24563a

Observation 2cbbefed-68ef-4c45-9eea-87278b528cb0 · outbound

This paper cites an unresolved cited work.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Unresolved cited work

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.792530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.792530Z digest=sha256:1379498e855aa981a8e453d16596281eab37fe2ed49f6ef07ac80bb3c95736b0

Observation 77bc36d6-7d6d-482c-a4b2-f9903d3755ec · outbound

This paper cites 2024 , publisher=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments 2024 , publisher=

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:34.991239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.796172Z digest=sha256:b78b1f240449e7510f1dff69fcaebce26d367840c264a4bbb142356734d761a3

Observation 89df69cd-4405-40d8-a20b-7102f040f3d0 · outbound

This paper cites Journal of Consumer Research , volume=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Journal of Consumer Research , volume=

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:34.979116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.799690Z digest=sha256:9dc87fca1ecccda8c31023549526b267b03f52db59ef0c37cc55ddbacdf7ab92

Pith citing papers

No inbound Pith citation observations are available.