Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T15:04:39.270261Z
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 1 inbound Pith citation observation for arXiv:2412.11373.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T15:04:39.270261Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-02T13:11:53.361408Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-02T13:16:58.186477Z
47 of 47 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation e1053adb-1045-4dbc-a31d-8c41abcc6de6 · outbound
Codenames as a Benchmark for Large Language Models A Survey of Large Language Models
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fddc3e02-0dd5-4f13-8059-5433409a8b32 · outbound
Codenames as a Benchmark for Large Language Models Gpt for games: A scoping review (2020-2023),
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af584fee-f280-479d-80a1-6178b5e0780f · outbound
Codenames as a Benchmark for Large Language Models Level generation through large language models,
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 3c72bc04-38ad-4691-a0f7-4c4a35c596c4 · outbound
Codenames as a Benchmark for Large Language Models Langbirds: An agent for angry birds using a large language model,
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 522511b7-a741-47d7-a14b-91d7241ce5bd · outbound
Codenames as a Benchmark for Large Language Models Playing NetHack with LLMs: Potential & Limitations as Zero-Shot Agents
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f14aad39-85a5-426c-b5eb-cf77cd390c4c · outbound
Codenames as a Benchmark for Large Language Models The Go Transformer: Natural Language Modeling for Game Play
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 512caf11-b5d6-49b4-931b-f5aecb328430 · outbound
Codenames as a Benchmark for Large Language Models Generative ai in mafia-like game simulation,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 106e53a6-eeec-4e37-a680-d4e50b2c2278 · outbound
Codenames as a Benchmark for Large Language Models Language-driven play: Large language models as game-playing agents in slay the spire,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation fe3d006f-bf5b-41d4-a4ea-eb4c92a19531 · outbound
Codenames as a Benchmark for Large Language Models GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5780f2a0-c1ad-4f0e-aeb1-95b8f1a95ceb · outbound
Codenames as a Benchmark for Large Language Models A general reinforcement learning algorithm that masters chess, shogi, and Go through self-play,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 958f0864-b30b-46df-81a4-5488638df476 · outbound
Codenames as a Benchmark for Large Language Models Mastering the game of Go without human knowledge,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 46e42308-dc23-4486-be9d-abf06c1d10d7 · outbound
Codenames as a Benchmark for Large Language Models TAG: Pandemic Competition,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation a02c347d-d6ce-4ffe-a0dd-e43c48bf2977 · outbound
Codenames as a Benchmark for Large Language Models Chv ´atil, Codenames
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation a0e76b01-32d9-4fbc-a544-abf71eec9020 · outbound
Codenames as a Benchmark for Large Language Models The codenames ai competition,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation e2c4dabd-82d5-4c9d-be51-7627e5fdcf9a · outbound
Codenames as a Benchmark for Large Language Models Cooperation and Codenames: Understanding Natural Language Processing via Code- names,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 3cd71485-9486-4e8d-8da2-36a8843fef01 · outbound
Codenames as a Benchmark for Large Language Models Evaluating Large Language Models in Theory of Mind Tasks
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29feeb47-a770-45e3-9b1a-71fe8b55b244 · outbound
Codenames as a Benchmark for Large Language Models Measuring Massive Multitask Language Understanding
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a898356-201a-4632-8634-32b466472bd5 · outbound
Codenames as a Benchmark for Large Language Models HellaSwag: Can a Machine Really Finish Your Sentence?
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b93d489-5b7d-40f2-9c71-449c20c378d5 · outbound
Codenames as a Benchmark for Large Language Models Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f077eae6-4009-4986-b626-899bc7d7a8e9 · outbound
Codenames as a Benchmark for Large Language Models Towards Reasoning in Large Language Models: A Survey
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa84398f-fb26-435b-a8aa-a1057377358d · outbound
Codenames as a Benchmark for Large Language Models LLM as a Mastermind: A Survey of Strategic Reasoning with Large Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6e0c79f-dd1f-409c-8eba-9e2633e7ea59 · outbound
Codenames as a Benchmark for Large Language Models Tydi qa: A benchmark for information-seeking question answering in typologically diverse languages,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 84a0128e-a2e0-4d33-b1e1-a7365a0a209a · outbound
Codenames as a Benchmark for Large Language Models Measuring Mathematical Problem Solving With the MATH Dataset
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42ed6593-d7f7-46eb-a0bb-90e0f5832853 · outbound
Codenames as a Benchmark for Large Language Models Neural Theory-of-Mind? On the Limits of Social Intelligence in Large LMs
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 561e2d28-16b3-4733-b2b7-ae5cf74167e5 · outbound
Codenames as a Benchmark for Large Language Models Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 5076321f-b35e-4b8f-9ecb-1623980f845f · outbound
Codenames as a Benchmark for Large Language Models Prompt Engineering ChatGPT for Co- denames,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 1445d912-3197-467c-8697-a92b6264b4b0 · outbound
Codenames as a Benchmark for Large Language Models Strategic Reasoning with Language Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5414ba8-bc6b-4770-bf81-bef79fc690ab · outbound
Codenames as a Benchmark for Large Language Models Human-AI Collaboration in Cooperative Games: A Study of Playing Codenames with an LLM Assistant,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 818a5283-b747-4866-b203-b8079a4c805b · outbound
Codenames as a Benchmark for Large Language Models LLMs achieve adult human performance on higher-order theory of mind tasks,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 804c3586-7bb7-4d04-88ea-cf0bd26d0cf4 · outbound
Codenames as a Benchmark for Large Language Models Word autobots: Using transformers for word association in the game codenames,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 6cc3fe07-0614-46ac-afcb-4c5d7813cc24 · outbound
Codenames as a Benchmark for Large Language Models Playing Codenames with Language Graphs and Word Embeddings,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 3e9de337-dff8-42b5-b725-c7918b2aafd6 · outbound
Codenames as a Benchmark for Large Language Models Adapting to teammates in a cooperative language game,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 279aa059-2e3c-4378-9d41-7da5308293e3 · outbound
Codenames as a Benchmark for Large Language Models Noisy communication modeling for improved cooperation in codenames,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 09c2e3b7-51e6-4632-bf22-2ab2f559585b · outbound
Codenames as a Benchmark for Large Language Models ThinkSum: Probabilis- tic reasoning over sets using large language models,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 7260ea2c-a4c9-452e-8af2-d17cc01ccf4a · outbound
Codenames as a Benchmark for Large Language Models Large Language Models are Zero-Shot Reasoners,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 1930016f-fccd-4c5a-b3ff-8443e1df8264 · outbound
Codenames as a Benchmark for Large Language Models Self-Refine: Iterative Refinement with Self-Feedback,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 7ffcb101-ed85-44d5-83b5-dc3d4b45f3c0 · outbound
Codenames as a Benchmark for Large Language Models Unleashing the Emergent Cognitive Synergy in Large Language Models: A Task- Solving Agent through Multi-Persona Self-Collaboration,
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation d1edced3-fcdd-4898-b7c7-b6051bd8cee2 · outbound
Codenames as a Benchmark for Large Language Models Efficient Estimation of Word Representations in Vector Space
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a47810d-927f-4ff0-bb75-89fb5f625ade · outbound
Codenames as a Benchmark for Large Language Models GloVe: Global vectors for word representation,
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd9f6a45-853d-4e21-a771-bdd9cd0eb6fd · outbound
Codenames as a Benchmark for Large Language Models Concatenated power mean word embeddings as universal cross-lingual sentence representations,
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 6a2722be-1261-4bf6-87c7-ae39d0a705b8 · outbound
Codenames as a Benchmark for Large Language Models Human-ai collaboration in cooperative games: A study of playing codenames with an llm assistant,
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9f0da6b-4d2d-483b-9394-80690c54fd50 · outbound
Codenames as a Benchmark for Large Language Models Semantic Priming Effects In Visual Word Recognition: A Selective Review Of Current Findings And Theories,
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 9f5d4445-7aba-46af-bbd3-3f0c3c8fee82 · outbound
Codenames as a Benchmark for Large Language Models Prototypes Revisited,
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation f6cdfe12-e54f-4743-88a0-24a5f8cae024 · outbound
Codenames as a Benchmark for Large Language Models Context-independent and context-dependent informa- tion in concepts,
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c46ee15c-2bcf-4864-ac5f-16a75f22cd1f · outbound
Codenames as a Benchmark for Large Language Models Human Learning from Artificial Intel- ligence: Evidence from Human Go Players’ Decisions after AlphaGo,
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 341c479d-6f9a-4f4a-aa8e-3fe80a6ca425 · outbound
Codenames as a Benchmark for Large Language Models Generative AI in Mafia-like Game Simulation
Reference 2023
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 6de52024-8aaf-4918-b232-a21e783e1faa · outbound
Codenames as a Benchmark for Large Language Models Available: https://doi.org/10.1145/3649921.3650013
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73901c8d-411a-4fc7-be78-0ef9fa024dac · inbound
"Don't Say It!": Constraints, Compliance, and Communication when Language Models Play Taboo Codenames as a Benchmark for Large Language Models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.