Pith. sign in

Paper Citation Record · LEDGER

Codenames as a Benchmark for Large Language Models

As of 19 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 1 inbound Pith citation observation for arXiv:2412.11373.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.11373 v2

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T15:04:39.270261Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-02T13:11:53.361408Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T13:16:58.186477Z

Reference resolution

47 of 47 outbound references displayed

  • verified exact1
  • verified fuzzy28
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e1053adb-1045-4dbc-a31d-8c41abcc6de6 · outbound

This paper cites A Survey of Large Language Models.

Codenames as a Benchmark for Large Language Models A Survey of Large Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T15:04:39.051311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:04:39.051311Z digest=sha256:765ee28fced3dcb417980152297aa1079a6fcf880d7a4d6138f957ff9f8b90c5

Observation fddc3e02-0dd5-4f13-8059-5433409a8b32 · outbound

This paper cites Gpt for games: A scoping review (2020-2023),.

Codenames as a Benchmark for Large Language Models Gpt for games: A scoping review (2020-2023),

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T15:04:39.056864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:04:39.056864Z digest=sha256:aa4258f9c8c18a5a64bd6ebac13d6f74377261ed8874c05a7fa7ba6129aa9d47

Observation af584fee-f280-479d-80a1-6178b5e0780f · outbound

This paper cites Level generation through large language models,.

Codenames as a Benchmark for Large Language Models Level generation through large language models,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:04:40.167761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T15:04:39.062190Z digest=sha256:49d2a24baecdbf2bc67e8c58d6f7808c96a7eba63baafcecd473daf5a29c43bd

Observation 3c72bc04-38ad-4691-a0f7-4c4a35c596c4 · outbound

This paper cites Langbirds: An agent for angry birds using a large language model,.

Codenames as a Benchmark for Large Language Models Langbirds: An agent for angry birds using a large language model,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:04:40.152805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T15:04:39.067697Z digest=sha256:71de02e9e10bba9d9dec0373ff713d281df1a2af6dc60c877dd011f0aee7f2fd

Observation 522511b7-a741-47d7-a14b-91d7241ce5bd · outbound

This paper cites Playing NetHack with LLMs: Potential & Limitations as Zero-Shot Agents.

Codenames as a Benchmark for Large Language Models Playing NetHack with LLMs: Potential & Limitations as Zero-Shot Agents

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T15:04:39.072749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:04:39.072749Z digest=sha256:d48862e9f677e109872f69d20be74abeba149a1250b1354332142807a3a3d3c1

Observation f14aad39-85a5-426c-b5eb-cf77cd390c4c · outbound

This paper cites The Go Transformer: Natural Language Modeling for Game Play.

Codenames as a Benchmark for Large Language Models The Go Transformer: Natural Language Modeling for Game Play

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-08-11T15:04:39.631742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T15:04:39.078134Z digest=sha256:2de5c7e9e4c09636b019a0889ee02f80bc8e804edc7823ed23acd659788a4884

Observation 512caf11-b5d6-49b4-931b-f5aecb328430 · outbound

This paper cites Generative ai in mafia-like game simulation,.

Codenames as a Benchmark for Large Language Models Generative ai in mafia-like game simulation,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:04:40.138208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T15:04:39.084726Z digest=sha256:80c1f5902688ac381c96a01950e118bfef30edc73e8f5f60d1568d3141ece1ad

Observation 106e53a6-eeec-4e37-a680-d4e50b2c2278 · outbound

This paper cites Language-driven play: Large language models as game-playing agents in slay the spire,.

Codenames as a Benchmark for Large Language Models Language-driven play: Large language models as game-playing agents in slay the spire,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:04:40.123563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T15:04:39.094186Z digest=sha256:b9923b122862056ca31f003723fb808977c08ba8c0133b5f86f565a090265edc

Observation fe3d006f-bf5b-41d4-a4ea-eb4c92a19531 · outbound

This paper cites GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents.

Codenames as a Benchmark for Large Language Models GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T15:04:39.102857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:04:39.102857Z digest=sha256:797bacc51d7915d9fdf1b5583747bffad9124980f086ec8c8da9f2a1e0576949

Observation 5780f2a0-c1ad-4f0e-aeb1-95b8f1a95ceb · outbound

This paper cites A general reinforcement learning algorithm that masters chess, shogi, and Go through self-play,.

Codenames as a Benchmark for Large Language Models A general reinforcement learning algorithm that masters chess, shogi, and Go through self-play,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:04:40.108499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T15:04:39.107738Z digest=sha256:4ab398c6426ffa37116844ca469d5ce7a87be5fd85b338d55cfdfae68973b767

Observation 958f0864-b30b-46df-81a4-5488638df476 · outbound

This paper cites Mastering the game of Go without human knowledge,.

Codenames as a Benchmark for Large Language Models Mastering the game of Go without human knowledge,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:04:40.093948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T15:04:39.112392Z digest=sha256:4da4ffed53f085e64aaac5361bd7f5f620906367cb67d09fd6f6ed6c7e0b2de0

Observation 46e42308-dc23-4486-be9d-abf06c1d10d7 · outbound

This paper cites TAG: Pandemic Competition,.

Codenames as a Benchmark for Large Language Models TAG: Pandemic Competition,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:04:40.080150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T15:04:39.117077Z digest=sha256:03204a01c462b83a8413835be605d9c91663c09135b5f08bd2587d9d5583a4b0

Observation a02c347d-d6ce-4ffe-a0dd-e43c48bf2977 · outbound

This paper cites Chv ´atil, Codenames.

Codenames as a Benchmark for Large Language Models Chv ´atil, Codenames

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:04:40.066129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T15:04:39.121622Z digest=sha256:9f3f110ec9e953b142e51359ba6ba2ecdbe6ae940a00f9c7529455159e8c5b9c

Observation a0e76b01-32d9-4fbc-a544-abf71eec9020 · outbound

This paper cites The codenames ai competition,.

Codenames as a Benchmark for Large Language Models The codenames ai competition,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:04:40.051931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T15:04:39.126199Z digest=sha256:a16df874fb601fdbb436dda87d2f500ac1f63fd0751d59b41e6137a9b55bcbd6

Observation e2c4dabd-82d5-4c9d-be51-7627e5fdcf9a · outbound

This paper cites Cooperation and Codenames: Understanding Natural Language Processing via Code- names,.

Codenames as a Benchmark for Large Language Models Cooperation and Codenames: Understanding Natural Language Processing via Code- names,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:04:40.036135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T15:04:39.130669Z digest=sha256:f8236b8c2fe11ebe82ff5098e168d0d8674a3a951eb56749ff245f7cfce74cf7

Observation 3cd71485-9486-4e8d-8da2-36a8843fef01 · outbound

This paper cites Evaluating Large Language Models in Theory of Mind Tasks.

Codenames as a Benchmark for Large Language Models Evaluating Large Language Models in Theory of Mind Tasks

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T15:04:39.135574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:04:39.135574Z digest=sha256:5481bd2b2e1caa7d2a208a566fcbfed53da92569d474af6384161fb72fc6ad26

Observation 29feeb47-a770-45e3-9b1a-71fe8b55b244 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Codenames as a Benchmark for Large Language Models Measuring Massive Multitask Language Understanding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T15:04:39.140371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:04:39.140371Z digest=sha256:ad54b2657aa4c9625d862a2ad76697d6fd568ccf90f8ead6aafc572021fbecde

Observation 6a898356-201a-4632-8634-32b466472bd5 · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

Codenames as a Benchmark for Large Language Models HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T15:04:39.145384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:04:39.145384Z digest=sha256:f047f5cf2d98c56afaf1b8acc9f2b73fe369775cd741a26ab7a793ce5eef37d0

Observation 5b93d489-5b7d-40f2-9c71-449c20c378d5 · outbound

This paper cites Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them.

Codenames as a Benchmark for Large Language Models Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T15:04:39.150171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:04:39.150171Z digest=sha256:2bf710e69b0402c970f07eb57bb4c95da4e2b14e8e1539f2b7d5fd1dab71ca5c

Observation f077eae6-4009-4986-b626-899bc7d7a8e9 · outbound

This paper cites Towards Reasoning in Large Language Models: A Survey.

Codenames as a Benchmark for Large Language Models Towards Reasoning in Large Language Models: A Survey

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T15:04:39.155107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:04:39.155107Z digest=sha256:435c89139336bd57373ea84e3bd445532171b87c317ebc41fb6699dee636284c

Observation fa84398f-fb26-435b-a8aa-a1057377358d · outbound

This paper cites LLM as a Mastermind: A Survey of Strategic Reasoning with Large Language Models.

Codenames as a Benchmark for Large Language Models LLM as a Mastermind: A Survey of Strategic Reasoning with Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T15:04:39.159692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:04:39.159692Z digest=sha256:7e54f7dac72f0bd9447f7dc51fcc4f411028c0b669542176a3fa5c37ab729d7a

Observation f6e0c79f-dd1f-409c-8eba-9e2633e7ea59 · outbound

This paper cites Tydi qa: A benchmark for information-seeking question answering in typologically diverse languages,.

Codenames as a Benchmark for Large Language Models Tydi qa: A benchmark for information-seeking question answering in typologically diverse languages,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:04:40.021270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T15:04:39.163755Z digest=sha256:c43eceafa3a75bd15d87a1693e20625ba2f8949b2e40cee37a9e129e2f51cae3

Observation 84a0128e-a2e0-4d33-b1e1-a7365a0a209a · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Codenames as a Benchmark for Large Language Models Measuring Mathematical Problem Solving With the MATH Dataset

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T15:04:39.167804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:04:39.167804Z digest=sha256:f4406c531fddb03025c970dd18d4914c1accd9e82479133fddef08f68bf88e34

Observation 42ed6593-d7f7-46eb-a0bb-90e0f5832853 · outbound

This paper cites Neural Theory-of-Mind? On the Limits of Social Intelligence in Large LMs.

Codenames as a Benchmark for Large Language Models Neural Theory-of-Mind? On the Limits of Social Intelligence in Large LMs

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T15:04:39.172302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:04:39.172302Z digest=sha256:149651777e6bf60049aaf679e942d213ba25bbcd036d4e7a602765c40368d388

Observation 561e2d28-16b3-4733-b2b7-ae5cf74167e5 · outbound

This paper cites Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks,.

Codenames as a Benchmark for Large Language Models Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:04:40.006222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T15:04:39.177381Z digest=sha256:a1f3563ed0c4afcd3cdb8566c88a7cf5c93047f001b5418ddc92082c69726752

Observation 5076321f-b35e-4b8f-9ecb-1623980f845f · outbound

This paper cites Prompt Engineering ChatGPT for Co- denames,.

Codenames as a Benchmark for Large Language Models Prompt Engineering ChatGPT for Co- denames,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:04:39.990849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T15:04:39.182207Z digest=sha256:502bcb2e2e4e45e0fbd096596100144ea5b41ad596b6fb65c442ccaf4bd876e5

Observation 1445d912-3197-467c-8697-a92b6264b4b0 · outbound

This paper cites Strategic Reasoning with Language Models.

Codenames as a Benchmark for Large Language Models Strategic Reasoning with Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T15:04:39.186524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:04:39.186524Z digest=sha256:7f86c13cdde397c4fa4d5024231d1df4de1b80fe7cf928be49854d93081cacf0

Observation e5414ba8-bc6b-4770-bf81-bef79fc690ab · outbound

This paper cites Human-AI Collaboration in Cooperative Games: A Study of Playing Codenames with an LLM Assistant,.

Codenames as a Benchmark for Large Language Models Human-AI Collaboration in Cooperative Games: A Study of Playing Codenames with an LLM Assistant,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:04:39.976103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T15:04:39.191199Z digest=sha256:066d5431f66bcd4f0c8d86b275fa13a89428ade07cd5b5e7db6cdc0b169cc1c9

Observation 818a5283-b747-4866-b203-b8079a4c805b · outbound

This paper cites LLMs achieve adult human performance on higher-order theory of mind tasks,.

Codenames as a Benchmark for Large Language Models LLMs achieve adult human performance on higher-order theory of mind tasks,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:04:39.961158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T15:04:39.195587Z digest=sha256:9d8819fac6ecac2c3e75e9e047523f4751d17d569e1ec3b2d23170e9a5d8239e

Observation 804c3586-7bb7-4d04-88ea-cf0bd26d0cf4 · outbound

This paper cites Word autobots: Using transformers for word association in the game codenames,.

Codenames as a Benchmark for Large Language Models Word autobots: Using transformers for word association in the game codenames,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:04:39.946218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T15:04:39.200090Z digest=sha256:421aad8b77b84632485004658945f5595ce8db875833f779238bd862b8dba4fa

Observation 6cc3fe07-0614-46ac-afcb-4c5d7813cc24 · outbound

This paper cites Playing Codenames with Language Graphs and Word Embeddings,.

Codenames as a Benchmark for Large Language Models Playing Codenames with Language Graphs and Word Embeddings,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:04:39.931713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T15:04:39.204484Z digest=sha256:d638a14f5ed3880acc5e2dc88ac03861bd400c78f681cc336bea51ee7d8970bf

Observation 3e9de337-dff8-42b5-b725-c7918b2aafd6 · outbound

This paper cites Adapting to teammates in a cooperative language game,.

Codenames as a Benchmark for Large Language Models Adapting to teammates in a cooperative language game,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:04:39.916655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T15:04:39.208956Z digest=sha256:970b0f31c7ff4318a7d0b9108e48cf5580769b0eef1ad9e8ff5d5f54f7656c04

Observation 279aa059-2e3c-4378-9d41-7da5308293e3 · outbound

This paper cites Noisy communication modeling for improved cooperation in codenames,.

Codenames as a Benchmark for Large Language Models Noisy communication modeling for improved cooperation in codenames,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:04:39.899931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T15:04:39.213267Z digest=sha256:7cf056666ba79fe5ef08cfa2d723ea2d987683ed0d93b7269c14a14496de31f3

Observation 09c2e3b7-51e6-4632-bf22-2ab2f559585b · outbound

This paper cites ThinkSum: Probabilis- tic reasoning over sets using large language models,.

Codenames as a Benchmark for Large Language Models ThinkSum: Probabilis- tic reasoning over sets using large language models,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:04:39.884924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T15:04:39.217646Z digest=sha256:437849ce49e33a95f1f7d44369cc6feb4f175601d0b148d3f57e8bbedd3eaac6

Observation 7260ea2c-a4c9-452e-8af2-d17cc01ccf4a · outbound

This paper cites Large Language Models are Zero-Shot Reasoners,.

Codenames as a Benchmark for Large Language Models Large Language Models are Zero-Shot Reasoners,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:04:39.870893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T15:04:39.222073Z digest=sha256:38eff17f723a596befde499963b147a6c79a10f900d9228c2d2dce460926685f

Observation 1930016f-fccd-4c5a-b3ff-8443e1df8264 · outbound

This paper cites Self-Refine: Iterative Refinement with Self-Feedback,.

Codenames as a Benchmark for Large Language Models Self-Refine: Iterative Refinement with Self-Feedback,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:04:39.857002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T15:04:39.226672Z digest=sha256:7ecb3d3f3d3d79d9253599802c3e458eb99b114e0a07546a1116107295628945

Observation 7ffcb101-ed85-44d5-83b5-dc3d4b45f3c0 · outbound

This paper cites Unleashing the Emergent Cognitive Synergy in Large Language Models: A Task- Solving Agent through Multi-Persona Self-Collaboration,.

Codenames as a Benchmark for Large Language Models Unleashing the Emergent Cognitive Synergy in Large Language Models: A Task- Solving Agent through Multi-Persona Self-Collaboration,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:04:39.842119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T15:04:39.231054Z digest=sha256:8ef1d71aea2e08350b2e7bc23158a3e673a2ad9870339840027b796faf9af215

Observation d1edced3-fcdd-4898-b7c7-b6051bd8cee2 · outbound

This paper cites Efficient Estimation of Word Representations in Vector Space.

Codenames as a Benchmark for Large Language Models Efficient Estimation of Word Representations in Vector Space

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T15:04:39.236693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:04:39.236693Z digest=sha256:fe9294f5f8eb40819e41f44313dad95fa9679665584acefee00a37e630b3d73b

Observation 7a47810d-927f-4ff0-bb75-89fb5f625ade · outbound

This paper cites GloVe: Global vectors for word representation,.

Codenames as a Benchmark for Large Language Models GloVe: Global vectors for word representation,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T15:04:39.240994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:04:39.240994Z digest=sha256:9a8c579ad93776449d2588a6a72bc55833bac8576e844e606d6851b76450e3ee

Observation cd9f6a45-853d-4e21-a771-bdd9cd0eb6fd · outbound

This paper cites Concatenated power mean word embeddings as universal cross-lingual sentence representations,.

Codenames as a Benchmark for Large Language Models Concatenated power mean word embeddings as universal cross-lingual sentence representations,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:04:39.814965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T15:04:39.246450Z digest=sha256:b7b89c6cc133929118684cda1656f49f79c86a9969d875e0c818d523185d0c4e

Observation 6a2722be-1261-4bf6-87c7-ae39d0a705b8 · outbound

This paper cites Human-ai collaboration in cooperative games: A study of playing codenames with an llm assistant,.

Codenames as a Benchmark for Large Language Models Human-ai collaboration in cooperative games: A study of playing codenames with an llm assistant,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T15:04:39.250962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:04:39.250962Z digest=sha256:5cae0e2e6483e2988202e5bdf6852068d21f5703d5917ea71c49bb8508306508

Observation a9f0da6b-4d2d-483b-9394-80690c54fd50 · outbound

This paper cites Semantic Priming Effects In Visual Word Recognition: A Selective Review Of Current Findings And Theories,.

Codenames as a Benchmark for Large Language Models Semantic Priming Effects In Visual Word Recognition: A Selective Review Of Current Findings And Theories,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:04:39.799558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T15:04:39.256460Z digest=sha256:d228518de06798f21aa837b651ae24fb623570e1c0ec77f706780d56399a32eb

Observation 9f5d4445-7aba-46af-bbd3-3f0c3c8fee82 · outbound

This paper cites Prototypes Revisited,.

Codenames as a Benchmark for Large Language Models Prototypes Revisited,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:04:39.784433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T15:04:39.261427Z digest=sha256:1a9b3f43689e401ae5315c6bf73574ebea1cd843f18f1c2a123eafa875875f5f

Observation f6cdfe12-e54f-4743-88a0-24a5f8cae024 · outbound

This paper cites Context-independent and context-dependent informa- tion in concepts,.

Codenames as a Benchmark for Large Language Models Context-independent and context-dependent informa- tion in concepts,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:04:39.769921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T15:04:39.265579Z digest=sha256:588ded08fe6b83783fde9677ef3c082ccb60ce6c8938d2f38c9ef074c6a37dc8

Observation c46ee15c-2bcf-4864-ac5f-16a75f22cd1f · outbound

This paper cites Human Learning from Artificial Intel- ligence: Evidence from Human Go Players’ Decisions after AlphaGo,.

Codenames as a Benchmark for Large Language Models Human Learning from Artificial Intel- ligence: Evidence from Human Go Players’ Decisions after AlphaGo,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:04:39.755098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T15:04:39.270261Z digest=sha256:c7f9b5150aa72598f42974bb068df6bd373b72991e785c683c9f5da15e3c3e57

Observation 341c479d-6f9a-4f4a-aa8e-3fe80a6ca425 · outbound

This paper cites Generative AI in Mafia-like Game Simulation.

Codenames as a Benchmark for Large Language Models Generative AI in Mafia-like Game Simulation

Reference 2023

Resolution
metadata mismatch
local_arxiv, observed 2026-08-11T15:04:39.610750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T15:04:39.089343Z digest=sha256:f97e52f345438942c9ccdc1478a10f352bca8b217693b15289d4984e4629d45d

Observation 6de52024-8aaf-4918-b232-a21e783e1faa · outbound

This paper cites Available: https://doi.org/10.1145/3649921.3650013.

Codenames as a Benchmark for Large Language Models Available: https://doi.org/10.1145/3649921.3650013

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T15:04:39.098567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:04:39.098567Z digest=sha256:44b25a23f109da33a334422f60868d22b73d2a67f9517dbaa731bed6f996959b

Pith citing papers

Observation 73901c8d-411a-4fc7-be78-0ef9fa024dac · inbound

"Don't Say It!": Constraints, Compliance, and Communication when Language Models Play Taboo cites this paper.

"Don't Say It!": Constraints, Compliance, and Communication when Language Models Play Taboo Codenames as a Benchmark for Large Language Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:16:58.188109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-02T13:11:53.361408Z digest=sha256:da03fb2faf0637208875aba6cd7b7bb6584f9687de240cfa5a593c776e420fbb