Pith. sign in

Paper Citation Record · LEDGER

Evaluating Large Language Models with Grid-Based Game Competitions: An Extensible LLM Benchmark and Leaderboard

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2407.07796.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.07796 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T22:50:32.995433Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T13:41:36.684597Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation dcfa133e-3d45-45ae-a553-635e64b4425c · inbound

Game Theory Meets Large Language Models: A Systematic Survey with Taxonomy and New Frontiers cites this paper.

Game Theory Meets Large Language Models: A Systematic Survey with Taxonomy and New Frontiers Evaluating Large Language Models with Grid-Based Game Competitions: An Extensible LLM Benchmark and Leaderboard

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T22:50:32.995433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T22:50:32.995433Z digest=sha256:9aaceb0b5d6c1091c19e7f872a60da315acec842dc2cfb7675730fd382289e0e

Observation bd94c50a-1a6f-407c-878b-f44e1c7a8789 · inbound

MTR-Bench: A Comprehensive Benchmark for Multi-Turn Reasoning Evaluation cites this paper.

MTR-Bench: A Comprehensive Benchmark for Multi-Turn Reasoning Evaluation Evaluating Large Language Models with Grid-Based Game Competitions: An Extensible LLM Benchmark and Leaderboard

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-22T13:41:36.687125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T13:37:49.475416Z digest=sha256:1d5b1b47e65a2fc1b0d1fcad666fa2140ab66081d19ceb5610c5ff93724781d4

Observation e773708e-4e7f-4d08-95e7-99412259d6b0 · inbound

MARBLE: A Hard Benchmark for Multimodal Spatial Reasoning and Planning cites this paper.

MARBLE: A Hard Benchmark for Multimodal Spatial Reasoning and Planning Evaluating Large Language Models with Grid-Based Game Competitions: An Extensible LLM Benchmark and Leaderboard

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T21:58:36.058002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:58:36.058002Z digest=sha256:24762982a92c1f655c5e94ce0727644074c2e1b9a809cdaa1201c90c63a0e30e

Observation 526200cd-e14d-497e-bfc1-7226018d7aa2 · inbound

Game Reasoning Arena: A Framework and Benchmark for Assessing Reasoning Capabilities of Large Language Models via Game Play cites this paper.

Game Reasoning Arena: A Framework and Benchmark for Assessing Reasoning Capabilities of Large Language Models via Game Play Evaluating Large Language Models with Grid-Based Game Competitions: An Extensible LLM Benchmark and Leaderboard

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T04:33:42.811925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T04:33:42.811925Z digest=sha256:757645bf170cdc727a1e7d83889258cd760836ef13b9c8676ff41ef05b55003e

Observation 347ec6d5-84a2-47d0-b02c-4a838aa54c0a · inbound

When Reasoning Narrows the Move: Diversity Collapse in LLM Game Play cites this paper.

When Reasoning Narrows the Move: Diversity Collapse in LLM Game Play Evaluating Large Language Models with Grid-Based Game Competitions: An Extensible LLM Benchmark and Leaderboard

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T12:36:12.257522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:36:12.257522Z digest=sha256:d66e247271f690e7fae36de1bc50a500de995c05ded2fb355c38fefa27679776