Pith. sign in

Paper Citation Record · LEDGER

SmartPlay: A Benchmark for LLMs as Intelligent Agents

As of 24 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 27 inbound Pith citation observations for arXiv:2310.01557.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.01557 v5

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 27 of 27 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:42:30.786968Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T10:58:02.902301Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 90270e29-6aa2-494f-9d05-ce4d4e24c221 · inbound

The Landscape of Emerging AI Agent Architectures for Reasoning, Planning, and Tool Calling: A Survey cites this paper.

The Landscape of Emerging AI Agent Architectures for Reasoning, Planning, and Tool Calling: A Survey SmartPlay: A Benchmark for LLMs as Intelligent Agents

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T23:16:41.811118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T23:16:41.679855Z digest=sha256:742674a297f6e9a6e104f82485c8b18b5ad94785431b166a286ac99ca31973dc

Observation f5de8c36-fab9-4f8c-a80b-90e33d3bf892 · inbound

Learning to Ask: When LLM Agents Meet Unclear Instruction cites this paper.

Learning to Ask: When LLM Agents Meet Unclear Instruction SmartPlay: A Benchmark for LLMs as Intelligent Agents

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-23T21:13:28.105584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-23T21:08:42.276002Z digest=sha256:ebfcf44fd2ba377fa4fc2382dabb6dc04fe78462a6941937dee073a2e75bc9b1

Observation 0f6ca42a-c266-406b-8868-9bf44571fc3b · inbound

VideoCogQA: A Controllable Benchmark for Evaluating Cognitive Abilities in Video-Language Models cites this paper.

VideoCogQA: A Controllable Benchmark for Evaluating Cognitive Abilities in Video-Language Models SmartPlay: A Benchmark for LLMs as Intelligent Agents

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T21:07:09.648291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:07:09.648291Z digest=sha256:e31e94962f1a45b1809a1fb9bb2fd3d38ad5fa653821daed7b6312857778a57d

Observation 9a1212a9-3faa-4bcb-bde0-a3a9a6fc2689 · inbound

Probing the Capacity of Language Model Agents to Operationalize Disparate Experiential Context Despite Distraction cites this paper.

Probing the Capacity of Language Model Agents to Operationalize Disparate Experiential Context Despite Distraction SmartPlay: A Benchmark for LLMs as Intelligent Agents

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T17:12:27.531912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:12:27.531912Z digest=sha256:fa0bd30695bb588829a27ae36c870bbd0242e27097170280266cc5b07b95e7fb

Observation 59022474-3e81-49ff-bd8b-ba06c6e099c1 · inbound

BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games cites this paper.

BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games SmartPlay: A Benchmark for LLMs as Intelligent Agents

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-12T16:21:08.429907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:21:08.429907Z digest=sha256:f1c5c10a690a5787650028fdfe1349698fca355959705d0c9563e0cde4cc5c2c

Observation cd2ce70c-ca21-4281-ade4-9ed29175bc4c · inbound

SPAgent: Adaptive Task Decomposition and Model Selection for General Video Generation and Editing cites this paper.

SPAgent: Adaptive Task Decomposition and Model Selection for General Video Generation and Editing SmartPlay: A Benchmark for LLMs as Intelligent Agents

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T10:44:14.831130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:44:14.831130Z digest=sha256:329e24736a532d41f39e672d02e169d0a5f60bb50c9f06862040edfdaed916df

Observation b6e492cf-cafb-4b0a-a162-0297a0209807 · inbound

Disentangling Exploration of Large Language Models by Optimal Exploitation cites this paper.

Disentangling Exploration of Large Language Models by Optimal Exploitation SmartPlay: A Benchmark for LLMs as Intelligent Agents

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-10T20:18:58.494246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:18:58.494246Z digest=sha256:e6921d0ea2adbf7479fb0f0523bef05a9c8c322d0eb185887c16bc74aa79d609

Observation 8e777ba9-046d-4891-8604-e21799a6a796 · inbound

Large Language Model-Enhanced Multi-Armed Bandits cites this paper.

Large Language Model-Enhanced Multi-Armed Bandits SmartPlay: A Benchmark for LLMs as Intelligent Agents

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-09T16:35:28.435740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:35:28.435740Z digest=sha256:d4a2765ff0c741ff84057a7bd94fc78b4370190e2921eb123f4b517efe4dd50c

Observation 8e5c67e6-ed17-4d1c-be77-e217a9e33843 · inbound

PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities cites this paper.

PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities SmartPlay: A Benchmark for LLMs as Intelligent Agents

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:30.786968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:30.786968Z digest=sha256:e74d0196feb798308f75e0d1364f14ccde572a743e14ac4404278325bd805cda

Observation 540c3282-0c12-4267-ad01-be8ff0d44ae8 · inbound

Generative to Agentic AI: Survey, Conceptualization, and Challenges cites this paper.

Generative to Agentic AI: Survey, Conceptualization, and Challenges SmartPlay: A Benchmark for LLMs as Intelligent Agents

Reference 154

Resolution
unresolved
no resolver link, observed 2026-08-16T10:10:25.290195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:10:25.290195Z digest=sha256:bd6b6f450de9704be6c2d51ac1b510f32787f820516fccdb94c670469a9a8ba4

Observation 9abf58a2-bfbc-4c3f-857b-07344db01db4 · inbound

Humanizing LLMs: A Survey of Psychological Measurements with Tools, Datasets, and Human-Agent Applications cites this paper.

Humanizing LLMs: A Survey of Psychological Measurements with Tools, Datasets, and Human-Agent Applications SmartPlay: A Benchmark for LLMs as Intelligent Agents

Reference 145

Resolution
unresolved
no resolver link, observed 2026-08-16T05:09:43.493662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:09:43.493662Z digest=sha256:e29720c17e9111718ff96495c723749a9480c737e09da921453d9e691d883e51

Observation 564526f6-4b44-4d2e-9eb5-c1230ed93105 · inbound

Reasoning Capabilities of Large Language Models on Dynamic Tasks cites this paper.

Reasoning Capabilities of Large Language Models on Dynamic Tasks SmartPlay: A Benchmark for LLMs as Intelligent Agents

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T21:12:01.502630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:12:01.502630Z digest=sha256:dd39791b8591331d9ae2ab507b1fe394046ca70ecf048e812c9d204303e4d22f

Observation 3c5628f0-b510-41b6-afb7-d5133c637cf5 · inbound

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning cites this paper.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning SmartPlay: A Benchmark for LLMs as Intelligent Agents

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:40.592511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:40.592511Z digest=sha256:4767a5118abada34abc7f99008a0257e4f624a64655f29e3faf61160a08d3486

Observation feb50f57-f0b3-4573-880d-350a08c99f3b · inbound

lmgame-Bench: How Good are LLMs at Playing Games? cites this paper.

lmgame-Bench: How Good are LLMs at Playing Games? SmartPlay: A Benchmark for LLMs as Intelligent Agents

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:57.020933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:57.020933Z digest=sha256:82efd241bdab0fb4d1bdf6ea797af8e50568a2f4d145aa170f5da43996247cda

Observation 30df3919-5f42-4860-af5b-83969385785e · inbound

LLM-Powered AI Agent Systems and Their Applications in Industry cites this paper.

LLM-Powered AI Agent Systems and Their Applications in Industry SmartPlay: A Benchmark for LLMs as Intelligent Agents

Reference 106

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T14:06:38.128795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T14:05:54.535411Z digest=sha256:ceb540cccf166eeb8013681cd35f419b806f8918fa3d83641d52ab0c0c8160c1

Observation f180e1ff-1a80-492e-9bf4-63fd792250d2 · inbound

Orak: A Foundational Benchmark for Training and Evaluating LLM Agents on Diverse Video Games cites this paper.

Orak: A Foundational Benchmark for Training and Evaluating LLM Agents on Diverse Video Games SmartPlay: A Benchmark for LLMs as Intelligent Agents

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T12:02:16.580550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T12:01:42.681135Z digest=sha256:c3cb173251a1f6b95de4ac57d2d1e1b65c006d2f2d13353cab20ff60600760d4

Observation 74aa5be8-8b80-47cf-a647-7f0c6a55d467 · inbound

WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models cites this paper.

WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models SmartPlay: A Benchmark for LLMs as Intelligent Agents

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:38.508070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:34:38.508070Z digest=sha256:23a02e467f721560d9656345df4bb56780248744ef1161a05edb8211f6056757

Observation 5f06527c-6227-443c-8469-78cd2b7844d5 · inbound

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey cites this paper.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey SmartPlay: A Benchmark for LLMs as Intelligent Agents

Reference 107

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.508591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.508591Z digest=sha256:73d36ef46951a8eb76b79dcd444487229f74a16ddcbe3861e22f988aae5e398c

Observation 90b1f3fa-4a8f-437e-850d-17dea2a6b230 · inbound

AgentCrypt: Advancing Privacy and (Secure) Computation in AI Agent Collaboration cites this paper.

AgentCrypt: Advancing Privacy and (Secure) Computation in AI Agent Collaboration SmartPlay: A Benchmark for LLMs as Intelligent Agents

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T23:58:42.681391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T23:57:19.757902Z digest=sha256:588e440bcd8052acf2277d4425f6e001c483604a4a7fb18121871cc4926a68a0

Observation 68ba3c28-5d13-4f4d-88d8-6289f65790c9 · inbound

MINT: Minimal Information Neuro-Symbolic Tree for Objective-Driven Knowledge-Gap Reasoning and Active Elicitation cites this paper.

MINT: Minimal Information Neuro-Symbolic Tree for Objective-Driven Knowledge-Gap Reasoning and Active Elicitation SmartPlay: A Benchmark for LLMs as Intelligent Agents

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T07:17:30.695755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T07:13:42.705619Z digest=sha256:21d9a0a4ff48ce73f510e71e0c2201ccbe2a94b9926498004f3c54c55398fbf2

Observation 709bfa64-b63c-4c12-ab0b-608126e359ed · inbound

LLM-WikiRace Benchmark: How Far Can LLMs Plan over Real-World Knowledge Graphs? cites this paper.

LLM-WikiRace Benchmark: How Far Can LLMs Plan over Real-World Knowledge Graphs? SmartPlay: A Benchmark for LLMs as Intelligent Agents

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T22:26:38.719140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T22:26:38.719140Z digest=sha256:514b9ac9cc27f71f5ca29e69be1ebbe8ed616043fa3260e4785b661134d91b64

Observation 73ef67f3-4894-4d52-a384-f33780a824f8 · inbound

Mastering PokeGym: Graph-Guided Multimodal Evolution at Test Time cites this paper.

Mastering PokeGym: Graph-Guided Multimodal Evolution at Test Time SmartPlay: A Benchmark for LLMs as Intelligent Agents

Reference 80

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:41:00.480146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T17:59:48.877783Z digest=sha256:731fa823836d99159b065c0111501c13b88c271a53427ded6f1e29cd3fc4e8f0

Observation 048b4870-2112-47d2-9c6c-7ad98b9910d2 · inbound

Mastering PokeGym: Graph-Guided Multimodal Evolution at Test Time cites this paper.

Mastering PokeGym: Graph-Guided Multimodal Evolution at Test Time SmartPlay: A Benchmark for LLMs as Intelligent Agents

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-04T05:32:20.455742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:32:20.455742Z digest=sha256:ab80efbf0bcadd7c6194277029ed688c5ce4917784cc8374abeeb50ee7e7d920

Observation 13b47207-a27c-4add-bd35-a40001ed8f36 · inbound

OmniGameArena: A Unified UE5 Benchmark for VLM Game Agents with Improvement Dynamics cites this paper.

OmniGameArena: A Unified UE5 Benchmark for VLM Game Agents with Improvement Dynamics SmartPlay: A Benchmark for LLMs as Intelligent Agents

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:07:29.918510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-27T16:50:36.194650Z digest=sha256:50a29f0a11c797ba0a15768d8eb575076334efc3fbb0270b27c71039bbab0b4c

Observation b21c4984-a132-4d4a-bd1c-6d813fcbbbd6 · inbound

Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application cites this paper.

Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application SmartPlay: A Benchmark for LLMs as Intelligent Agents

Reference 102

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T10:58:02.903663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-27T09:46:30.702256Z digest=sha256:daaeb9618d22204619acef446307c536e7a47a18463dc96d87dce44ff0985bda

Observation 068817dc-a68c-4149-bc26-7f43309f6363 · inbound

Spatial Reasoning in LLM Game Agents: Impact of Causal Context and Multi-Step Planning cites this paper.

Spatial Reasoning in LLM Game Agents: Impact of Causal Context and Multi-Step Planning SmartPlay: A Benchmark for LLMs as Intelligent Agents

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T10:58:16.489441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:58:16.489441Z digest=sha256:1445b4923f4e69ebe19028a7ee77472e9d5437a7183af0bf0b6bf7d77ea5fafd

Observation aa1993d8-c6eb-4706-9ca9-77f389addc25 · inbound

DungeonBench: A Benchmark for Rules-Rich Tactical Reasoning in Dungeons & Dragons Combat cites this paper.

DungeonBench: A Benchmark for Rules-Rich Tactical Reasoning in Dungeons & Dragons Combat SmartPlay: A Benchmark for LLMs as Intelligent Agents

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T04:15:06.235494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:15:06.235494Z digest=sha256:accabc0eef40128f0d42b90425ed4f29943169bd1e9d3a63a6253856577bbb0b