Pith. sign in

Paper Citation Record · LEDGER

BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games

As of 4 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 28 inbound Pith citation observations for arXiv:2411.13543.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.13543 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 28 of 28 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T05:32:20.374639Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-10T06:15:00.866473Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b0f44da1-aa5d-489c-9942-f79c07a69933 · inbound

Orak: A Foundational Benchmark for Training and Evaluating LLM Agents on Diverse Video Games cites this paper.

Orak: A Foundational Benchmark for Training and Evaluating LLM Agents on Diverse Video Games BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-19T12:02:16.708806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T12:01:42.681135Z digest=sha256:f102c92eb31caed465edec683a021a1fc95c7a1eea2e593fec4f934e350f0113

Observation 4b2f3b3e-49b7-459a-ad5b-4048162b3566 · inbound

Towards Understanding, Analyzing, and Optimizing Agentic AI Execution: A CPU-Centric Perspective cites this paper.

Towards Understanding, Analyzing, and Optimizing Agentic AI Execution: A CPU-Centric Perspective BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:55:33.414341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T00:55:14.213746Z digest=sha256:b168d69f7f2c8c96f9623850fe963010f1278c279110840f119d11b945d44da0

Observation 97c2fc39-6466-4cf3-a527-32042fbe3344 · inbound

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization cites this paper.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:27.793264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:27.793264Z digest=sha256:37bfad88fd69ab3b1841f521922c5169f072ceb6d0ee4e81fb3ae968a8a6425a

Observation 947c1b49-e8f2-402d-b200-91c5492f9ae9 · inbound

LLM-WikiRace Benchmark: How Far Can LLMs Plan over Real-World Knowledge Graphs? cites this paper.

LLM-WikiRace Benchmark: How Far Can LLMs Plan over Real-World Knowledge Graphs? BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T22:26:37.148467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T22:26:37.148467Z digest=sha256:c759b5db524995c3f03ae5a0f74a5cdeae296dbf84c87723a5b6e49bacab3b8d

Observation be4c6fd2-0156-40a1-9993-3bf9027a06ee · inbound

Mastering PokeGym: Graph-Guided Multimodal Evolution at Test Time cites this paper.

Mastering PokeGym: Graph-Guided Multimodal Evolution at Test Time BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:41:00.454858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T17:59:48.877783Z digest=sha256:a734c2bda196d7cda3769266b01c765f7e756dcdb6a4ca292e9dc32653bdb837

Observation b3475f7c-cfe5-4a63-b749-b21218b45b5d · inbound

Mastering PokeGym: Graph-Guided Multimodal Evolution at Test Time cites this paper.

Mastering PokeGym: Graph-Guided Multimodal Evolution at Test Time BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T05:32:20.374639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:32:20.374639Z digest=sha256:2d5fa0b910334d9b67d76ff88c8e152c43d621c8a99a27997742df86bf4819b7

Observation ca240d33-c54f-4067-89ec-dc6749ba7bd9 · inbound

Co-Evolving LLM Decision and Skill Bank Agents for Long-Horizon Tasks cites this paper.

Co-Evolving LLM Decision and Skill Bank Agents for Long-Horizon Tasks BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T00:14:46.454940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:14:07.017420Z digest=sha256:d7f6f7bff6d66e0d6fe43bd1b4ed621253f68ba8f995cba8c330b6463d628921

Observation 95274a99-b203-4a2b-97fa-9ba2d1802eb3 · inbound

Hierarchical Behaviour Spaces cites this paper.

Hierarchical Behaviour Spaces BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:01:12.700310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T03:33:48.527621Z digest=sha256:9b3332d9b8a06632377e738f6014e95a10325dc71f959660729b3556b3b5eb53

Observation dbd0bfa9-aaea-431a-a21c-32d6c9e874ee · inbound

Agentick: A Unified Benchmark for General Sequential Decision-Making Agents cites this paper.

Agentick: A Unified Benchmark for General Sequential Decision-Making Agents BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:25:58.628439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T01:22:32.713175Z digest=sha256:e97c1489734ef5e19fb5803ea65dec0a57ea9ab387fd1a67f3ad071a8ad1af9d

Observation de8bb243-1d72-4652-9007-582c9d8d1060 · inbound

Agentick: A Unified Benchmark for General Sequential Decision-Making Agents cites this paper.

Agentick: A Unified Benchmark for General Sequential Decision-Making Agents BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T20:52:57.892871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-14T20:51:26.063471Z digest=sha256:01e98bd9b8ac0a365336c3d7271154b6f1c82f6a39e14af8938fe443b946b5ad

Observation dac7dc5c-5c4d-4969-961b-e11ebdffe1e2 · inbound

Behavior Cue Reasoning: Monitorable Reasoning Improves Efficiency and Safety through Oversight cites this paper.

Behavior Cue Reasoning: Monitorable Reasoning Improves Efficiency and Safety through Oversight BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:00:56.528187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T00:54:25.549158Z digest=sha256:8b13da43e67f96769ef7afdfceff0f3f10cf233880fec0cd457cb28646b62d67

Observation 8960fb7d-4e4f-4233-aa2d-c03a3bfc03aa · inbound

Behavior Cue Reasoning: Monitorable Reasoning Improves Efficiency and Safety through Oversight cites this paper.

Behavior Cue Reasoning: Monitorable Reasoning Improves Efficiency and Safety through Oversight BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:29:52.732321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:29:09.122055Z digest=sha256:c6f063b9177147b55108df2f7cfe02eb8cfc8f6c9895bf06d994958238549c9e

Observation db4efc6b-1dc2-4fe6-8df1-481204036066 · inbound

Towards Generalist Game Players: An Investigation of Foundation Models in the Game Multiverse cites this paper.

Towards Generalist Game Players: An Investigation of Foundation Models in the Game Multiverse BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games

Reference 128

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:26:19.238017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T03:25:24.844859Z digest=sha256:4d90582054a5914b79ed26ede53005988a27f4063a3ae5ff4ab845a61e00edb6

Observation 08a1e57d-28ee-429e-9add-6411e0557f3f · inbound

Towards Generalist Game Players: An Investigation of Foundation Models in the Game Multiverse cites this paper.

Towards Generalist Game Players: An Investigation of Foundation Models in the Game Multiverse BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games

Reference 128

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:47:27.232473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T06:44:28.552513Z digest=sha256:a889590967fb6d44f26f09640f0b3d5669e7fedbfe4cd809337d40b78d96439c

Observation e32d29a5-3527-4721-b0f1-63d19cbcfee8 · inbound

Do Vision-Language-Models show human-like logical problem-solving capability in point and click puzzle games? cites this paper.

Do Vision-Language-Models show human-like logical problem-solving capability in point and click puzzle games? BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T02:02:05.521656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-13T01:58:39.476408Z digest=sha256:45bf975f1e1ae671806e2b77fc2aa0e5d7a7e4e7d2ce376394db59276079bb87

Observation 3f1cd5a9-a3de-47d8-b583-e05ab03d8ada · inbound

Do Vision-Language-Models show human-like logical problem-solving capability in point and click puzzle games? cites this paper.

Do Vision-Language-Models show human-like logical problem-solving capability in point and click puzzle games? BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T02:07:08.791635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-13T01:58:39.476408Z digest=sha256:56a5eb95eb17efd620d58f7eb875372a0c2d9a60d6e4f4f63e4087c88478bf41

Observation 076b6c1a-aa9f-40c5-8d4b-97b2bb7c1fee · inbound

Common-agency Games for Multi-Objective Test-Time Alignment cites this paper.

Common-agency Games for Multi-Objective Test-Time Alignment BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games

Reference 52

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T06:15:06.433297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-15T06:14:53.685486Z digest=sha256:f975ed3f3c69084f3d8cc5cec5da65e86b8533c6d8b544ad3b101d68a309a429

Observation 3078f913-8c75-40fd-b21b-dd69bd2f9666 · inbound

AtlasVA: Self-Evolving Visual Skill Memory for Teacher-Free VLM Agents cites this paper.

AtlasVA: Self-Evolving Visual Skill Memory for Teacher-Free VLM Agents BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T11:28:14.500972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T11:24:48.558423Z digest=sha256:2555cef445c94430a74e8bd9e663849a9998407f0b8d909f9201b27166743aff

Observation 9a439c1a-4186-4b59-bc87-810f66df3813 · inbound

DemoEvolve: Overcoming Sparse Feedback in Agentic Harness Evolution with Demonstrations cites this paper.

DemoEvolve: Overcoming Sparse Feedback in Agentic Harness Evolution with Demonstrations BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:14:40.644739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T13:12:46.927103Z digest=sha256:8f436736a63f6d9452596c606b0ff479879d993239fdd0bedc45883df8b12622

Observation 677af669-2782-48b4-85fa-82fe1bbc86c3 · inbound

Resonant Minds: Closed-Loop Social Avatars with Theory of Mind cites this paper.

Resonant Minds: Closed-Loop Social Avatars with Theory of Mind BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T12:36:56.229554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T02:04:39.753443Z digest=sha256:73f6f30dd3633a51807c1de81924bad803ecd194ab6567f0c42ddb6a89b0db8a

Observation 6419eb26-cc55-40a9-af78-ea8bd5f3c81a · inbound

OmniGameArena: A Unified UE5 Benchmark for VLM Game Agents with Improvement Dynamics cites this paper.

OmniGameArena: A Unified UE5 Benchmark for VLM Game Agents with Improvement Dynamics BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:07:29.936088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-27T16:50:36.194650Z digest=sha256:85f118905521b539114586a2e6728a6c7d21816b73f875d9167ccaaf9f600611

Observation 36304c02-9475-4d6a-9927-07b9e08406fd · inbound

Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application cites this paper.

Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games

Reference 114

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:58:02.898623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-27T09:46:30.702256Z digest=sha256:f60a1dcbecf846e841e3065a7383c9bb864554c7af524ffd0355074fcf05b261

Observation 7af6216c-f02c-45ab-bd11-77827173cf68 · inbound

AgentBeats: Agentifying Agent Assessment for Openness, Standardization, and Reproducibility cites this paper.

AgentBeats: Agentifying Agent Assessment for Openness, Standardization, and Reproducibility BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:08:33.114440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-27T06:41:41.799596Z digest=sha256:ad9118b812d6dd972dc0a578802de712ed6dabe54f4d0d309bb6369c8f4f030f

Observation e19c1c44-97b3-4b3e-bd77-ee3fafaabd4b · inbound

RTSGameBench: An RTS Benchmark for Strategic Reasoning by Vision-Language Models cites this paper.

RTSGameBench: An RTS Benchmark for Strategic Reasoning by Vision-Language Models BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T00:39:17.286611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T21:02:26.081166Z digest=sha256:36043c3d2ee5ba091daacd596ce270f6646e4b9bdb8b796288b13996d77e3b94

Observation ddf4e3cd-fae7-4de7-9a7a-3ccc871109a4 · inbound

Beyond the Current Observation: Evaluating Multimodal Large Language Models in Controllable Non-Markov Games cites this paper.

Beyond the Current Observation: Evaluating Multimodal Large Language Models in Controllable Non-Markov Games BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:19:13.650685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T21:17:02.332687Z digest=sha256:e2c0f7e9eb6883f0cfd39f2b1bea13f63439b5753888d3c71080917d7a046f16

Observation 21f38e44-3766-45fb-99a2-d083342a44eb · inbound

AutoMem: Automated Learning of Memory as a Cognitive Skill cites this paper.

AutoMem: Automated Learning of Memory as a Cognitive Skill BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:16:56.514755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-02T12:09:32.053764Z digest=sha256:ce40b3298f4e098a0e052b5b23ca87142ebc500e241d6374e65d8d9106bc8e02

Observation a20c9f69-bbab-485a-8f42-99b3d89d43fd · inbound

Spatial Reasoning in LLM Game Agents: Impact of Causal Context and Multi-Step Planning cites this paper.

Spatial Reasoning in LLM Game Agents: Impact of Causal Context and Multi-Step Planning BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T10:58:16.786880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:58:16.786880Z digest=sha256:b5baaf37f6845f39308a7470d9f7eeae4eae3c9e224bfd6052276cbd5520c9a3

Observation c3cfa10d-5cd9-48a2-8036-69dc7653fd14 · inbound

DungeonBench: A Benchmark for Rules-Rich Tactical Reasoning in Dungeons & Dragons Combat cites this paper.

DungeonBench: A Benchmark for Rules-Rich Tactical Reasoning in Dungeons & Dragons Combat BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T04:15:06.000334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:15:06.000334Z digest=sha256:afcf164a005d7791dfa0613741176c7b8f772d3a8a019cf5cbf06070aca4dc60