Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 4 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 28 inbound Pith citation observations for arXiv:2411.13543.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T05:32:20.374639Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-10T06:15:00.866473Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation b0f44da1-aa5d-489c-9942-f79c07a69933 · inbound
Orak: A Foundational Benchmark for Training and Evaluating LLM Agents on Diverse Video Games BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4b2f3b3e-49b7-459a-ad5b-4048162b3566 · inbound
Towards Understanding, Analyzing, and Optimizing Agentic AI Execution: A CPU-Centric Perspective BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 97c2fc39-6466-4cf3-a527-32042fbe3344 · inbound
Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 947c1b49-e8f2-402d-b200-91c5492f9ae9 · inbound
LLM-WikiRace Benchmark: How Far Can LLMs Plan over Real-World Knowledge Graphs? BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be4c6fd2-0156-40a1-9993-3bf9027a06ee · inbound
Mastering PokeGym: Graph-Guided Multimodal Evolution at Test Time BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b3475f7c-cfe5-4a63-b749-b21218b45b5d · inbound
Mastering PokeGym: Graph-Guided Multimodal Evolution at Test Time BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca240d33-c54f-4067-89ec-dc6749ba7bd9 · inbound
Co-Evolving LLM Decision and Skill Bank Agents for Long-Horizon Tasks BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 95274a99-b203-4a2b-97fa-9ba2d1802eb3 · inbound
Hierarchical Behaviour Spaces BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation dbd0bfa9-aaea-431a-a21c-32d6c9e874ee · inbound
Agentick: A Unified Benchmark for General Sequential Decision-Making Agents BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation de8bb243-1d72-4652-9007-582c9d8d1060 · inbound
Agentick: A Unified Benchmark for General Sequential Decision-Making Agents BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation dac7dc5c-5c4d-4969-961b-e11ebdffe1e2 · inbound
Behavior Cue Reasoning: Monitorable Reasoning Improves Efficiency and Safety through Oversight BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8960fb7d-4e4f-4233-aa2d-c03a3bfc03aa · inbound
Behavior Cue Reasoning: Monitorable Reasoning Improves Efficiency and Safety through Oversight BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation db4efc6b-1dc2-4fe6-8df1-481204036066 · inbound
Towards Generalist Game Players: An Investigation of Foundation Models in the Game Multiverse BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games
Reference 128
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 08a1e57d-28ee-429e-9add-6411e0557f3f · inbound
Towards Generalist Game Players: An Investigation of Foundation Models in the Game Multiverse BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games
Reference 128
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e32d29a5-3527-4721-b0f1-63d19cbcfee8 · inbound
Do Vision-Language-Models show human-like logical problem-solving capability in point and click puzzle games? BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3f1cd5a9-a3de-47d8-b583-e05ab03d8ada · inbound
Do Vision-Language-Models show human-like logical problem-solving capability in point and click puzzle games? BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 076b6c1a-aa9f-40c5-8d4b-97b2bb7c1fee · inbound
Common-agency Games for Multi-Objective Test-Time Alignment BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3078f913-8c75-40fd-b21b-dd69bd2f9666 · inbound
AtlasVA: Self-Evolving Visual Skill Memory for Teacher-Free VLM Agents BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9a439c1a-4186-4b59-bc87-810f66df3813 · inbound
DemoEvolve: Overcoming Sparse Feedback in Agentic Harness Evolution with Demonstrations BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 677af669-2782-48b4-85fa-82fe1bbc86c3 · inbound
Resonant Minds: Closed-Loop Social Avatars with Theory of Mind BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6419eb26-cc55-40a9-af78-ea8bd5f3c81a · inbound
OmniGameArena: A Unified UE5 Benchmark for VLM Game Agents with Improvement Dynamics BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 36304c02-9475-4d6a-9927-07b9e08406fd · inbound
Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games
Reference 114
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7af6216c-f02c-45ab-bd11-77827173cf68 · inbound
AgentBeats: Agentifying Agent Assessment for Openness, Standardization, and Reproducibility BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e19c1c44-97b3-4b3e-bd77-ee3fafaabd4b · inbound
RTSGameBench: An RTS Benchmark for Strategic Reasoning by Vision-Language Models BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ddf4e3cd-fae7-4de7-9a7a-3ccc871109a4 · inbound
Beyond the Current Observation: Evaluating Multimodal Large Language Models in Controllable Non-Markov Games BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 21f38e44-3766-45fb-99a2-d083342a44eb · inbound
AutoMem: Automated Learning of Memory as a Cognitive Skill BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a20c9f69-bbab-485a-8f42-99b3d89d43fd · inbound
Spatial Reasoning in LLM Game Agents: Impact of Causal Context and Multi-Step Planning BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3cfa10d-5cd9-48a2-8036-69dc7653fd14 · inbound
DungeonBench: A Benchmark for Rules-Rich Tactical Reasoning in Dungeons & Dragons Combat BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.