Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T22:36:46.166368Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 0 inbound Pith citation observations for arXiv:2506.21252.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T22:36:46.166368Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
47 of 47 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation f06f09b6-722d-4be3-b938-e09a910cd995 · outbound
Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9725d029-a434-408e-b4f7-1a3a251fada7 · outbound
Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents GPT-4 Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c030d5ff-4f4e-4a31-95cd-2127a6cf995a · outbound
Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents Pixtral 12B
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69c1f3e3-825d-45de-a713-bd71074f7788 · outbound
Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 17a169ef-baee-4141-a20b-587776fef169 · outbound
Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents Large Language Models for Planning: A Comprehensive and Systematic Survey
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cfbb0d56-e472-4344-ab30-6a7a09779e7c · outbound
Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents FireAct: Toward Language Agent Fine-tuning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d722460-63ac-4ef0-bdcd-672c1891413e · outbound
Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents Towards End-to-End Embodied Decision Making via Multi-modal Large Language Model: Explorations with GPT4-Vision and Beyond
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b15e6fa-9b30-40d4-8df5-08b7af498f81 · outbound
Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents PCA-Bench: Evaluating Multimodal Large Language Models in Perception-Cognition-Action Chain
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bdb8f448-6613-409d-8391-bf6d51358543 · outbound
Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents Unresolved cited work
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 40593145-71eb-4183-9f85-5f38e66d7558 · outbound
Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e650d70f-b58b-4084-ab85-5bdf84073b81 · outbound
Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents Unresolved cited work
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c4f224b8-47a4-4c60-867b-e6dd639c5596 · outbound
Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d349826a-2549-4668-ab70-27a71dc6a5f6 · outbound
Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents The Llama 3 Herd of Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12488757-e78c-4a47-bf9a-8b41e6bbb716 · outbound
Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents Unresolved cited work
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation af810d31-6b93-466a-aae4-81e72e25c7fa · outbound
Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents Is Your LLM Secretly a World Model of the Internet? Model-Based Planning for Web Agents
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11592a68-2150-4cfa-9ff9-147ac89d1dbf · outbound
Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents GPT-4o System Card
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 563cd57d-4319-48dc-a29a-5bdf892a5138 · outbound
Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents RAG-RewardBench: Benchmarking Reward Models in Retrieval Augmented Generation for Preference Alignment
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55131c7c-5202-4653-8ad2-711d937d028a · outbound
Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 168d3974-78aa-4789-8369-2f46f9d735fc · outbound
Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents Unresolved cited work
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d3bb893-9a25-4e4d-848c-60ef6b79823b · outbound
Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents Gonzalez, Hao Zhang, and Ion Stoica
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c9e2b1f-c263-474e-aa11-0fce5b7db5a9 · outbound
Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents LLaVA-OneVision: Easy Visual Task Transfer
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4bef8c0-b8c5-474b-8a7d-b23863924a88 · outbound
Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents VL-RewardBench: A Challenging Benchmark for Vision-Language Generative Reward Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c681772-ba8e-4653-83e6-f73694c759db · outbound
Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents QLASS: Boosting Language Agent Inference via Q-Guided Stepwise Search
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 027d2cb7-8b0b-4b7a-a572-cc99be388e3e · outbound
Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents Unlocking the Future: Exploring Look-Ahead Planning Mechanistic Interpretability in Large Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0cb8f075-81df-44af-a819-b23870384605 · outbound
Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36cc2361-8787-4bb2-b086-cdebe9bba711 · outbound
Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de0dac99-505e-4fda-9f2a-d93267561f42 · outbound
Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents Trial and Error: Exploration-Based Trajectory Optimization for LLM Agents
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3dc86184-7470-46bd-ae30-9af0b0e30b48 · outbound
Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task Synthesis
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3bd24895-869d-4623-ae1a-74c79815fa6f · outbound
Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eaf61ff5-95e3-42b3-8bb5-622ab65d8b6c · outbound
Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c46d6f02-f238-456e-8f5c-127d2a1518c0 · outbound
Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents Describe, Explain, Plan and Select: Interactive Planning with Large Language Models Enables Open-World Multi-Task Agents
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9bd1a4cf-3b2f-4787-ad6c-73ff556a3023 · outbound
Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents Unresolved cited work
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97b2a971-9d40-41bc-86c7-83e970d60b63 · outbound
Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents VSP: Assessing the dual challenges of perception and reasoning in spatial planning tasks for VLMs
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e0486c0-ab10-4aaf-a464-eb7c2fc8e9d9 · outbound
Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents TravelPlanner: A Benchmark for Real-World Planning with Language Agents
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eac92648-2a35-4996-8114-fa90e82c1a78 · outbound
Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a412a3c8-72dc-466b-9fde-208523838d65 · outbound
Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents Unresolved cited work
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5d209ea6-adfe-47c5-bf66-a0419f304ee3 · outbound
Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents Unresolved cited work
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7f309774-b59f-4f7f-be72-d7885afdbca8 · outbound
Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents AgentTuning: Enabling Generalized Agent Abilities for LLMs
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72928ad8-48a3-482c-9f98-3429532b762e · outbound
Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents Enhancing Decision-Making for LLM Agents via Step-Level Q-Value Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b81944de-802d-48d4-8d7c-31505eeed8c7 · outbound
Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents MFE-ETP: A Comprehensive Evaluation Benchmark for Multi-modal Foundation Models on Embodied Task Planning
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a5f0f52-67b2-4554-a6b8-26260ebf08cb · outbound
Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents Attacking Vision-Language Computer Agents via Pop-ups
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f1ab5c3-d448-454f-9bf4-2ccd1f9dac4a · outbound
Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents WebPilot: A Versatile and Autonomous Multi-Agent System for Web Task Execution with Strategic Exploration
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be8ff1ed-8400-433b-aa1e-e13b8a79b760 · outbound
Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents RMB: Comprehensively Benchmarking Reward Models in LLM Alignment
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59d6978a-d62d-4dbf-9edc-cd76423a9675 · outbound
Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents Multimodal Situational Safety
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83c0accd-e77a-4a31-b5f9-5fa06e6e5ad1 · outbound
Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents WebArena: A Realistic Web Environment for Building Autonomous Agents
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3493dd4f-8f64-4b55-97a7-1aaf7321bd13 · outbound
Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents online" 'onlinestring :=
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1d9ba60-f5b1-496e-bbd3-e045859f0fc6 · outbound
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.