Pith. sign in

Paper Citation Record · LEDGER

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents

As of 7 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 0 inbound Pith citation observations for arXiv:2506.21252.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.21252 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:36:46.166368Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

47 of 47 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved47
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f06f09b6-722d-4be3-b938-e09a910cd995 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:41.369605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:41.369605Z digest=sha256:86812036096de3a85d0456313c59d09d82a7d6feb648b23de32d9a8824c21688

Observation 9725d029-a434-408e-b4f7-1a3a251fada7 · outbound

This paper cites GPT-4 Technical Report.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:41.438405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:41.438405Z digest=sha256:022d917de5ccee4b838a27469f7f9a5cff0aeab71c924df911db2442041417dd

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:41.685574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:41.685574Z digest=sha256:5d31f1196ff45e7dd5af35bcb2a888c1f39b65c4a5487a806fb40de2a7501d28

Observation 69c1f3e3-825d-45de-a713-bd71074f7788 · outbound

This paper cites an unresolved cited work.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:36:48.332392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:36:41.826052Z digest=sha256:bd8ba9a7b39c82b5368933f8a598ceb4fdc5abe05fea39c2c7414820f1e04afe

Observation 17a169ef-baee-4141-a20b-587776fef169 · outbound

This paper cites Large Language Models for Planning: A Comprehensive and Systematic Survey.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents Large Language Models for Planning: A Comprehensive and Systematic Survey

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:41.978125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:41.978125Z digest=sha256:cc2bfec69e0f81996723355b4c43d5d3d9fb0e59b17b8bd346ebcdcd0da2f8f3

Observation cfbb0d56-e472-4344-ab30-6a7a09779e7c · outbound

This paper cites FireAct: Toward Language Agent Fine-tuning.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents FireAct: Toward Language Agent Fine-tuning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:42.146724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:42.146724Z digest=sha256:d499da9c595d6de0e6b96bc95ffdd3ff158854bf33df5407d7aa5356ea8d160a

Observation 2d722460-63ac-4ef0-bdcd-672c1891413e · outbound

This paper cites Towards End-to-End Embodied Decision Making via Multi-modal Large Language Model: Explorations with GPT4-Vision and Beyond.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents Towards End-to-End Embodied Decision Making via Multi-modal Large Language Model: Explorations with GPT4-Vision and Beyond

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:42.313802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:42.313802Z digest=sha256:05b31320d173311f51135e8d12a3f17bdcc4de7206bbcb0902f23505576045ac

Observation 8b15e6fa-9b30-40d4-8df5-08b7af498f81 · outbound

This paper cites PCA-Bench: Evaluating Multimodal Large Language Models in Perception-Cognition-Action Chain.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents PCA-Bench: Evaluating Multimodal Large Language Models in Perception-Cognition-Action Chain

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:42.406234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:42.406234Z digest=sha256:9022906432b03c8dd8d3f32b0e78d77310cbe9d5b58a314f27e2796261ce1c11

Observation bdb8f448-6613-409d-8391-bf6d51358543 · outbound

This paper cites an unresolved cited work.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:36:48.031390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:36:42.497421Z digest=sha256:30a9ef824785a0fbb5e5a711e2f6d9f1e436ae36ce92172019f617843d2ad9e5

Observation 40593145-71eb-4183-9f85-5f38e66d7558 · outbound

This paper cites SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:42.561879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:42.561879Z digest=sha256:79b1b1315a1c1610c2d41efdb451aee8508f46549540a0adeac7b00a40c8cbfb

Observation e650d70f-b58b-4084-ab85-5bdf84073b81 · outbound

This paper cites an unresolved cited work.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:36:47.726733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:36:42.625120Z digest=sha256:e6e0e37c33abaa6471a871ec3dffd21a206f5a4d5fccb1d4680825b28eb05337

Observation c4f224b8-47a4-4c60-867b-e6dd639c5596 · outbound

This paper cites an unresolved cited work.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:36:47.536393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:36:42.687111Z digest=sha256:2d8f0959f941705c8ed36177b7e7de8d8dad22e6f71c93dc1eb5325dfa19a292

Observation d349826a-2549-4668-ab70-27a71dc6a5f6 · outbound

This paper cites The Llama 3 Herd of Models.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents The Llama 3 Herd of Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:42.774037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:42.774037Z digest=sha256:dff325069a1fb0aaa9c30cc5557abf7f964a6e9c18292363ea4982816a26bac4

Observation 12488757-e78c-4a47-bf9a-8b41e6bbb716 · outbound

This paper cites an unresolved cited work.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:36:47.296199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:36:42.844408Z digest=sha256:4469aab3d8ceca4ddbd30f47a27ee2439f36e0eb1aafbe47f9416a0b91784ee9

Observation af810d31-6b93-466a-aae4-81e72e25c7fa · outbound

This paper cites Is Your LLM Secretly a World Model of the Internet? Model-Based Planning for Web Agents.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents Is Your LLM Secretly a World Model of the Internet? Model-Based Planning for Web Agents

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:42.927894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:42.927894Z digest=sha256:407b39b14152fa6cae26add1fbab42f4e05e45cf41387bab8f6ed82ae12ba946

Observation 11592a68-2150-4cfa-9ff9-147ac89d1dbf · outbound

This paper cites GPT-4o System Card.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents GPT-4o System Card

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:43.013554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:43.013554Z digest=sha256:9ca610966deb949e43c92618a9af02101971a880971f9582f31a075b527f5957

Observation 563cd57d-4319-48dc-a29a-5bdf892a5138 · outbound

This paper cites RAG-RewardBench: Benchmarking Reward Models in Retrieval Augmented Generation for Preference Alignment.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents RAG-RewardBench: Benchmarking Reward Models in Retrieval Augmented Generation for Preference Alignment

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:43.134536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:43.134536Z digest=sha256:018cfbc00fa9234633890a90c03651d35de82d5e453560402e0dfd4ce66cb65f

Observation 55131c7c-5202-4653-8ad2-711d937d028a · outbound

This paper cites VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:43.222340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:43.222340Z digest=sha256:52c74dc806e9037409b2d54f57c1b73ffd90f5ff2aa1dca39cdbee53e3215e16

Observation 168d3974-78aa-4789-8369-2f46f9d735fc · outbound

This paper cites an unresolved cited work.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:43.283968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:43.283968Z digest=sha256:6332ed4b25296e990004477f1fd820b82995bac97920df353e20469f144047a5

Observation 2d3bb893-9a25-4e4d-848c-60ef6b79823b · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents Gonzalez, Hao Zhang, and Ion Stoica

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:43.360429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:43.360429Z digest=sha256:33a56cbbd329abbd11d609340b01e78ee035c144237210e6eaaeb09ac40086cc

Observation 6c9e2b1f-c263-474e-aa11-0fce5b7db5a9 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents LLaVA-OneVision: Easy Visual Task Transfer

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:43.438047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:43.438047Z digest=sha256:a96d692937fd1d00ad3faee5cf36d49fcb5aa08279f3c3740f1ef3abaa976eca

Observation c4bef8c0-b8c5-474b-8a7d-b23863924a88 · outbound

This paper cites VL-RewardBench: A Challenging Benchmark for Vision-Language Generative Reward Models.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents VL-RewardBench: A Challenging Benchmark for Vision-Language Generative Reward Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:43.531337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:43.531337Z digest=sha256:7dbae714409ec84e66bc148165a5df0cd6b84e9061b9409ff3902a8de969f4f3

Observation 2c681772-ba8e-4653-83e6-f73694c759db · outbound

This paper cites QLASS: Boosting Language Agent Inference via Q-Guided Stepwise Search.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents QLASS: Boosting Language Agent Inference via Q-Guided Stepwise Search

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:43.620059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:43.620059Z digest=sha256:ea4b7f10db3f2623ef24ab84541ffbaa0a0e19ce62081fd60c909f5dd240c1eb

Observation 027d2cb7-8b0b-4b7a-a572-cc99be388e3e · outbound

This paper cites Unlocking the Future: Exploring Look-Ahead Planning Mechanistic Interpretability in Large Language Models.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents Unlocking the Future: Exploring Look-Ahead Planning Mechanistic Interpretability in Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:43.709254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:43.709254Z digest=sha256:82309299e460210caa8fb12bdab4814a2fcfe98cdbe20ac6ef92386f4b9bdea4

Observation 0cb8f075-81df-44af-a819-b23870384605 · outbound

This paper cites EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:43.779856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:43.779856Z digest=sha256:6a7ec700eb78f59fd00c90f2381f0e8eb38271ee8629bb47b6e9fd06400be158

Observation 36cc2361-8787-4bb2-b086-cdebe9bba711 · outbound

This paper cites AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:43.891326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:43.891326Z digest=sha256:22c09edb55ee68e16b9f413fbab751aca2d5b6f1293e11b9c3a03d4fcabae02b

Observation de0dac99-505e-4fda-9f2a-d93267561f42 · outbound

This paper cites Trial and Error: Exploration-Based Trajectory Optimization for LLM Agents.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents Trial and Error: Exploration-Based Trajectory Optimization for LLM Agents

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:43.969795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:43.969795Z digest=sha256:2bfcd47af9d34f998c57186aaa2a72a19d5350b4e1fe03ee3e0c1ee34652ee9d

Observation 3dc86184-7470-46bd-ae30-9af0b0e30b48 · outbound

This paper cites OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task Synthesis.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task Synthesis

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:44.079391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:44.079391Z digest=sha256:a578b1dca48c34a3e2c24224c45b2496985968a629dca4a83744a19cff0706df

Observation 3bd24895-869d-4623-ae1a-74c79815fa6f · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:44.193779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:44.193779Z digest=sha256:917535db52a50a036f1a3aedeba703a4d218be20a12b2013ec2dfb7e246f7b98

Observation eaf61ff5-95e3-42b3-8bb5-622ab65d8b6c · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:44.304809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:44.304809Z digest=sha256:b1dd8b5960c52acf672dcab66c3548bc9ca846ef7a6eaae79e0e26ea407de1bf

Observation c46d6f02-f238-456e-8f5c-127d2a1518c0 · outbound

This paper cites Describe, Explain, Plan and Select: Interactive Planning with Large Language Models Enables Open-World Multi-Task Agents.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents Describe, Explain, Plan and Select: Interactive Planning with Large Language Models Enables Open-World Multi-Task Agents

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:44.366489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:44.366489Z digest=sha256:93dbb0a4d2f0ad99d4399aa2d190000a89d62285fb1ffc49a39870bf032fab04

Observation 9bd1a4cf-3b2f-4787-ad6c-73ff556a3023 · outbound

This paper cites an unresolved cited work.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:44.432191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:44.432191Z digest=sha256:f04d0878d2e77803f7ea900c34a6fdff9e4944aa9af9d083e70419e4e941d432

Observation 97b2a971-9d40-41bc-86c7-83e970d60b63 · outbound

This paper cites VSP: Assessing the dual challenges of perception and reasoning in spatial planning tasks for VLMs.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents VSP: Assessing the dual challenges of perception and reasoning in spatial planning tasks for VLMs

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:44.501851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:44.501851Z digest=sha256:bf231473c718a3a80b29f01c64da56d02ed01e85cde15550e1f05006fe71e398

Observation 8e0486c0-ab10-4aaf-a464-eb7c2fc8e9d9 · outbound

This paper cites TravelPlanner: A Benchmark for Real-World Planning with Language Agents.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents TravelPlanner: A Benchmark for Real-World Planning with Language Agents

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:44.505963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:44.505963Z digest=sha256:b36d4d771ece3879a3fc7793eabcdd87cda86446c9c61e5dc7321c89a296239a

Observation eac92648-2a35-4996-8114-fa90e82c1a78 · outbound

This paper cites OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:44.585127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:44.585127Z digest=sha256:ce07524962c97e1aa4533b6626dc6d7b48643106d1703018b09be00a44a49371

Observation a412a3c8-72dc-466b-9fde-208523838d65 · outbound

This paper cites an unresolved cited work.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:36:47.087466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:36:44.721148Z digest=sha256:ff143c1725fbac0ae2ceb47cc94eee5b9c8070c26ae5165396b82ffb697f355c

Observation 5d209ea6-adfe-47c5-bf66-a0419f304ee3 · outbound

This paper cites an unresolved cited work.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:36:46.860630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:36:44.901469Z digest=sha256:359b0cd07126b023e524ef0679f735ece8d98884ca15f6af7206e1915590a36b

Observation 7f309774-b59f-4f7f-be72-d7885afdbca8 · outbound

This paper cites AgentTuning: Enabling Generalized Agent Abilities for LLMs.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents AgentTuning: Enabling Generalized Agent Abilities for LLMs

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:45.082095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:45.082095Z digest=sha256:82368b28c9cebd098b1a1db0ad92cbfea26405d0e15434d98535db41ee4a7281

Observation 72928ad8-48a3-482c-9f98-3429532b762e · outbound

This paper cites Enhancing Decision-Making for LLM Agents via Step-Level Q-Value Models.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents Enhancing Decision-Making for LLM Agents via Step-Level Q-Value Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:45.285843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:45.285843Z digest=sha256:4611209bbea4695e8d36ece96a0561b065c0aa15d0f9b4dffece81870bde8655

Observation b81944de-802d-48d4-8d7c-31505eeed8c7 · outbound

This paper cites MFE-ETP: A Comprehensive Evaluation Benchmark for Multi-modal Foundation Models on Embodied Task Planning.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents MFE-ETP: A Comprehensive Evaluation Benchmark for Multi-modal Foundation Models on Embodied Task Planning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:45.463037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:45.463037Z digest=sha256:60b17cda62d4cb0eb5234e297cf93150c579b9b42b36a4bb81ed702662171d0e

Observation 8a5f0f52-67b2-4554-a6b8-26260ebf08cb · outbound

This paper cites Attacking Vision-Language Computer Agents via Pop-ups.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents Attacking Vision-Language Computer Agents via Pop-ups

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:45.668547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:45.668547Z digest=sha256:c6102a378a8d250a9808eb41b30cc8d2792886d439dff53503a8d024443b333a

Observation 9f1ab5c3-d448-454f-9bf4-2ccd1f9dac4a · outbound

This paper cites WebPilot: A Versatile and Autonomous Multi-Agent System for Web Task Execution with Strategic Exploration.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents WebPilot: A Versatile and Autonomous Multi-Agent System for Web Task Execution with Strategic Exploration

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:45.814121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:45.814121Z digest=sha256:1977a703a85b5978ba3cd8b92a0befa07518ac665188f2c12b4ccf493288d02c

Observation be8ff1ed-8400-433b-aa1e-e13b8a79b760 · outbound

This paper cites RMB: Comprehensively Benchmarking Reward Models in LLM Alignment.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents RMB: Comprehensively Benchmarking Reward Models in LLM Alignment

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:45.894072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:45.894072Z digest=sha256:44973b110f0fb64db1f4b63e3fb3bacbf80a8a63efd39eb4daabfb39a36a97ad

Observation 59d6978a-d62d-4dbf-9edc-cd76423a9675 · outbound

This paper cites Multimodal Situational Safety.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents Multimodal Situational Safety

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:45.933900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:45.933900Z digest=sha256:41b19e4ae3a081ff442988ab5efa1b7d6c023ee2e7c1fb04da317c9ad86bdb4c

Observation 83c0accd-e77a-4a31-b5f9-5fa06e6e5ad1 · outbound

This paper cites WebArena: A Realistic Web Environment for Building Autonomous Agents.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents WebArena: A Realistic Web Environment for Building Autonomous Agents

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:45.975203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:45.975203Z digest=sha256:e997a2cfd0fa477d2edf4270b7e92d3cce800568a1edcea73df5617eb694d680

Observation 3493dd4f-8f64-4b55-97a7-1aaf7321bd13 · outbound

This paper cites online" 'onlinestring :=.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents online" 'onlinestring :=

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:46.052241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:46.052241Z digest=sha256:5137928110767e76b3429b9fb26b41bffd2c87fa7b74dbc90afca4db6c4c075c

Observation b1d9ba60-f5b1-496e-bbd3-e045859f0fc6 · outbound

This paper cites write newline.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents write newline

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:46.166368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:46.166368Z digest=sha256:1948e9ef5c8386165f93a0a83a589c814a12c882b9a79111c112d156990120ec

Pith citing papers

No inbound Pith citation observations are available.