Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 15 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 23 inbound Pith citation observations for arXiv:2310.12921.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-12T16:45:47.211087Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
6
pith, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation 0db6c1cc-34a3-4cf8-a65a-5ede233e8ad4 · inbound
VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 1c616136-d84a-46b6-ad8c-e02bdf1af33c · inbound
ViSTa Dataset: Do vision-language models understand sequential tasks? Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6765abe-bdc9-49d1-a9c3-8e254d7608b9 · inbound
LLM-Based Offline Learning for Embodied Agents via Consistency-Guided Reward Ensemble Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 809d90c1-9ee9-4b2c-9e03-0a98c76fa23c · inbound
YingSound: Video-Guided Sound Effects Generation with Multi-modal Chain-of-Thought Controls Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5df7d45-d240-484d-a2a1-5d28d26010c9 · inbound
CLIP-RLDrive: Human-Aligned Autonomous Driving via CLIP-Based Reward Shaping in Reinforcement Learning Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5df212af-6864-4c05-bcde-c381e175389c · inbound
Preference VLM: Leveraging VLMs for Scalable Preference-Based Reinforcement Learning Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0df10cc-6fa3-4edf-bb43-6e49a0db3b95 · inbound
Advancing Autonomous VLM Agents via Variational Subgoal-Conditioned Reinforcement Learning Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6282dfeb-c2a8-4da9-88bf-5d1d1f081959 · inbound
RobotSmith: Generative Robotic Tool Design for Acquisition of Complex Manipulation Skills Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8732b67-dd5e-4cdb-ba09-daf077b9ee65 · inbound
FOUNDER: Grounding Foundation Models in World Models for Open-Ended Embodied Decision Making Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7be5cd62-64b7-4087-8c1e-9aa52721930d · inbound
A Survey on Generative Recommendation: Data, Model, and Tasks Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning
Reference 152
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation b3117d81-fdb0-4c25-8f81-157f93d8ecd7 · inbound
TOPReward: Token Probabilities as Hidden Zero-Shot Rewards for Robotics Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning
Reference 1991
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 326ddb2c-398e-47b5-8d62-9041f74f5f33 · inbound
SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3929100-c867-41f6-b196-3babab1b8bd3 · inbound
AtlasVA: Self-Evolving Visual Skill Memory for Teacher-Free VLM Agents Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 7525be9d-48ed-4c8c-92e6-7b6dd4f9ed10 · inbound
OrganicHAR: Towards Activity Discovery in Organic Settings for Privacy Preserving Sensors Using Efficient Video Analysis Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 611fa5dd-27c9-484c-a022-5b5050bec545 · inbound
World Model Self-Distillation: Training World Models to Solve General Tasks Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 06b70759-44ff-45b7-9b5d-d60bab24f3e9 · inbound
Improving Robotic Generalist Policies via Flow Reversal Steering Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 9a3f352e-611b-48fa-a509-8c82461e9ba1 · inbound
Learning Process Rewards via Success Visitation Matching for Efficient RL Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation ed8f13a1-8017-4500-965e-fe11ff42e7c2 · inbound
Vision-Language Models for Deployable Social Robot Navigation: Bridging Semantic Reasoning and Low-Level Control Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation bb0cf977-af77-4644-a776-9e4236e5ddd0 · inbound
LLM-as-a-Verifier: A General-Purpose Verification Framework Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation d27b0c92-b711-49ec-a398-c26a00ba82a1 · inbound
LLM-as-a-Verifier: A General-Purpose Verification Framework Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73bebea9-1c91-4db5-ae22-76f5ed84a3c7 · inbound
PhyAgentOS: A Self-Evolving Operating System for Embodied Agents with Decoupled Cognitive Planning and Physical Execution Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9eff0a4f-c6d2-4a5d-b544-92bb455fb36f · inbound
MARS-RA: Rank Aggregation for Credit Assignment via Multimodal Comparisons in Embodied Multi-Agent Cooperation Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad616c7d-68a3-4d77-9ff1-1bf7b6e5b624 · inbound
Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning
Reference 216
Source-reported events for the cited work
Unavailable: canonical work link unavailable.