Pith. sign in

Paper Citation Record · LEDGER

Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning

As of 15 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 23 inbound Pith citation observations for arXiv:2310.12921.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.12921 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 23 of 23 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T16:45:47.211087Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

6
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0db6c1cc-34a3-4cf8-a65a-5ede233e8ad4 · inbound

VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models cites this paper.

VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning

Reference 87

Resolution
verified exact
arxiv_id, observed 2026-05-13T08:57:22.363001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-13T08:57:22.299028Z digest=sha256:869ad7885019d4d99a1562014438b6454159bb892631f7a30d79c9d47bc13ac8

Observation 1c616136-d84a-46b6-ad8c-e02bdf1af33c · inbound

ViSTa Dataset: Do vision-language models understand sequential tasks? cites this paper.

ViSTa Dataset: Do vision-language models understand sequential tasks? Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:47.211087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:47.211087Z digest=sha256:4b2fe7cdf6fe2734980ea2dc9f52ecec81b73e9e3c7d1dce4a26d83ea38a211a

Observation a6765abe-bdc9-49d1-a9c3-8e254d7608b9 · inbound

LLM-Based Offline Learning for Embodied Agents via Consistency-Guided Reward Ensemble cites this paper.

LLM-Based Offline Learning for Embodied Agents via Consistency-Guided Reward Ensemble Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T12:33:07.124273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:33:07.124273Z digest=sha256:43f2c1803c9dfb3e18323f4a11eea4f1b83b3f3b0f533fc63a4c599d81e60df7

Observation 809d90c1-9ee9-4b2c-9e03-0a98c76fa23c · inbound

YingSound: Video-Guided Sound Effects Generation with Multi-modal Chain-of-Thought Controls cites this paper.

YingSound: Video-Guided Sound Effects Generation with Multi-modal Chain-of-Thought Controls Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-11T17:17:49.468728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:17:49.468728Z digest=sha256:9d6c755dded14f7acada3ff7b08cb499e525b13ea5bf2b4bfc75a203206ec993

Observation c5df7d45-d240-484d-a2a1-5d28d26010c9 · inbound

CLIP-RLDrive: Human-Aligned Autonomous Driving via CLIP-Based Reward Shaping in Reinforcement Learning cites this paper.

CLIP-RLDrive: Human-Aligned Autonomous Driving via CLIP-Based Reward Shaping in Reinforcement Learning Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T14:10:08.719867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:10:08.719867Z digest=sha256:e2f9636c72252ef3061c5ff70064c19f9f4f92158ed839f33e285b9e55b8332a

Observation 5df212af-6864-4c05-bcde-c381e175389c · inbound

Preference VLM: Leveraging VLMs for Scalable Preference-Based Reinforcement Learning cites this paper.

Preference VLM: Leveraging VLMs for Scalable Preference-Based Reinforcement Learning Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-09T14:52:27.274360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T14:52:27.274360Z digest=sha256:3a799f9e7ac6e8159236af81a8eaa939b149cae0848a2a0961d1ee5d04b525a4

Observation c0df10cc-6fa3-4edf-bb43-6e49a0db3b95 · inbound

Advancing Autonomous VLM Agents via Variational Subgoal-Conditioned Reinforcement Learning cites this paper.

Advancing Autonomous VLM Agents via Variational Subgoal-Conditioned Reinforcement Learning Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T11:25:49.172358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T11:25:49.172358Z digest=sha256:e206836f8aa3d55f165bb8cda4a9e2f2391427380fa97af982858541453b63e5

Observation 6282dfeb-c2a8-4da9-88bf-5d1d1f081959 · inbound

RobotSmith: Generative Robotic Tool Design for Acquisition of Complex Manipulation Skills cites this paper.

RobotSmith: Generative Robotic Tool Design for Acquisition of Complex Manipulation Skills Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T00:13:35.014210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:13:35.014210Z digest=sha256:8153f87e9e0f1cbd1d810ca165723e4dd456793fca4b2d17703a36549764514d

Observation e8732b67-dd5e-4cdb-ba09-daf077b9ee65 · inbound

FOUNDER: Grounding Foundation Models in World Models for Open-Ended Embodied Decision Making cites this paper.

FOUNDER: Grounding Foundation Models in World Models for Open-Ended Embodied Decision Making Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T17:12:29.623006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:12:29.623006Z digest=sha256:191665ce528162f4eff9697f887c35417ad9c25a9d3396333399eefabf3321f2

Observation 7be5cd62-64b7-4087-8c1e-9aa52721930d · inbound

A Survey on Generative Recommendation: Data, Model, and Tasks cites this paper.

A Survey on Generative Recommendation: Data, Model, and Tasks Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning

Reference 152

Resolution
verified exact
arxiv_id, observed 2026-05-18T03:50:52.205396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T03:47:08.208082Z digest=sha256:2300a2afab41602adaacdc2a5483bdaccc7358b67bfdae49d6015948a0edb991

Observation b3117d81-fdb0-4c25-8f81-157f93d8ecd7 · inbound

TOPReward: Token Probabilities as Hidden Zero-Shot Rewards for Robotics cites this paper.

TOPReward: Token Probabilities as Hidden Zero-Shot Rewards for Robotics Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning

Reference 1991

Resolution
unresolved
no resolver link, observed 2026-08-02T21:43:09.722349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:43:09.722349Z digest=sha256:0db81ae2fe00f5524a28f5e3226fbf9b34f023a58dd9104729f0ad17e9951b4f

Observation 326ddb2c-398e-47b5-8d62-9041f74f5f33 · inbound

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning cites this paper.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:fff91ae6d8efbc4c0ee25c2a5892b6eedd99b68cee89ea1f7ebb2267d2000959

Observation a3929100-c867-41f6-b196-3babab1b8bd3 · inbound

AtlasVA: Self-Evolving Visual Skill Memory for Teacher-Free VLM Agents cites this paper.

AtlasVA: Self-Evolving Visual Skill Memory for Teacher-Free VLM Agents Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:28:14.532171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T11:24:48.558423Z digest=sha256:abd92c43176b1f8abe9b859650e09fb5abcc780efc003418a5f4da4ccc2fc0cf

Observation 7525be9d-48ed-4c8c-92e6-7b6dd4f9ed10 · inbound

OrganicHAR: Towards Activity Discovery in Organic Settings for Privacy Preserving Sensors Using Efficient Video Analysis cites this paper.

OrganicHAR: Towards Activity Discovery in Organic Settings for Privacy Preserving Sensors Using Efficient Video Analysis Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-20T08:38:08.968547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T08:37:44.181899Z digest=sha256:a5f5ff616ea585faf4fbc8baf7eff44b104cbef24d645b9cf1f4c6f838ee2ba6

Observation 611fa5dd-27c9-484c-a022-5b5050bec545 · inbound

World Model Self-Distillation: Training World Models to Solve General Tasks cites this paper.

World Model Self-Distillation: Training World Models to Solve General Tasks Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-06-27T10:20:48.821380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-27T10:16:35.511426Z digest=sha256:b1ac145aab297d1f2019216d0b135c5b279c5c15299f23394f07aa01fea81f17

Observation 06b70759-44ff-45b7-9b5d-d60bab24f3e9 · inbound

Improving Robotic Generalist Policies via Flow Reversal Steering cites this paper.

Improving Robotic Generalist Policies via Flow Reversal Steering Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:48:35.734667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-27T06:20:19.209180Z digest=sha256:474c34d9aa11341faec5e0f0e3a16992662e98cb4e4c59e1835ff9b5484f5b5c

Observation 9a3f352e-611b-48fa-a509-8c82461e9ba1 · inbound

Learning Process Rewards via Success Visitation Matching for Efficient RL cites this paper.

Learning Process Rewards via Success Visitation Matching for Efficient RL Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:59:44.662173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T09:20:35.062060Z digest=sha256:e16ad63091ef60c43bd74ceda24204f81e3bfe8bd295e9789a62a31d3f4c398a

Observation ed8f13a1-8017-4500-965e-fe11ff42e7c2 · inbound

Vision-Language Models for Deployable Social Robot Navigation: Bridging Semantic Reasoning and Low-Level Control cites this paper.

Vision-Language Models for Deployable Social Robot Navigation: Bridging Semantic Reasoning and Low-Level Control Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-06-30T09:54:34.569731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T09:52:51.549135Z digest=sha256:528efc562e3bf59506d4e2c129e37296dffdcacabd88818d21aa2dc0e371aaaa

Observation bb0cf977-af77-4644-a776-9e4236e5ddd0 · inbound

LLM-as-a-Verifier: A General-Purpose Verification Framework cites this paper.

LLM-as-a-Verifier: A General-Purpose Verification Framework Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning

Reference 79

Resolution
metadata mismatch
local_arxiv, observed 2026-07-07T12:53:50.203074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-07T12:47:29.552283Z digest=sha256:4bd1a0508ad0ab633593e009e3f3274a01494356cd793b5a70d16954a1ed0f2f

Observation d27b0c92-b711-49ec-a398-c26a00ba82a1 · inbound

LLM-as-a-Verifier: A General-Purpose Verification Framework cites this paper.

LLM-as-a-Verifier: A General-Purpose Verification Framework Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning

Reference 79

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:9a0f9ffd15076cd3609c626f020947b8148ab831e491bb23369ef685d49af923

Observation 73bebea9-1c91-4db5-ae22-76f5ed84a3c7 · inbound

PhyAgentOS: A Self-Evolving Operating System for Embodied Agents with Decoupled Cognitive Planning and Physical Execution cites this paper.

PhyAgentOS: A Self-Evolving Operating System for Embodied Agents with Decoupled Cognitive Planning and Physical Execution Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-01T20:28:20.711404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T20:28:20.711404Z digest=sha256:e57df56f24a64924aaf74b33630adf234fd889c3273c50a7cb301230f45f9cb9

Observation 9eff0a4f-c6d2-4a5d-b544-92bb455fb36f · inbound

MARS-RA: Rank Aggregation for Credit Assignment via Multimodal Comparisons in Embodied Multi-Agent Cooperation cites this paper.

MARS-RA: Rank Aggregation for Credit Assignment via Multimodal Comparisons in Embodied Multi-Agent Cooperation Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning

Reference 67

Resolution
unresolved
no resolver link, observed 2026-07-31T21:55:17.391043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T21:55:17.391043Z digest=sha256:a71dbe21c35c289933f0c448a225ae6e71079d529bff7695680cb1f2ea0ea900

Observation ad616c7d-68a3-4d77-9ff1-1bf7b6e5b624 · inbound

Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills cites this paper.

Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning

Reference 216

Resolution
unresolved
no resolver link, observed 2026-08-04T19:45:35.172574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:45:35.172574Z digest=sha256:517f05766831d519fda0546386a39503ba18269ac1d841b19e9d27360219ee63