Pith. sign in

Paper Citation Record · LEDGER

Vision-Language Models as a Source of Rewards

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2312.09187.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.09187 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:40:11.923452Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T09:59:44.592995Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 6f759d14-e0f2-4a46-984d-4cf3474c1a56 · inbound

LLM-Based Offline Learning for Embodied Agents via Consistency-Guided Reward Ensemble cites this paper.

LLM-Based Offline Learning for Embodied Agents via Consistency-Guided Reward Ensemble Vision-Language Models as a Source of Rewards

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T12:33:07.000394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:33:07.000394Z digest=sha256:721dda1fe9836e61602f6a48e55d42ee428beb5a89180770f7a9dd1377c96eed

Observation f14040fc-c140-4551-af25-53001d9673c0 · inbound

STEVE-Audio: Expanding the Goal Conditioning Modalities of Embodied Agents in Minecraft cites this paper.

STEVE-Audio: Expanding the Goal Conditioning Modalities of Embodied Agents in Minecraft Vision-Language Models as a Source of Rewards

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T04:54:54.711221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:54:54.711221Z digest=sha256:709973a093732dce1f84c36a5a606b77ebde44cda55c7794cc183dab8f9f3a4f

Observation bcb625d3-95da-4688-af60-76847536ef5c · inbound

YingSound: Video-Guided Sound Effects Generation with Multi-modal Chain-of-Thought Controls cites this paper.

YingSound: Video-Guided Sound Effects Generation with Multi-modal Chain-of-Thought Controls Vision-Language Models as a Source of Rewards

Reference 102

Resolution
unresolved
no resolver link, observed 2026-08-11T17:17:49.504330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:17:49.504330Z digest=sha256:5cc5c26513764526b0f0cf8ca36c6d1e8b73c5e4231b55ca2b98894c1c760c11

Observation b83bc419-0129-477e-90b2-e188a5197b0d · inbound

RobotSmith: Generative Robotic Tool Design for Acquisition of Complex Manipulation Skills cites this paper.

RobotSmith: Generative Robotic Tool Design for Acquisition of Complex Manipulation Skills Vision-Language Models as a Source of Rewards

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T00:13:28.793016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:13:28.793016Z digest=sha256:fc3b94b1f59015190251591b795818030a7ab59978bc6bfe9be8c263377feb8a

Observation 393ff599-eb2f-415e-9f3d-295793eb7661 · inbound

FOUNDER: Grounding Foundation Models in World Models for Open-Ended Embodied Decision Making cites this paper.

FOUNDER: Grounding Foundation Models in World Models for Open-Ended Embodied Decision Making Vision-Language Models as a Source of Rewards

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T17:12:27.133061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:12:27.133061Z digest=sha256:5f85f5cdab4d481d052dda9e9d16edf339d4b7f52e78137bb3ab221d6da00117

Observation 9ff64ef4-43cb-4187-8a91-b9730e3e8d1b · inbound

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning cites this paper.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning Vision-Language Models as a Source of Rewards

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:dd591be5b7464ad0b72b78458660a257454baa54f79ba8762db1120571bf4d64

Observation 8f96b5e0-a3ca-448d-a974-045bb3a59bb0 · inbound

Learning Process Rewards via Success Visitation Matching for Efficient RL cites this paper.

Learning Process Rewards via Success Visitation Matching for Efficient RL Vision-Language Models as a Source of Rewards

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:59:44.594485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T09:20:35.062060Z digest=sha256:6c87e1cdca7e560714267240ee1b46440140f246561b74ec06bb9d43792f80f6

Observation 5e767bed-48c0-49f0-b88f-1a9050aebd65 · inbound

QVal: Cheaply Evaluating Dense Supervision Signals for Long-Horizon LLM Agents cites this paper.

QVal: Cheaply Evaluating Dense Supervision Signals for Long-Horizon LLM Agents Vision-Language Models as a Source of Rewards

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-07-01T09:55:41.445185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-07-01T05:59:13.631078Z digest=sha256:d3a6b759971eb94dfed9970cdc69b101fe1f39ee2a393bcc90106aaa918b1625

Observation f7291943-708c-4944-9fe9-03b5eeb91fad · inbound

Progress Reward Modeling for Robotic Learning: A Comprehensive Survey cites this paper.

Progress Reward Modeling for Robotic Learning: A Comprehensive Survey Vision-Language Models as a Source of Rewards

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T09:42:55.547292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:42:55.547292Z digest=sha256:fe08fcab65aeaadca2f115c6e4b0e6fca87fdff2e47286401a6c5decf2bffd5d

Observation fccf0884-caea-4685-8570-debabaf3e1fa · inbound

A Unified Framework for Dynamic Reward Shaping in Reinforcement Learning cites this paper.

A Unified Framework for Dynamic Reward Shaping in Reinforcement Learning Vision-Language Models as a Source of Rewards

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-12T00:30:33.353399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:30:33.353399Z digest=sha256:47e3e4e51ecab5cbceac294d28668563c800a4c8bb8ba9b2f36f1d58c11e6ae3

Observation 0884635c-e168-4f7e-8c97-f4efbcac098c · inbound

TrustRoboReward: Preference-Ordered Isotonic Score Editing for Multi-Paradigm Robot Reward Models cites this paper.

TrustRoboReward: Preference-Ordered Isotonic Score Editing for Multi-Paradigm Robot Reward Models Vision-Language Models as a Source of Rewards

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-14T04:40:11.923452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:40:11.923452Z digest=sha256:c0d414a035c99dd500764e49aac6c2443b3e0b60291b8d03e191cc56b7077895