Pith. sign in

Paper Citation Record · LEDGER

Training-free Generation of Temporally Consistent Rewards from VLMs

As of 8 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 0 inbound Pith citation observations for arXiv:2507.04789.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.04789 v1

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:47:44.014826Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

55 of 55 outbound references displayed

  • verified exact0
  • verified fuzzy33
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f134ada1-780f-48a4-be4d-1f63cc164808 · outbound

This paper cites Do As I Can, Not As I Say: Grounding Language in Robotic Affordances.

Training-free Generation of Temporally Consistent Rewards from VLMs Do As I Can, Not As I Say: Grounding Language in Robotic Affordances

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:40.030491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:47:40.030491Z digest=sha256:94b7a28055f9d2e44933431dc091d8d4d08243b2c5c45e6caf01471ff15af98b

Observation a287c5c4-e6d8-4dfd-8dfd-0407e8d7bb64 · outbound

This paper cites Vision-language models as a source of rewards.

Training-free Generation of Temporally Consistent Rewards from VLMs Vision-language models as a source of rewards

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:51.972386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:47:40.080984Z digest=sha256:1be080add229affb537c4841bde4ac7db5e958b09e809f91eb0aac1f6beed2d2

Observation 4b02908f-fac7-4629-99ba-9633105e81b9 · outbound

This paper cites Spatialvlm: Endow- ing vision-language models with spatial reasoning capabili- ties.

Training-free Generation of Temporally Consistent Rewards from VLMs Spatialvlm: Endow- ing vision-language models with spatial reasoning capabili- ties

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:40.142313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:47:40.142313Z digest=sha256:a167df8d7fecf3a459e0aa4c5496767d020377fb43a1b67e63a7f2b694b90775

Observation f1d4c3a6-72d9-47a9-be9a-8ed80b7cc642 · outbound

This paper cites Towards a unified agent with foundation models.

Training-free Generation of Temporally Consistent Rewards from VLMs Towards a unified agent with foundation models

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:51.809780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:47:40.179746Z digest=sha256:1190fc7e5dd429fb28c3f01bc398273c8fcf0d800aff9aee82180e2989a5b111

Observation 23a79ce5-2d5d-46e7-8591-be13e497db7e · outbound

This paper cites Video language planning.

Training-free Generation of Temporally Consistent Rewards from VLMs Video language planning

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:51.651023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:47:40.263441Z digest=sha256:8dd929c357c5aa6b266bb48e105e10f8e14528739f76f14e87ae19a2639c0f19

Observation a4937b0b-b98f-47d0-b8da-e5b9fa2e992d · outbound

This paper cites Manipulate- anything: Automating real-world robots using vision- language models.

Training-free Generation of Temporally Consistent Rewards from VLMs Manipulate- anything: Automating real-world robots using vision- language models

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:51.455521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:47:40.331962Z digest=sha256:5bbe2c7997ddfdaabef912b26e29daa1ebed39596a8f33224210dd1e539c3d13

Observation f0f9be62-7248-4a83-b51e-bff7ba31dddf · outbound

This paper cites Phys- ically grounded vision-language models for robotic manip- ulation.

Training-free Generation of Temporally Consistent Rewards from VLMs Phys- ically grounded vision-language models for robotic manip- ulation

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:51.267732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:47:40.421741Z digest=sha256:68e9f6cb92c042a59c0b4b48af80b9397d8861271b3e677a7a6f75ee16ff707d

Observation 46fd493f-6084-4191-90ea-d1646bb4cc50 · outbound

This paper cites Doremi: Grounding language model by detecting and recov- ering from plan-execution misalignment.

Training-free Generation of Temporally Consistent Rewards from VLMs Doremi: Grounding language model by detecting and recov- ering from plan-execution misalignment

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:51.078631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:47:40.507364Z digest=sha256:05ba2da369029401b04ec9bf0aae22749fcb9920b10754ab9968f0e5c9cffa80

Observation 909f63de-b7b0-47fa-86d9-2c31311c91f9 · outbound

This paper cites Mixgen: A new multi- modal data augmentation.

Training-free Generation of Temporally Consistent Rewards from VLMs Mixgen: A new multi- modal data augmentation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:40.622191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:47:40.622191Z digest=sha256:3d09fafabe6e270227a74cd05abb2064b65b766ae633c09b06cbad2824bfbcac

Observation a69a73d2-41fb-4632-add1-e48336247fac · outbound

This paper cites V oxposer: Composable 3d value maps for robotic manipulation with language models.

Training-free Generation of Temporally Consistent Rewards from VLMs V oxposer: Composable 3d value maps for robotic manipulation with language models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:50.920534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:47:40.668131Z digest=sha256:bbcd7875ae016a34224675014162142c0ddbe3ccca3f25b4103bf9ddac22f6db

Observation a2f238db-8987-4812-911c-21b459dfa134 · outbound

This paper cites Robobrain: A unified brain model for robotic manipulation from abstract to concrete.

Training-free Generation of Temporally Consistent Rewards from VLMs Robobrain: A unified brain model for robotic manipulation from abstract to concrete

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:50.753133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:47:40.739804Z digest=sha256:9b02366c8ffa65e449bb31d989a30e9d0d5b60dba326dc5aa86f553f9fcb47e9

Observation 1736832a-c8e6-4a07-bbb1-29512151ca3f · outbound

This paper cites Scaling up visual and vision-language representa- tion learning with noisy text supervision.

Training-free Generation of Temporally Consistent Rewards from VLMs Scaling up visual and vision-language representa- tion learning with noisy text supervision

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:50.534051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:47:40.807513Z digest=sha256:1b33ab45975da565d5d4723b44e1fd1205d3fdfb1806d673d31f22b07e0a380b

Observation 6bbb4804-88f2-4259-bfe3-7ac2b93ef29f · outbound

This paper cites De- composed prompting: A modular approach for solving com- plex tasks.

Training-free Generation of Temporally Consistent Rewards from VLMs De- composed prompting: A modular approach for solving com- plex tasks

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:50.298515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:47:40.867070Z digest=sha256:3180a2f5b3bca4f0bae4d8e6a5bbf0c34c5ef05a6b18f6bbab54289c5f6c5103

Observation 768e7873-9fd7-4bfe-9814-3b021123782b · outbound

This paper cites Segment any- thing.

Training-free Generation of Temporally Consistent Rewards from VLMs Segment any- thing

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:40.928757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:47:40.928757Z digest=sha256:61d25083fb96ece1be962dc8bafe5985a61673714cf17261688073995268e5b3

Observation 3360e7b3-df2e-4695-8faf-1822753c8fbe · outbound

This paper cites Multimodal sensor fusion with differentiable filters.

Training-free Generation of Temporally Consistent Rewards from VLMs Multimodal sensor fusion with differentiable filters

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:50.117560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:47:40.986288Z digest=sha256:de7021dc216b44c55eb5db9920601398f45853762caf108165ec43364784a84b

Observation 74fe2b5e-accf-4561-8206-e9516c154ca0 · outbound

This paper cites What foundation models can bring for robot learning in manipulation: A survey.

Training-free Generation of Temporally Consistent Rewards from VLMs What foundation models can bring for robot learning in manipulation: A survey

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:41.045499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:47:41.045499Z digest=sha256:9d2332a7520e4a01bdd88e3fa328d3e09686b4e2a0d485af37867943eca00199

Observation f6987d0b-d893-4fc4-86ff-5bfd90c13d27 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Training-free Generation of Temporally Consistent Rewards from VLMs Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:49.856803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:47:41.100895Z digest=sha256:ee8ec1b0533b7f8e5bbf747e8de969b2d0f3378148772e4d73936efb08733604

Observation 469d33dc-beb8-457e-b52c-98824cd12f77 · outbound

This paper cites Code as policies: Language model programs for embodied con- trol.

Training-free Generation of Temporally Consistent Rewards from VLMs Code as policies: Language model programs for embodied con- trol

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:49.716556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:47:41.135694Z digest=sha256:e05141ab51ddc23e2b9c6c93646331da1490d3f0866d905c1fb915e93a0f188a

Observation be3dc320-6d76-4df1-8112-9b1642ac2f8d · outbound

This paper cites Reflect: Summa- rizing robot experiences for failure explanation and correc- tion.

Training-free Generation of Temporally Consistent Rewards from VLMs Reflect: Summa- rizing robot experiences for failure explanation and correc- tion

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:49.579110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:47:41.187899Z digest=sha256:750de207745e89f5ae4efa25f5ca9725789bf7eef7ddd7ddcc9178941eb31bd8

Observation 41423878-e643-4d71-a570-f1945e5c1435 · outbound

This paper cites ExploRLLM: Guiding Exploration in Reinforcement Learning with Large Language Models.

Training-free Generation of Temporally Consistent Rewards from VLMs ExploRLLM: Guiding Exploration in Reinforcement Learning with Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:41.236915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:47:41.236915Z digest=sha256:1f4ad05c75f58a360df2037681c532b70c0d208bb8aa5be0220e8983f5a61b62

Observation a3bcff94-6a78-43d9-9dba-0aea345740ba · outbound

This paper cites Calvin: A benchmark for language- conditioned policy learning for long-horizon robot manip- ulation tasks.

Training-free Generation of Temporally Consistent Rewards from VLMs Calvin: A benchmark for language- conditioned policy learning for long-horizon robot manip- ulation tasks

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:49.277646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:47:41.292412Z digest=sha256:98715d11acbab40d0891677fb8e3745df48464445e81af5934475dba5008acfc

Observation baaa893d-4bfb-4997-956a-a2b1f8e47150 · outbound

This paper cites The deep latent space particle filter for real-time data assimilation with uncertainty quantification.

Training-free Generation of Temporally Consistent Rewards from VLMs The deep latent space particle filter for real-time data assimilation with uncertainty quantification

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:49.089889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:47:41.337357Z digest=sha256:1e2345b9ba65c63c185d1ea07947105420d023a390d128347fa3aaaec80d0f7a

Observation bb729636-72e8-4237-8032-cbae885bfbfc · outbound

This paper cites an unresolved cited work.

Training-free Generation of Temporally Consistent Rewards from VLMs Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:47:48.903738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:47:41.384122Z digest=sha256:183678895f42cba98a1264c5d9d011cedec2e5fa562eb3aa21ab3ba0e824ec42

Observation ef9ca880-c2e3-4d85-9b2c-7aed41380125 · outbound

This paper cites A real-to-sim-to-real approach to robotic manip- ulation with VLM-generated iterative keypoint rewards.

Training-free Generation of Temporally Consistent Rewards from VLMs A real-to-sim-to-real approach to robotic manip- ulation with VLM-generated iterative keypoint rewards

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:48.721460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:47:41.447612Z digest=sha256:2e21bb359bab79fff4d5b6a5f71a8245ad6d1d159f2b4065582f3ef292724cd1

Observation 99187c9c-c470-45b2-b0aa-b6df514ad934 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

Training-free Generation of Temporally Consistent Rewards from VLMs Learn- ing transferable visual models from natural language super- vision

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:48.543814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:47:41.501344Z digest=sha256:2595cd572bc76b11b31c0ef4d2e0d5ead3e02d045b2089bd6b424fc97cf485b6

Observation f1662e6a-7c88-47f7-b24c-85f183ef3f94 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

Training-free Generation of Temporally Consistent Rewards from VLMs SAM 2: Segment Anything in Images and Videos

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:41.565323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:47:41.565323Z digest=sha256:c23b8bc5b6c50c11ca209e826efe74f3e2b4ea31bf24d62bc683642c3db6e7ac

Observation 23706177-03b2-4fbf-8698-4cba3fcd5a9d · outbound

This paper cites Vision-language models are zero- shot reward models for reinforcement learning.

Training-free Generation of Temporally Consistent Rewards from VLMs Vision-language models are zero- shot reward models for reinforcement learning

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:48.405348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:47:41.604704Z digest=sha256:257d8ef93d296bc2b30cf2859f1f529848e5dfd0282c380c5160302e7e223c44

Observation b25714cc-0c6e-47d5-8b07-a3043dcb36dd · outbound

This paper cites Chinchali.

Training-free Generation of Temporally Consistent Rewards from VLMs Chinchali

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:48.164013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:47:41.664896Z digest=sha256:3fb3d856dbced6ed6ddd260b97a20bd30e3133dab92b8e7e66aac07fa862029a

Observation 10d0986b-78cb-4f54-971c-ac9ea122b65e · outbound

This paper cites Cliport: What and where pathways for robotic manipulation.

Training-free Generation of Temporally Consistent Rewards from VLMs Cliport: What and where pathways for robotic manipulation

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:47.942743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:47:41.728121Z digest=sha256:cd1ba5af75cfa0a613e1da25a64ea997b22ea39271134c423c8bd9adc02bbe20

Observation b1f21a69-874f-41b1-b189-c8ff53451e41 · outbound

This paper cites Perceiver- actor: A multi-task transformer for robotic manipulation.

Training-free Generation of Temporally Consistent Rewards from VLMs Perceiver- actor: A multi-task transformer for robotic manipulation

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:47.735982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:47:41.817923Z digest=sha256:847d9a2149090bc0c80e61d67a4953b41c55409cd4f13f8e3ae32c0aa3c0b423

Observation b4c1f25f-8f94-4208-8291-1bdc5dcf0673 · outbound

This paper cites Progprompt: Generating situ- ated robot task plans using large language models.

Training-free Generation of Temporally Consistent Rewards from VLMs Progprompt: Generating situ- ated robot task plans using large language models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:47.543615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:47:41.891782Z digest=sha256:3c028694c9a1a6aef6c7ed1eba20fa835eb2badf6ebb6daa3241041ea9e41b23

Observation c54b97b5-9066-497b-999e-157c0dd7a806 · outbound

This paper cites Reason-rft: Reinforcement fine-tuning for visual reasoning.

Training-free Generation of Temporally Consistent Rewards from VLMs Reason-rft: Reinforcement fine-tuning for visual reasoning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:41.937872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:47:41.937872Z digest=sha256:1f6e1a3fe9fbeae7484f997efdf7058b12e039d2213483013c84cf9d85891d64

Observation 38d4dd40-061d-4311-b806-f3d669eca342 · outbound

This paper cites Cotdet: Affordance knowledge prompting for task driven object de- tection.

Training-free Generation of Temporally Consistent Rewards from VLMs Cotdet: Affordance knowledge prompting for task driven object de- tection

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:47.390815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:47:41.990751Z digest=sha256:198bc88d2cd32a1b183210695a4ccd8d9c4c951adda7c369912508085fadf367

Observation 97be814c-a937-4040-b3cd-c67d0f46dc28 · outbound

This paper cites AffordGrasp: In-Context Affordance Reasoning for Open-Vocabulary Task-Oriented Grasping in Clutter.

Training-free Generation of Temporally Consistent Rewards from VLMs AffordGrasp: In-Context Affordance Reasoning for Open-Vocabulary Task-Oriented Grasping in Clutter

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:42.042018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:47:42.042018Z digest=sha256:031c56ac469008a9d9f68e226a6e08d2343f63c00995264efe48829516316fac

Observation 54bc3c62-7248-4cab-8587-cf1fc9edbe6f · outbound

This paper cites Real-World Offline Reinforcement Learning from Vision Language Model Feedback.

Training-free Generation of Temporally Consistent Rewards from VLMs Real-World Offline Reinforcement Learning from Vision Language Model Feedback

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:42.135720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:47:42.135720Z digest=sha256:628d7033cb8a19585dfc9fa1f233850581155b4690cac2a59f293bd79c9ef4da

Observation 27fd4f0c-f04d-427c-98df-383ab589ccce · outbound

This paper cites Code as reward: Empowering reinforcement learning with vlms.

Training-free Generation of Temporally Consistent Rewards from VLMs Code as reward: Empowering reinforcement learning with vlms

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:47.174138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:47:42.250922Z digest=sha256:97513ca14ff5cbf1331fbf9ec820e059ee75f42bf179909b7dd7ac2ef35eefb9

Observation bf4a0aaf-3f49-4f04-8f85-551259fb6f6f · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Training-free Generation of Temporally Consistent Rewards from VLMs Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:42.366889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:47:42.366889Z digest=sha256:83c0c292f629aad574b1dfb17a94888b265a80538493f5c4709984ea809e7fee

Observation 20d27e66-116d-4cbc-85aa-aea9cdc545fd · outbound

This paper cites Rl-vlm-f: Rein- forcement learning from vision language foundation model feedback.

Training-free Generation of Temporally Consistent Rewards from VLMs Rl-vlm-f: Rein- forcement learning from vision language foundation model feedback

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:46.922946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:47:42.528976Z digest=sha256:af6f0d12ca33234c1c599ca540abab8702fd55490d175fcf509a1867a887b490

Observation fb08866f-8e6e-4dcc-a15c-23a99cda2ebb · outbound

This paper cites Any-point Trajectory Modeling for Policy Learning.

Training-free Generation of Temporally Consistent Rewards from VLMs Any-point Trajectory Modeling for Policy Learning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:42.614493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:47:42.614493Z digest=sha256:5dde9c09d580a91b411924492f3383ca9b1f1f386f56a0e8b8d2232a7d60de23

Observation 7dce3e65-a42f-40f7-8ed2-0869618216e2 · outbound

This paper cites RoboMIND: Benchmark on Multi-embodiment Intelligence Normative Data for Robot Manipulation.

Training-free Generation of Temporally Consistent Rewards from VLMs RoboMIND: Benchmark on Multi-embodiment Intelligence Normative Data for Robot Manipulation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:42.681720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:47:42.681720Z digest=sha256:c5fe4465d0e900dc0c3a81120805d4aa7b264a41aa7c3ec4e8f431fb12360549

Observation 63959c63-c49c-4b44-90e8-5d05e92514a1 · outbound

This paper cites Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V.

Training-free Generation of Temporally Consistent Rewards from VLMs Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:42.750529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:47:42.750529Z digest=sha256:1f002af5a0f36c063be0be7af4436516b6abdc807641cff61c0b1e5c5dce722a

Observation 7d982682-9648-4ece-baef-88452c82b263 · outbound

This paper cites Robot fine- tuning made easy: Pre-training rewards and policies for au- tonomous real-world reinforcement learning.

Training-free Generation of Temporally Consistent Rewards from VLMs Robot fine- tuning made easy: Pre-training rewards and policies for au- tonomous real-world reinforcement learning

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:46.783281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:47:42.851403Z digest=sha256:a81bab4cace42fb4f0c5bfaeb904bf6249465d036ee0f16a95fa8d2c81ac37b8

Observation e2bb6c5d-f225-4e7b-806a-51239a3d6acf · outbound

This paper cites Par- ticle filters in latent space for robust deformable linear ob- ject tracking.

Training-free Generation of Temporally Consistent Rewards from VLMs Par- ticle filters in latent space for robust deformable linear ob- ject tracking

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:46.603324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:47:42.923587Z digest=sha256:c2738dc7f93e4769dfbc9158b6b610db23f417dd02a0f7a297de356aa0113a98

Observation 4f28e9b1-6896-45fa-8085-92ec2fb70bd8 · outbound

This paper cites Sornet: Spatial object-centric representations for se- quential manipulation.

Training-free Generation of Temporally Consistent Rewards from VLMs Sornet: Spatial object-centric representations for se- quential manipulation

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:46.378775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:47:42.987023Z digest=sha256:c71939aac207236528bb89f2a053a214e209de5ae2815001bcd3584cfd28c8ea

Observation 218006e5-280a-4e00-837c-38e8e384aa60 · outbound

This paper cites Robopoint: A vision-language model for spatial affordance prediction in robotics.

Training-free Generation of Temporally Consistent Rewards from VLMs Robopoint: A vision-language model for spatial affordance prediction in robotics

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:46.158520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:47:43.078554Z digest=sha256:5276a279ca438d3e9d7d28a07245a2f7be7ba8fbc34634d69c4e00ef2f1e2a5b

Observation c571f2e5-095b-4f16-bfaf-fea0bb00de83 · outbound

This paper cites Sam-e: Leveraging visual foundation model with sequence imitation for embodied ma- nipulation.

Training-free Generation of Temporally Consistent Rewards from VLMs Sam-e: Leveraging visual foundation model with sequence imitation for embodied ma- nipulation

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:46.014193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:47:43.151441Z digest=sha256:987ab4126a3a7013901114c95e92dd066be106909354ada94d9c567dc8b64276

Observation 429af5d3-1448-4a2c-a090-499261c9dcb6 · outbound

This paper cites MapNav: A Novel Memory Representation via Annotated Semantic Maps for Vision-and-Language Navigation.

Training-free Generation of Temporally Consistent Rewards from VLMs MapNav: A Novel Memory Representation via Annotated Semantic Maps for Vision-and-Language Navigation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:43.218177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:47:43.218177Z digest=sha256:8ea7ff271d9c97d134d7bf3c9a1ffd5adac9286afa2e0a7925dad2610d588e3e

Observation e2bde3eb-7b89-43d8-ac34-52fb8082f21b · outbound

This paper cites TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies.

Training-free Generation of Temporally Consistent Rewards from VLMs TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:43.290054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:47:43.290054Z digest=sha256:357495ac5b0a19e2605e004ee13a5d1be973530c392ff9951f3a8841e66c5dd3

Observation e9e5d94c-0ddc-42ff-9e4a-a575fa5fcfce · outbound

This paper cites an unresolved cited work.

Training-free Generation of Temporally Consistent Rewards from VLMs Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:47:45.846680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:47:43.353180Z digest=sha256:df7bb9b0dfb1f70ba5b3ff6bce84bba4051260b32e0db108bda214dfb30a4928

Observation 6a8e0e12-be39-4bf6-86ce-c1690d17dbb5 · outbound

This paper cites an unresolved cited work.

Training-free Generation of Temporally Consistent Rewards from VLMs Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:47:45.286813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:47:43.546848Z digest=sha256:0cc4601a5093e0ab91bd9488632f54fd9b786058e5aa80664ada11f1503bf599

Observation 95211ab0-e2b3-4a49-8da8-7f736371a782 · outbound

This paper cites Figure 11.

Training-free Generation of Temporally Consistent Rewards from VLMs Figure 11

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:45.081039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:47:43.639443Z digest=sha256:1e20ea289ad8862c822f71ff761fb25f6d1ab29ce0291223f6d2e204f770d9d2

Observation a9656fe6-cc2f-4b2f-bdfb-2fb0cc0ead1c · outbound

This paper cites an unresolved cited work.

Training-free Generation of Temporally Consistent Rewards from VLMs Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:47:44.831500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:47:43.723254Z digest=sha256:8dd2a9701a4d36d4e706812e7de6f783c11368b59202a6d78538ee0616958033

Observation d2b99751-8743-4fa9-a03e-4a64aaf32e54 · outbound

This paper cites an unresolved cited work.

Training-free Generation of Temporally Consistent Rewards from VLMs Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:47:45.446685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:47:43.790046Z digest=sha256:abc3572f06054a6cd3c96dc9b5de387628aea9247d2cb8eb754a98a3520fafc1

Observation db95df95-ace1-4a26-899d-51c16c8dae9c · outbound

This paper cites an unresolved cited work.

Training-free Generation of Temporally Consistent Rewards from VLMs Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:47:45.620439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:47:43.909897Z digest=sha256:6e0cd0d5c27a157111d26da9b36e9a26f760c845b4adcab89d9e8cf18023cbbd

Observation 4fe6474c-0e82-4c38-a003-dcb777e0049d · outbound

This paper cites Completion Status Identification of Sub-goals: System prompt: Detailed in Fig.

Training-free Generation of Temporally Consistent Rewards from VLMs Completion Status Identification of Sub-goals: System prompt: Detailed in Fig

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:44.614948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:47:44.014826Z digest=sha256:5772977f5c14d65f6de859a5c99430160207e9a700566fdf39781dd22b30716d

Pith citing papers

No inbound Pith citation observations are available.