Pith. sign in

Paper Citation Record · LEDGER

Training-free Generation of Temporally Consistent Rewards from VLMs

As of 21 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 0 inbound Pith citation observations for arXiv:2507.04789.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.04789 v1

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:47:44.014826Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

55 of 55 outbound references displayed

  • verified exact0
  • verified fuzzy33
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f134ada1-780f-48a4-be4d-1f63cc164808 · outbound

This paper cites Do As I Can, Not As I Say: Grounding Language in Robotic Affordances.

Training-free Generation of Temporally Consistent Rewards from VLMs Do As I Can, Not As I Say: Grounding Language in Robotic Affordances

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:40.030491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:47:40.030491Z digest=sha256:a3adda1eaa65e41321bdf9cb973922be724f43acfedfe3662595a970ca5781a1

Observation a287c5c4-e6d8-4dfd-8dfd-0407e8d7bb64 · outbound

This paper cites Vision-language models as a source of rewards.

Training-free Generation of Temporally Consistent Rewards from VLMs Vision-language models as a source of rewards

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:51.972386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:47:40.080984Z digest=sha256:138c894793370b9a5dc0a87f93ccb7047ae6f3734136ccd7306155d125017bcd

Observation 4b02908f-fac7-4629-99ba-9633105e81b9 · outbound

This paper cites Spatialvlm: Endow- ing vision-language models with spatial reasoning capabili- ties.

Training-free Generation of Temporally Consistent Rewards from VLMs Spatialvlm: Endow- ing vision-language models with spatial reasoning capabili- ties

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:40.142313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:47:40.142313Z digest=sha256:208e52937f4d6fb1fefbf47709f5748dfbd31b0dc0b6874dd7eb82da05a511be

Observation f1d4c3a6-72d9-47a9-be9a-8ed80b7cc642 · outbound

This paper cites Towards a unified agent with foundation models.

Training-free Generation of Temporally Consistent Rewards from VLMs Towards a unified agent with foundation models

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:51.809780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:47:40.179746Z digest=sha256:3ef00292777ba7fb6085df0641a6b5ce457d1cb10dfbd5af9e27190df537ec59

Observation 23a79ce5-2d5d-46e7-8591-be13e497db7e · outbound

This paper cites Video language planning.

Training-free Generation of Temporally Consistent Rewards from VLMs Video language planning

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:51.651023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:47:40.263441Z digest=sha256:0820ec4304cb9c88666482c6e694f8f785c207f1fbde8f075ebfc98ac99a48a2

Observation a4937b0b-b98f-47d0-b8da-e5b9fa2e992d · outbound

This paper cites Manipulate- anything: Automating real-world robots using vision- language models.

Training-free Generation of Temporally Consistent Rewards from VLMs Manipulate- anything: Automating real-world robots using vision- language models

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:51.455521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:47:40.331962Z digest=sha256:892d36523d5013f834ae5f8ce9a3847adfe57cdfea6d3c899a2901ff4265a254

Observation f0f9be62-7248-4a83-b51e-bff7ba31dddf · outbound

This paper cites Phys- ically grounded vision-language models for robotic manip- ulation.

Training-free Generation of Temporally Consistent Rewards from VLMs Phys- ically grounded vision-language models for robotic manip- ulation

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:51.267732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:47:40.421741Z digest=sha256:22bccdd4c078b12b4598df6cce5f156cec0b9cedeb42ceb5ab30999f771bfe9d

Observation 46fd493f-6084-4191-90ea-d1646bb4cc50 · outbound

This paper cites Doremi: Grounding language model by detecting and recov- ering from plan-execution misalignment.

Training-free Generation of Temporally Consistent Rewards from VLMs Doremi: Grounding language model by detecting and recov- ering from plan-execution misalignment

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:51.078631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:47:40.507364Z digest=sha256:16211998352adfb60496f9344a023267d004c670856203fd84b3caad15888d10

Observation 909f63de-b7b0-47fa-86d9-2c31311c91f9 · outbound

This paper cites Mixgen: A new multi- modal data augmentation.

Training-free Generation of Temporally Consistent Rewards from VLMs Mixgen: A new multi- modal data augmentation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:40.622191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:47:40.622191Z digest=sha256:cb024af8831247e0c2a3a10d33a7b6a6cd995984ff6c77659c90f1f597746d62

Observation a69a73d2-41fb-4632-add1-e48336247fac · outbound

This paper cites V oxposer: Composable 3d value maps for robotic manipulation with language models.

Training-free Generation of Temporally Consistent Rewards from VLMs V oxposer: Composable 3d value maps for robotic manipulation with language models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:50.920534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:47:40.668131Z digest=sha256:db590d78aa26440bba7840c3181d3be34b878354a10c48efee5fdab28b74abd2

Observation a2f238db-8987-4812-911c-21b459dfa134 · outbound

This paper cites Robobrain: A unified brain model for robotic manipulation from abstract to concrete.

Training-free Generation of Temporally Consistent Rewards from VLMs Robobrain: A unified brain model for robotic manipulation from abstract to concrete

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:50.753133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:47:40.739804Z digest=sha256:9b1d935d788c38ca888eb54de1aa5ac9c97b1df807d8318a6199038f0992494c

Observation 1736832a-c8e6-4a07-bbb1-29512151ca3f · outbound

This paper cites Scaling up visual and vision-language representa- tion learning with noisy text supervision.

Training-free Generation of Temporally Consistent Rewards from VLMs Scaling up visual and vision-language representa- tion learning with noisy text supervision

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:50.534051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:47:40.807513Z digest=sha256:a2c3ff9e34b8730728358b2cf1b091038a144a1d24e2bad5ae466a14689a8656

Observation 6bbb4804-88f2-4259-bfe3-7ac2b93ef29f · outbound

This paper cites De- composed prompting: A modular approach for solving com- plex tasks.

Training-free Generation of Temporally Consistent Rewards from VLMs De- composed prompting: A modular approach for solving com- plex tasks

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:50.298515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:47:40.867070Z digest=sha256:f765c86f83b4e13d489c52c3f9ae0d72a983ffe9756b00b803290adf3cd38450

Observation 768e7873-9fd7-4bfe-9814-3b021123782b · outbound

This paper cites Segment any- thing.

Training-free Generation of Temporally Consistent Rewards from VLMs Segment any- thing

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:40.928757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:47:40.928757Z digest=sha256:e114f8c74257c54eb20a18eb86c7802cf91fcf5ae39a89c3e0699fa149302a9e

Observation 3360e7b3-df2e-4695-8faf-1822753c8fbe · outbound

This paper cites Multimodal sensor fusion with differentiable filters.

Training-free Generation of Temporally Consistent Rewards from VLMs Multimodal sensor fusion with differentiable filters

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:50.117560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:47:40.986288Z digest=sha256:ec21d776625472f3ec9f1258ffcd6e089b5df7d56744c0f532e729de4db986c5

Observation 74fe2b5e-accf-4561-8206-e9516c154ca0 · outbound

This paper cites What foundation models can bring for robot learning in manipulation: A survey.

Training-free Generation of Temporally Consistent Rewards from VLMs What foundation models can bring for robot learning in manipulation: A survey

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:41.045499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:47:41.045499Z digest=sha256:1d0193a84d4792c9ecf9389fc83466347e4705d0448d2edb5aed7624fc1c0de7

Observation f6987d0b-d893-4fc4-86ff-5bfd90c13d27 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Training-free Generation of Temporally Consistent Rewards from VLMs Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:49.856803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:47:41.100895Z digest=sha256:bb89b5602b7a7e256b0da8f7b9a964928ad67c1d60bf9b0002edbc28c63be2f2

Observation 469d33dc-beb8-457e-b52c-98824cd12f77 · outbound

This paper cites Code as policies: Language model programs for embodied con- trol.

Training-free Generation of Temporally Consistent Rewards from VLMs Code as policies: Language model programs for embodied con- trol

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:49.716556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:47:41.135694Z digest=sha256:f0c6518363a364a591c8d5c376d6293f849f26ba1e48c4235253df320de2244c

Observation be3dc320-6d76-4df1-8112-9b1642ac2f8d · outbound

This paper cites Reflect: Summa- rizing robot experiences for failure explanation and correc- tion.

Training-free Generation of Temporally Consistent Rewards from VLMs Reflect: Summa- rizing robot experiences for failure explanation and correc- tion

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:49.579110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:47:41.187899Z digest=sha256:9b19f855e4c5cc36a6312f475ac2c16d5572daf6203ac86958981c22ee3532c9

Observation 41423878-e643-4d71-a570-f1945e5c1435 · outbound

This paper cites ExploRLLM: Guiding Exploration in Reinforcement Learning with Large Language Models.

Training-free Generation of Temporally Consistent Rewards from VLMs ExploRLLM: Guiding Exploration in Reinforcement Learning with Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:41.236915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:47:41.236915Z digest=sha256:fbe302230889a452363a94b53cf36a1540e8a9c24bd76dba09854b9e7d22a407

Observation a3bcff94-6a78-43d9-9dba-0aea345740ba · outbound

This paper cites Calvin: A benchmark for language- conditioned policy learning for long-horizon robot manip- ulation tasks.

Training-free Generation of Temporally Consistent Rewards from VLMs Calvin: A benchmark for language- conditioned policy learning for long-horizon robot manip- ulation tasks

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:49.277646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:47:41.292412Z digest=sha256:4618a607450d2be6bb1fe90e08cb19706e0d6fed46ceecb757bd490067705bf1

Observation baaa893d-4bfb-4997-956a-a2b1f8e47150 · outbound

This paper cites The deep latent space particle filter for real-time data assimilation with uncertainty quantification.

Training-free Generation of Temporally Consistent Rewards from VLMs The deep latent space particle filter for real-time data assimilation with uncertainty quantification

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:49.089889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:47:41.337357Z digest=sha256:0fe70a7ed0a7e8496ce12c6f96b0811b2c0bdee77738868d4d722df1850e1ab1

Observation bb729636-72e8-4237-8032-cbae885bfbfc · outbound

This paper cites an unresolved cited work.

Training-free Generation of Temporally Consistent Rewards from VLMs Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:47:48.903738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:47:41.384122Z digest=sha256:579ce411b155859409959528334f07fb33caabb61190e392363e8c5e6fc7dda2

Observation ef9ca880-c2e3-4d85-9b2c-7aed41380125 · outbound

This paper cites A real-to-sim-to-real approach to robotic manip- ulation with VLM-generated iterative keypoint rewards.

Training-free Generation of Temporally Consistent Rewards from VLMs A real-to-sim-to-real approach to robotic manip- ulation with VLM-generated iterative keypoint rewards

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:48.721460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:47:41.447612Z digest=sha256:5aac3cb0005a3561ecd135e240b723198dc4c1269d9b965e245934c50a3091f1

Observation 99187c9c-c470-45b2-b0aa-b6df514ad934 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

Training-free Generation of Temporally Consistent Rewards from VLMs Learn- ing transferable visual models from natural language super- vision

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:48.543814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:47:41.501344Z digest=sha256:8788ea638beb39064c81a1729c12dd9cc9dbb68d2004a9f002f581cd1f29e26d

Observation f1662e6a-7c88-47f7-b24c-85f183ef3f94 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

Training-free Generation of Temporally Consistent Rewards from VLMs SAM 2: Segment Anything in Images and Videos

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:41.565323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:47:41.565323Z digest=sha256:42b85295e4d1af548b5c4a783d276fb5daa85a8854e2e863fe45d8052e9bbe8c

Observation 23706177-03b2-4fbf-8698-4cba3fcd5a9d · outbound

This paper cites Vision-language models are zero- shot reward models for reinforcement learning.

Training-free Generation of Temporally Consistent Rewards from VLMs Vision-language models are zero- shot reward models for reinforcement learning

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:48.405348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:47:41.604704Z digest=sha256:92f74fa67ccf98f9a16a5e63922af0e571fb9d3cf163c2718b5e6701b1fdd1be

Observation b25714cc-0c6e-47d5-8b07-a3043dcb36dd · outbound

This paper cites Chinchali.

Training-free Generation of Temporally Consistent Rewards from VLMs Chinchali

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:48.164013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:47:41.664896Z digest=sha256:515d6e388192a3e4975db5dd1b5fd497fc55c3900d82ec17840acd436c70c53e

Observation 10d0986b-78cb-4f54-971c-ac9ea122b65e · outbound

This paper cites Cliport: What and where pathways for robotic manipulation.

Training-free Generation of Temporally Consistent Rewards from VLMs Cliport: What and where pathways for robotic manipulation

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:47.942743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:47:41.728121Z digest=sha256:2be0816c066ec8274edea0a1de1543718606272bad8e2281183230401f19748e

Observation b1f21a69-874f-41b1-b189-c8ff53451e41 · outbound

This paper cites Perceiver- actor: A multi-task transformer for robotic manipulation.

Training-free Generation of Temporally Consistent Rewards from VLMs Perceiver- actor: A multi-task transformer for robotic manipulation

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:47.735982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:47:41.817923Z digest=sha256:50e75fb082390c06c9f74c0b5faa4106ae5795a24a3bdecae27e8fe6e235d954

Observation b4c1f25f-8f94-4208-8291-1bdc5dcf0673 · outbound

This paper cites Progprompt: Generating situ- ated robot task plans using large language models.

Training-free Generation of Temporally Consistent Rewards from VLMs Progprompt: Generating situ- ated robot task plans using large language models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:47.543615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:47:41.891782Z digest=sha256:fcaa58d266451dbd9d2e4f9311e73e919b992d063984eff539594064060da8fc

Observation c54b97b5-9066-497b-999e-157c0dd7a806 · outbound

This paper cites Reason-rft: Reinforcement fine-tuning for visual reasoning.

Training-free Generation of Temporally Consistent Rewards from VLMs Reason-rft: Reinforcement fine-tuning for visual reasoning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:41.937872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:47:41.937872Z digest=sha256:790c83d2d001d634f47a8332c662d5a1ca9122c4318e6107f7f1ceff78b1b24b

Observation 38d4dd40-061d-4311-b806-f3d669eca342 · outbound

This paper cites Cotdet: Affordance knowledge prompting for task driven object de- tection.

Training-free Generation of Temporally Consistent Rewards from VLMs Cotdet: Affordance knowledge prompting for task driven object de- tection

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:47.390815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:47:41.990751Z digest=sha256:db3edee920f4c604be10d86f392521527bf8d47cfc22e253e4449194fe20b443

Observation 97be814c-a937-4040-b3cd-c67d0f46dc28 · outbound

This paper cites AffordGrasp: In-Context Affordance Reasoning for Open-Vocabulary Task-Oriented Grasping in Clutter.

Training-free Generation of Temporally Consistent Rewards from VLMs AffordGrasp: In-Context Affordance Reasoning for Open-Vocabulary Task-Oriented Grasping in Clutter

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:42.042018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:47:42.042018Z digest=sha256:8136ca5bb154b284ab9cb911b532035b08ddb88b4f6a3a3c2991afc965bbab34

Observation 54bc3c62-7248-4cab-8587-cf1fc9edbe6f · outbound

This paper cites Real-World Offline Reinforcement Learning from Vision Language Model Feedback.

Training-free Generation of Temporally Consistent Rewards from VLMs Real-World Offline Reinforcement Learning from Vision Language Model Feedback

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:42.135720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:47:42.135720Z digest=sha256:13d628adf91d6df6a2f6e9f92502bbe152b3937cf1054d1cc1eaffa1b0735c9f

Observation 27fd4f0c-f04d-427c-98df-383ab589ccce · outbound

This paper cites Code as reward: Empowering reinforcement learning with vlms.

Training-free Generation of Temporally Consistent Rewards from VLMs Code as reward: Empowering reinforcement learning with vlms

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:47.174138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:47:42.250922Z digest=sha256:0959db682a59fb203561ab074da2819a77a36ffe676bf035e4cd2e2649d59a89

Observation bf4a0aaf-3f49-4f04-8f85-551259fb6f6f · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Training-free Generation of Temporally Consistent Rewards from VLMs Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:42.366889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:47:42.366889Z digest=sha256:ae117a34bb46453318a9e130f501029087ec22a88bbc477dbcbdf2de7bfa8c89

Observation 20d27e66-116d-4cbc-85aa-aea9cdc545fd · outbound

This paper cites Rl-vlm-f: Rein- forcement learning from vision language foundation model feedback.

Training-free Generation of Temporally Consistent Rewards from VLMs Rl-vlm-f: Rein- forcement learning from vision language foundation model feedback

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:46.922946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:47:42.528976Z digest=sha256:9c0499fad84bfa6121e3a2f50a89b973cb4f8fae715cc1d576041ba2c07b7c72

Observation fb08866f-8e6e-4dcc-a15c-23a99cda2ebb · outbound

This paper cites Any-point Trajectory Modeling for Policy Learning.

Training-free Generation of Temporally Consistent Rewards from VLMs Any-point Trajectory Modeling for Policy Learning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:42.614493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:47:42.614493Z digest=sha256:74efc9965ad3d4fdbc32288925cdf8cb233c393e8aefd7e0bd568b244ee844df

Observation 7dce3e65-a42f-40f7-8ed2-0869618216e2 · outbound

This paper cites RoboMIND: Benchmark on Multi-embodiment Intelligence Normative Data for Robot Manipulation.

Training-free Generation of Temporally Consistent Rewards from VLMs RoboMIND: Benchmark on Multi-embodiment Intelligence Normative Data for Robot Manipulation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:42.681720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:47:42.681720Z digest=sha256:e73b44fa40a03d68757e8be79f49e0a6c8149b53627f8e9f759fd7e2e827c01c

Observation 63959c63-c49c-4b44-90e8-5d05e92514a1 · outbound

This paper cites Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V.

Training-free Generation of Temporally Consistent Rewards from VLMs Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:42.750529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:47:42.750529Z digest=sha256:137fca52acc1246b002319e24932b8595120d942386df9e708af0e49a9decae4

Observation 7d982682-9648-4ece-baef-88452c82b263 · outbound

This paper cites Robot fine- tuning made easy: Pre-training rewards and policies for au- tonomous real-world reinforcement learning.

Training-free Generation of Temporally Consistent Rewards from VLMs Robot fine- tuning made easy: Pre-training rewards and policies for au- tonomous real-world reinforcement learning

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:46.783281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:47:42.851403Z digest=sha256:6f420b1297ff7f8d0a884242f38ad83f458d8c81f288cbd395679dcd5feb5a07

Observation e2bb6c5d-f225-4e7b-806a-51239a3d6acf · outbound

This paper cites Par- ticle filters in latent space for robust deformable linear ob- ject tracking.

Training-free Generation of Temporally Consistent Rewards from VLMs Par- ticle filters in latent space for robust deformable linear ob- ject tracking

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:46.603324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:47:42.923587Z digest=sha256:d50175b2e3800b0111d420e3b4ebcbf8c0889c04f6560c0194dbd2de62f03c2a

Observation 4f28e9b1-6896-45fa-8085-92ec2fb70bd8 · outbound

This paper cites Sornet: Spatial object-centric representations for se- quential manipulation.

Training-free Generation of Temporally Consistent Rewards from VLMs Sornet: Spatial object-centric representations for se- quential manipulation

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:46.378775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:47:42.987023Z digest=sha256:c2fb72e91fedf507d6329d8642cd5b2836bf2f1b67782720c55ab6e23ef54096

Observation 218006e5-280a-4e00-837c-38e8e384aa60 · outbound

This paper cites Robopoint: A vision-language model for spatial affordance prediction in robotics.

Training-free Generation of Temporally Consistent Rewards from VLMs Robopoint: A vision-language model for spatial affordance prediction in robotics

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:46.158520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:47:43.078554Z digest=sha256:46b4287170d0184010c773e3d7631d42b8bd979c6e096fc531a0a11263ad4d18

Observation c571f2e5-095b-4f16-bfaf-fea0bb00de83 · outbound

This paper cites Sam-e: Leveraging visual foundation model with sequence imitation for embodied ma- nipulation.

Training-free Generation of Temporally Consistent Rewards from VLMs Sam-e: Leveraging visual foundation model with sequence imitation for embodied ma- nipulation

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:46.014193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:47:43.151441Z digest=sha256:23dd65ad9361c9305d29631c9d8ba20f1e8aa61e3467bac53ba91a008ab4bffc

Observation 429af5d3-1448-4a2c-a090-499261c9dcb6 · outbound

This paper cites MapNav: A Novel Memory Representation via Annotated Semantic Maps for Vision-and-Language Navigation.

Training-free Generation of Temporally Consistent Rewards from VLMs MapNav: A Novel Memory Representation via Annotated Semantic Maps for Vision-and-Language Navigation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:43.218177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:47:43.218177Z digest=sha256:2893b24f1e0736687ea93a5c5112b810839261cbabddcbea336e9e1562c0da03

Observation e2bde3eb-7b89-43d8-ac34-52fb8082f21b · outbound

This paper cites TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies.

Training-free Generation of Temporally Consistent Rewards from VLMs TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:43.290054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:47:43.290054Z digest=sha256:205d4f487999cc0e19c462d3aee11bff99c06f2d0064a92658f6162dd03f900d

Observation e9e5d94c-0ddc-42ff-9e4a-a575fa5fcfce · outbound

This paper cites an unresolved cited work.

Training-free Generation of Temporally Consistent Rewards from VLMs Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:47:45.846680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:47:43.353180Z digest=sha256:4220524c4eda1fbcef10e6cb7a427c3d017cb281acda1fff8e35c36c9b8f4c7a

Observation 6a8e0e12-be39-4bf6-86ce-c1690d17dbb5 · outbound

This paper cites an unresolved cited work.

Training-free Generation of Temporally Consistent Rewards from VLMs Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:47:45.286813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:47:43.546848Z digest=sha256:b562d7add71b6dc24cead896a56b646ffca254ce42ee8847eae42bad331f4c6f

Observation 95211ab0-e2b3-4a49-8da8-7f736371a782 · outbound

This paper cites Figure 11.

Training-free Generation of Temporally Consistent Rewards from VLMs Figure 11

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:45.081039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:47:43.639443Z digest=sha256:e2254bf8a06ed62914d34c0a829e3ab4fe2af890920b7ff443ff5787b9097a03

Observation a9656fe6-cc2f-4b2f-bdfb-2fb0cc0ead1c · outbound

This paper cites an unresolved cited work.

Training-free Generation of Temporally Consistent Rewards from VLMs Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:47:44.831500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:47:43.723254Z digest=sha256:5faba33a03e2170911c07ba2e14588c3320eecd45b4f01377934f4096962b31b

Observation d2b99751-8743-4fa9-a03e-4a64aaf32e54 · outbound

This paper cites an unresolved cited work.

Training-free Generation of Temporally Consistent Rewards from VLMs Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:47:45.446685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:47:43.790046Z digest=sha256:65a6e0651d84236c2597714e9bd5af92e93a77b10988f3ae5a5d281cc97922c2

Observation db95df95-ace1-4a26-899d-51c16c8dae9c · outbound

This paper cites an unresolved cited work.

Training-free Generation of Temporally Consistent Rewards from VLMs Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:47:45.620439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:47:43.909897Z digest=sha256:c88810c925d90b3b60e109d5a33af30a29b361058e7250f53a035185017b3ad7

Observation 4fe6474c-0e82-4c38-a003-dcb777e0049d · outbound

This paper cites Completion Status Identification of Sub-goals: System prompt: Detailed in Fig.

Training-free Generation of Temporally Consistent Rewards from VLMs Completion Status Identification of Sub-goals: System prompt: Detailed in Fig

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:44.614948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:47:44.014826Z digest=sha256:5f00118a23951a5088821527dc6f99164c5a27b91c7e007af80f96247d345d99

Pith citing papers

No inbound Pith citation observations are available.