Pith. sign in

Paper Citation Record · LEDGER

Update-Free On-Policy Steering via Verifiers

As of 4 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 2 inbound Pith citation observations for arXiv:2603.10282.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2603.10282 v2

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-14T23:44:37.937519Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-30T07:49:04.825693Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-04T09:59:44.623252Z

Reference resolution

52 of 52 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved52
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b458c64e-d34e-4084-bbc3-8dc903467df9 · outbound

This paper cites Rt-1: Robotics transformer for real-world control at scale,.

Update-Free On-Policy Steering via Verifiers Rt-1: Robotics transformer for real-world control at scale,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-14T23:44:37.937519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:44:37.937519Z digest=sha256:fa9e97d0283ed8e86265f7d5ccbd9c7d21a3552dd5eb6ef6f5e6c3f320525cf0

Observation ec083a6f-bb15-4d7c-b2af-3a7d0a568bf0 · outbound

This paper cites Rt-2: Vision-language-action models transfer web knowledge to robotic control,.

Update-Free On-Policy Steering via Verifiers Rt-2: Vision-language-action models transfer web knowledge to robotic control,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-14T23:44:37.937519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:44:37.937519Z digest=sha256:1586a8e5ce486a14592eac1d2fcc1206063e8f34abe0e9739c8801659df497df

Observation 251fe879-9c30-4fb8-a078-3081649ed30f · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action diffusion,.

Update-Free On-Policy Steering via Verifiers Diffusion policy: Visuomotor policy learning via action diffusion,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-14T23:44:37.937519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:44:37.937519Z digest=sha256:b575a5572e24647ee64f3c0414b9bcffc607a758bd7fd8eaaf69b4cc64c13ecc

Observation 4352bd92-01bf-48d6-8248-5a25c8b1c5be · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

Update-Free On-Policy Steering via Verifiers $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-14T23:44:37.937519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:44:37.937519Z digest=sha256:e293bd19138a16e28d92a721f182b1d672709b0a07937f7c25154f297b8c1bef

Observation 99d45fd4-3fee-4dc3-b3a3-55fd57257ecb · outbound

This paper cites Openvla: An open-source vision-language-action model,.

Update-Free On-Policy Steering via Verifiers Openvla: An open-source vision-language-action model,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-14T23:44:37.937519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:44:37.937519Z digest=sha256:a77bab9eae71c3c75db5b2eb8ee8f89f07ebdf70ed3b909c80fe1fcee503bf95

Observation e30a4cdb-9f72-4481-8535-86653da50461 · outbound

This paper cites What matters in learning from offline human demonstrations for robot manipulation,.

Update-Free On-Policy Steering via Verifiers What matters in learning from offline human demonstrations for robot manipulation,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-14T23:44:37.937519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:44:37.937519Z digest=sha256:bd0e6f0c7f75e182588c60f19cb0ec5019f8cf3198217020aeaa870d24c705a4

Observation 8e5db4ea-44f2-4bd4-9dbb-4b3a3387be0f · outbound

This paper cites A reduction of imitation learning and structured pre- diction to no-regret online learning,.

Update-Free On-Policy Steering via Verifiers A reduction of imitation learning and structured pre- diction to no-regret online learning,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-14T23:44:37.937519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:44:37.937519Z digest=sha256:6a00bd2aa453a0408c599297c76ae96d5b6a2351f3d130f24315d9f25116fab2

Observation d5b73c3a-838c-44c6-88b8-2c2a465a8337 · outbound

This paper cites Fighting copycat agents in behavioral cloning from ob- servation histories,.

Update-Free On-Policy Steering via Verifiers Fighting copycat agents in behavioral cloning from ob- servation histories,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-14T23:44:37.937519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:44:37.937519Z digest=sha256:2950ca544b6b2f1870a252fdc6b31831593d4104dd912c83a1225290e046cf9a

Observation 96a4faa8-1147-44b4-88eb-7cdd86131131 · outbound

This paper cites Implicit behavioral cloning,.

Update-Free On-Policy Steering via Verifiers Implicit behavioral cloning,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-14T23:44:37.937519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:44:37.937519Z digest=sha256:ce2e037414a6025f28bbc00b555d68de936392bd6f1e14c88ea2086bffca66ac

Observation a3a8be01-3ea7-4dd0-bda0-397501c4f50c · outbound

This paper cites Human-in-the-Loop Imitation Learning using Remote Teleoperation.

Update-Free On-Policy Steering via Verifiers Human-in-the-Loop Imitation Learning using Remote Teleoperation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-14T23:44:37.937519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:44:37.937519Z digest=sha256:531b73f46474188541ca25d735be6470c54da64e2acfb12c043fc5f81e70e820

Observation 0b3a935e-b90a-4643-9a07-42843266dd71 · outbound

This paper cites Behavior transformers: Cloningkmodes with one stone,.

Update-Free On-Policy Steering via Verifiers Behavior transformers: Cloningkmodes with one stone,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-14T23:44:37.937519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:44:37.937519Z digest=sha256:10f1c1ec2f15f8f98b35489b34e41cd34ddfd7f7bbf719b774f5d82e364692c4

Observation 6bb49b78-be5f-4876-b2cd-498b3ac10b03 · outbound

This paper cites Learning fine-grained bimanual manipulation with low-cost hardware,.

Update-Free On-Policy Steering via Verifiers Learning fine-grained bimanual manipulation with low-cost hardware,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-14T23:44:37.937519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:44:37.937519Z digest=sha256:cbfeb58eb4392bd07be815223ec1d7a585c76a7834b437b37d3758d66b899e58

Observation 9c994b07-11c8-4b0a-93d3-b178d2bbed71 · outbound

This paper cites Data quality in imitation learning,.

Update-Free On-Policy Steering via Verifiers Data quality in imitation learning,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-14T23:44:37.937519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:44:37.937519Z digest=sha256:2db9e0dc0dbb0efc86cfcb70a38916703aae5f2ffc77692de29575b25c9df141

Observation 65a131d0-61fa-4890-8c20-fe387551de5f · outbound

This paper cites Hg-dagger: Interactive imitation learning with human experts,.

Update-Free On-Policy Steering via Verifiers Hg-dagger: Interactive imitation learning with human experts,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-14T23:44:37.937519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:44:37.937519Z digest=sha256:fb5a7077532cc97269a5c732bdb2bd08253ca6eac9b9d980108654449495f9d1

Observation 9fe8249a-0028-45ec-a2e1-fffd94c92ce8 · outbound

This paper cites Dart: Noise injection for robust imitation learning,.

Update-Free On-Policy Steering via Verifiers Dart: Noise injection for robust imitation learning,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-14T23:44:37.937519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:44:37.937519Z digest=sha256:105f66a837bbfdc38fed047ebfbf7c3dd29a403b537a33e2327153b9085896c8

Observation a103dd7d-0ca2-437e-b755-e5519c68a125 · outbound

This paper cites Inference-Time Scaling for Diffusion Models beyond Scaling Denoising Steps.

Update-Free On-Policy Steering via Verifiers Inference-Time Scaling for Diffusion Models beyond Scaling Denoising Steps

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-14T23:44:37.937519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:44:37.937519Z digest=sha256:9dfe32bd6a5c74864ff365f2fab3f03f34e9b25b2f483faca56d3042c083dd41

Observation 81fd325d-c2c4-4af3-9915-e57bacce70b6 · outbound

This paper cites Inference-time scaling of diffusion models through classical search,.

Update-Free On-Policy Steering via Verifiers Inference-time scaling of diffusion models through classical search,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-14T23:44:37.937519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:44:37.937519Z digest=sha256:b77a5710e6bf973b68ecfe86360b85d590ac5168feb11ec1df1374aa5ec46490

Observation a9b1c9e6-4caa-4a2f-a85b-4257d267a118 · outbound

This paper cites Universal guidance for diffusion models,.

Update-Free On-Policy Steering via Verifiers Universal guidance for diffusion models,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-14T23:44:37.937519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:44:37.937519Z digest=sha256:35503a5a0e615a489cdc1864dc0e2c57be000829c5f06b44cae659ecb6e5ca2f

Observation 136da44f-7fcf-4eb0-97a1-c3d8e39f21fb · outbound

This paper cites Dynaguide: Steering diffusion polices with active dynamic guidance,.

Update-Free On-Policy Steering via Verifiers Dynaguide: Steering diffusion polices with active dynamic guidance,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-14T23:44:37.937519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:44:37.937519Z digest=sha256:260ee93a74932df49c9152e6a9d87dde4f5e7db2702477334fc518d6cd1ccfa7

Observation 2d2003d0-bd8d-4867-b0ae-fd22179a5421 · outbound

This paper cites Inference-time policy steering through human interac- tions,.

Update-Free On-Policy Steering via Verifiers Inference-time policy steering through human interac- tions,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-14T23:44:37.937519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:44:37.937519Z digest=sha256:7431c569db9ae11e1843da080073bb2003acf2c4eec86da959ff5750628de7e2

Observation 0223b433-cf1f-4b7f-9c24-937c7dce20cd · outbound

This paper cites From foresight to forethought: Vlm-in-the-loop policy steering via latent alignment,.

Update-Free On-Policy Steering via Verifiers From foresight to forethought: Vlm-in-the-loop policy steering via latent alignment,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-14T23:44:37.937519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:44:37.937519Z digest=sha256:1ce599aefa8b7229708602835e0493163433ea03af27db00ee2b43360dc69063

Observation 333d83c1-d898-4070-9cff-ddcff4b97d7e · outbound

This paper cites Fine-tuning reinforcement learning models is secretly a forgetting mitigation problem,.

Update-Free On-Policy Steering via Verifiers Fine-tuning reinforcement learning models is secretly a forgetting mitigation problem,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-14T23:44:37.937519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:44:37.937519Z digest=sha256:50263c2d9114010ca89b15f3934facf602a5c847416aa0c9436d325a031a03f3

Observation 6768c286-cdfa-4714-aa92-7e99f1afcc3c · outbound

This paper cites Robocat: A self-improving generalist agent for robotic manipulation,.

Update-Free On-Policy Steering via Verifiers Robocat: A self-improving generalist agent for robotic manipulation,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-14T23:44:37.937519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:44:37.937519Z digest=sha256:cfedd92cf2cb09e3a0cd70debc1e48924b721a95327f1bbbe6a89968944bd020

Observation d19371eb-a5cb-4f21-81de-07130a9a0bf5 · outbound

This paper cites Sime: Enhancing policy self-improvement with modal- level exploration,.

Update-Free On-Policy Steering via Verifiers Sime: Enhancing policy self-improvement with modal- level exploration,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-14T23:44:37.937519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:44:37.937519Z digest=sha256:7af1384ab2efa37821290106f4d5aa59d6b014a98a61eac0bc67e63d9307ac45

Observation f1792a51-8ef0-4483-9753-e4280d59978b · outbound

This paper cites Soe: Sample-efficient robot policy self-improvement via on-manifold exploration,.

Update-Free On-Policy Steering via Verifiers Soe: Sample-efficient robot policy self-improvement via on-manifold exploration,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-14T23:44:37.937519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:44:37.937519Z digest=sha256:ac6c93b9ae1724f3d3385850b2e1485533721516e803fe7dc70e96589f568673

Observation 26de344e-0a88-4b24-b07d-a11605ab8029 · outbound

This paper cites Selfi: Autonomous self-improvement with reinforce- ment learning for social navigation,.

Update-Free On-Policy Steering via Verifiers Selfi: Autonomous self-improvement with reinforce- ment learning for social navigation,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-14T23:44:37.937519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:44:37.937519Z digest=sha256:6761d1471cd5324e94bd2d192f3c5a6302c2e7c7a60f9514bb8b85f14002e040

Observation bec5b23b-0178-42f0-ac97-d33b2a8c5717 · outbound

This paper cites Awac: Accelerating online reinforcement learning with offline datasets,.

Update-Free On-Policy Steering via Verifiers Awac: Accelerating online reinforcement learning with offline datasets,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-14T23:44:37.937519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:44:37.937519Z digest=sha256:82a20745dcc1265d49c731b5b87302031f32aaaaf8155256857b2aef1b84292f

Observation 177b5958-2f6d-4caf-bd64-c000e2a447fb · outbound

This paper cites Self-improving embodied foundation models,.

Update-Free On-Policy Steering via Verifiers Self-improving embodied foundation models,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-14T23:44:37.937519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:44:37.937519Z digest=sha256:1b0fd6b87990bc986dc38b78a6a15faf5dd458b03ceae1a57d22c33e65d1db06

Observation f3f34b0d-c82e-4882-8232-b0ced5d43091 · outbound

This paper cites Scaling Instructable Agents Across Many Simulated Worlds.

Update-Free On-Policy Steering via Verifiers Scaling Instructable Agents Across Many Simulated Worlds

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-14T23:44:37.937519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:44:37.937519Z digest=sha256:38cfa5f878090f75dca027a4fbbf74908792fada7e85e1d7f6e7fe858775e0dc

Observation 0a91c47b-ddc0-4bb8-886a-204ae4ff0e0f · outbound

This paper cites Seil: Simulation-augmented equivariant imitation learn- ing,.

Update-Free On-Policy Steering via Verifiers Seil: Simulation-augmented equivariant imitation learn- ing,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-14T23:44:37.937519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:44:37.937519Z digest=sha256:4e429ef7d692266ccf442876852bb576427337708785203d042aae724b92e807

Observation d544eb13-b496-4c01-b33c-8b239665a5c8 · outbound

This paper cites Self-augmented robot trajectory: Efficient imitation learn- ing via safe self-augmentation with demonstrator-annotated precision,.

Update-Free On-Policy Steering via Verifiers Self-augmented robot trajectory: Efficient imitation learn- ing via safe self-augmentation with demonstrator-annotated precision,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-14T23:44:37.937519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:44:37.937519Z digest=sha256:d90bc918b8dcade328599b0cd2e1878a3a9d6ab642227c4cfd71ffd451cd4402

Observation bd9fcc78-4b49-476c-a6e3-fec1e2facd9b · outbound

This paper cites Curating demonstrations using online experience,.

Update-Free On-Policy Steering via Verifiers Curating demonstrations using online experience,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-14T23:44:37.937519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:44:37.937519Z digest=sha256:72ec17027a7f937dfc0b0bf4f07f7d52ea601c0a13343a7258111c935fe66439

Observation bfa3b500-bf5c-4b2a-950e-7c42c4aca8d6 · outbound

This paper cites Is conditional generative modeling all you need for decision-making?.

Update-Free On-Policy Steering via Verifiers Is conditional generative modeling all you need for decision-making?

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-14T23:44:37.937519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:44:37.937519Z digest=sha256:76527c2a029d8af9cf751f093b247ca565b422d3b6eda9b8c539f925299cf760

Observation 99916ee1-1754-44d0-ac18-3fa2ad6c56d5 · outbound

This paper cites Diffusion Guidance Is a Controllable Policy Improvement Operator.

Update-Free On-Policy Steering via Verifiers Diffusion Guidance Is a Controllable Policy Improvement Operator

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-14T23:44:37.937519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:44:37.937519Z digest=sha256:cc4603981dd383e9af05726d9b046d48c498fba94cc1890b79547f399872854e

Observation bc5198ba-3e43-4280-8cea-1eeee4990551 · outbound

This paper cites Bidirectional decoding: Improving action chunking via guided test-time sampling,.

Update-Free On-Policy Steering via Verifiers Bidirectional decoding: Improving action chunking via guided test-time sampling,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-14T23:44:37.937519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:44:37.937519Z digest=sha256:9d6f04eb0f96d3c66ace4bd5f317fc79f0f5fa2b123ed1ef04840e0a207d6616

Observation 95cb89bc-bd91-4549-980e-0c8003013f87 · outbound

This paper cites Steering your generalists: Improving robotic foun- dation models via value guidance,.

Update-Free On-Policy Steering via Verifiers Steering your generalists: Improving robotic foun- dation models via value guidance,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-14T23:44:37.937519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:44:37.937519Z digest=sha256:86bc59b5b34f4f1b1e6d4dc898b4fa8e0c3c3f19e6b7b0198c071a8f03468115

Observation 815357fc-fe62-4997-abb1-ac14439895e2 · outbound

This paper cites Steering your diffusion policy with latent space reinforcement learning,.

Update-Free On-Policy Steering via Verifiers Steering your diffusion policy with latent space reinforcement learning,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-14T23:44:37.937519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:44:37.937519Z digest=sha256:c5f19bcab34a5d1490a14a56bc872638382b2dd3009d753f85ca88cf90c2e312

Observation 88977c50-aab6-4d56-b40a-eba76f4b4f04 · outbound

This paper cites Behavioral cloning from observation,.

Update-Free On-Policy Steering via Verifiers Behavioral cloning from observation,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-14T23:44:37.937519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:44:37.937519Z digest=sha256:0253eb32b0793d5b1bce60cd3addab7dc05b05c290ba64f53e7192e1283be436

Observation 2f832458-e998-4444-9632-c5872fa4ea9a · outbound

This paper cites Diffusion models beat gans on image synthesis,.

Update-Free On-Policy Steering via Verifiers Diffusion models beat gans on image synthesis,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-14T23:44:37.937519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:44:37.937519Z digest=sha256:31060d2638ea018fbedd1cc43d076d3154da9ca79606d8de706f84f2a1178ce4

Observation 7a428a7c-712a-4176-953d-daa0d2d510f2 · outbound

This paper cites Denoising diffusion probabilistic models,.

Update-Free On-Policy Steering via Verifiers Denoising diffusion probabilistic models,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-14T23:44:37.937519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:44:37.937519Z digest=sha256:f6bca09ccffc3cb4cee04b2a58349e98dfd7d2e65dcf0748995b79dc4341c983

Observation 8e9bd8bf-a950-4289-b188-a9d51f0deec5 · outbound

This paper cites Denoising diffusion implicit models,.

Update-Free On-Policy Steering via Verifiers Denoising diffusion implicit models,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-14T23:44:37.937519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:44:37.937519Z digest=sha256:616db96a22102271a6ee2aa9706ee0c162af6bfaa80f48dca334b275ca168737

Observation c47ee81d-eab6-48a2-8679-a87ef6beb72f · outbound

This paper cites an unresolved cited work.

Update-Free On-Policy Steering via Verifiers Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-14T23:44:37.937519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:44:37.937519Z digest=sha256:a3e4cfa35de7dd0f5d20ff833bb75f5eb51f7d0d0d1c84bd81ecc51e874ac57e

Observation cf62e0ac-6c58-417e-b5ff-85a99dfabc72 · outbound

This paper cites Implicit generation and generalization in energy-based models,.

Update-Free On-Policy Steering via Verifiers Implicit generation and generalization in energy-based models,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-14T23:44:37.937519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:44:37.937519Z digest=sha256:74df121d68a0140c9df31b067c47066be5c92164d41fb270eead5150d64c2006

Observation d6c36d89-dcc2-475b-b267-023495955fb2 · outbound

This paper cites Attention is all you need,.

Update-Free On-Policy Steering via Verifiers Attention is all you need,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-14T23:44:37.937519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:44:37.937519Z digest=sha256:123a44db1054a2797e7bd2249977397cede97da31bab0d167cdda34fc031fda5

Observation 574db489-54dc-426f-9ce1-3fc2ca63b8cb · outbound

This paper cites Dimensionality reduction by learning an invariant mapping,.

Update-Free On-Policy Steering via Verifiers Dimensionality reduction by learning an invariant mapping,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-14T23:44:37.937519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:44:37.937519Z digest=sha256:583a3f86d617ef46a3bf6b3b909deef7dc1d95e367c216fbd76ae35387a1e235

Observation 243b8656-6784-4354-a18a-a96244682ddf · outbound

This paper cites A smooth sea never made a skilled SAILOR: Robust imitation via learning to search,.

Update-Free On-Policy Steering via Verifiers A smooth sea never made a skilled SAILOR: Robust imitation via learning to search,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-07-14T23:44:37.937519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:44:37.937519Z digest=sha256:752b8fe7593d7510305bf57ec086abfb06841e8df4125283a888b587ba3d0230

Observation 5ca6822d-0c68-49c1-adfe-7083820cd263 · outbound

This paper cites Libero: Benchmarking knowledge transfer for lifelong robot learning,.

Update-Free On-Policy Steering via Verifiers Libero: Benchmarking knowledge transfer for lifelong robot learning,

Reference 47

Resolution
unresolved
no resolver link, observed 2026-07-14T23:44:37.937519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:44:37.937519Z digest=sha256:181eb1722d93cbc3045f7ed39fa1b5c3a53fcb39a2dfa6223600791508826597

Observation 3bc3b7c0-3cff-48f0-b901-bc1f141e678b · outbound

This paper cites Flow matching for generative modeling,.

Update-Free On-Policy Steering via Verifiers Flow matching for generative modeling,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-07-14T23:44:37.937519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:44:37.937519Z digest=sha256:ecaeabe25b9acd876829e9738b1f46640edf8193888dd2be0375b5175d5a8eda

Observation 5505ee8e-646f-4dc9-999a-d6e526b0fb0f · outbound

This paper cites Unsupervised Machine Translation Using Monolingual Corpora Only.

Update-Free On-Policy Steering via Verifiers Unsupervised Machine Translation Using Monolingual Corpora Only

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-14T23:44:37.937519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:44:37.937519Z digest=sha256:503c00b0b5cdafacf5a549693b30844928b8221dfe07d564ad31ceab8b6fae83

Observation 136942d0-f594-4853-a099-165c3cac9fee · outbound

This paper cites Interval estimation for the difference between independent proportions: Comparison of eleven methods,.

Update-Free On-Policy Steering via Verifiers Interval estimation for the difference between independent proportions: Comparison of eleven methods,

Reference 50

Resolution
unresolved
no resolver link, observed 2026-07-14T23:44:37.937519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:44:37.937519Z digest=sha256:bc298204fc2448be3885166a5e480dc2e23037e67e0253fa9cef995f96c241d1

Observation 78780e70-87c6-4ce7-8b3d-444b53de9727 · outbound

This paper cites U-net: Convolutional networks for biomedical image segmentation,.

Update-Free On-Policy Steering via Verifiers U-net: Convolutional networks for biomedical image segmentation,

Reference 51

Resolution
unresolved
no resolver link, observed 2026-07-14T23:44:37.937519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:44:37.937519Z digest=sha256:33e2d95f659baf78a63dff6e7eef8de636b118a01a7619ee405a5920253a0f5e

Observation 9ffd1e20-75a3-429c-ac98-4640ad659f07 · outbound

This paper cites Deep residual learning for image recognition,.

Update-Free On-Policy Steering via Verifiers Deep residual learning for image recognition,

Reference 52

Resolution
unresolved
no resolver link, observed 2026-07-14T23:44:37.937519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:44:37.937519Z digest=sha256:6d22b63ec35f957ddcb73bae7eb170d876aaf227f78b145cfd4a5c5a2cbb6315

Pith citing papers

Observation b7b92bc1-0050-48c4-acbe-3b3b9dc122a4 · inbound

Learning Process Rewards via Success Visitation Matching for Efficient RL cites this paper.

Learning Process Rewards via Success Visitation Matching for Efficient RL Update-Free On-Policy Steering via Verifiers

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-04T09:59:44.624465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T09:20:35.062060Z digest=sha256:4a3496a5fe31f53062a9a0aa0f8d6b25761fa74e350290e0e50f5924656c2ebb

Observation 5cc91ec5-46c6-491c-8dc0-3db97183c7cf · inbound

Behavior Uncloning: Distilling Mode Redirection into Policy Weights without Inference-Time Steering cites this paper.

Behavior Uncloning: Distilling Mode Redirection into Policy Weights without Inference-Time Steering Update-Free On-Policy Steering via Verifiers

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-06-30T07:54:22.221272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T07:49:04.825693Z digest=sha256:f686a2e565d5df79bea9a2176ea650fcf545547b7b7edae9d46bffedcf260d61