Pith. sign in

Paper Citation Record · LEDGER

Refined Policy Distillation: From VLA Generalists to RL Experts

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2503.05833.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.05833 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:47:41.904894Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T09:59:44.657362Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f796dcec-51f1-4328-b176-1228a1f59f21 · inbound

Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success cites this paper.

Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success Refined Policy Distillation: From VLA Generalists to RL Experts

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:35:32.529359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T04:35:31.914360Z digest=sha256:484a89587c8864d5076d74a434740616a164479765dfa2d3984848e03d041731

Observation abb0bb71-57f2-43ce-9940-7aa8cc843015 · inbound

Reinforcement Learning for Flow-Matching Policies cites this paper.

Reinforcement Learning for Flow-Matching Policies Refined Policy Distillation: From VLA Generalists to RL Experts

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T15:47:41.904894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:47:41.904894Z digest=sha256:09da0155c3459558ae490428e67115dcf2162c3247b754813fa9ab6de605440c

Observation b83d3683-e736-4568-b56e-a6281430d8fb · inbound

DyGRO-VLA: Cross-Task Scaling of Vision-Language-Action Models via Dynamic Grouped Residual Optimization cites this paper.

DyGRO-VLA: Cross-Task Scaling of Vision-Language-Action Models via Dynamic Grouped Residual Optimization Refined Policy Distillation: From VLA Generalists to RL Experts

Reference 96

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:43:17.053707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-20T12:39:50.004269Z digest=sha256:ff0c523cdf1f5a55f950623570fe0ed054cad98778e94552f95907c774204f8e

Observation d05f9e5d-5475-452f-b74c-dd15a729dd97 · inbound

Object-Centric Residual RL for Zero-Shot Sim-to-Real VLA Enhancement cites this paper.

Object-Centric Residual RL for Zero-Shot Sim-to-Real VLA Enhancement Refined Policy Distillation: From VLA Generalists to RL Experts

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:39:16.668191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T21:06:48.539679Z digest=sha256:7e2f159f0b739c033c009a546024725769c40d215b63bdeadf49c0459157cfa8

Observation 7374c3ab-d5bb-48ab-bb1d-86fdc0af1c0a · inbound

Learning Process Rewards via Success Visitation Matching for Efficient RL cites this paper.

Learning Process Rewards via Success Visitation Matching for Efficient RL Refined Policy Distillation: From VLA Generalists to RL Experts

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:59:44.659192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T09:20:35.062060Z digest=sha256:960a87b3268e56f838a7e27677bfd0fb122b9e3cddaeaa45cd9a43123f2b8c9f

Observation 52ffddce-aa70-4d9e-a9c8-5d04a09dfcbe · inbound

Adapting Generalist Robot Policies with Semantic Reinforcement Learning cites this paper.

Adapting Generalist Robot Policies with Semantic Reinforcement Learning Refined Policy Distillation: From VLA Generalists to RL Experts

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-01T10:45:42.686683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-01T05:09:29.625066Z digest=sha256:c5718a141f6119f75372cadf9cba1a961059e2296f7a15f75e76e38ea8bb2264

Observation 23ea7129-caa0-4f88-976b-fe4ff7462614 · inbound

Teach it to stop, not just to click cites this paper.

Teach it to stop, not just to click Refined Policy Distillation: From VLA Generalists to RL Experts

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T18:59:06.138498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:59:06.138498Z digest=sha256:f42863db0d6759cbd337b71c96c6ccef5d80d85f0e92fca02a44b614e1c399b2