Pith. sign in

Paper Citation Record · LEDGER

IRL-VLA: Training an Vision-Language-Action Policy via Reward World Model

As of 7 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 18 inbound Pith citation observations for arXiv:2508.06571.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.06571 v3

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T23:41:31.838029Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T16:27:24.257956Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T08:36:59.779636Z

Reference resolution

31 of 31 outbound references displayed

  • verified exact0
  • verified fuzzy12
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 19a43edb-9e70-49d0-8bd3-e205ac0bd3a0 · outbound

This paper cites GPT-4 Technical Report.

IRL-VLA: Training an Vision-Language-Action Policy via Reward World Model GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T23:41:31.725595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:41:31.725595Z digest=sha256:7b0e406866dc8a74e0aee03a9c4f770cae8373559486b827ecc68d77eb4527ba

Observation 73b426f7-a4dc-4022-a6de-0d63445c8c70 · outbound

This paper cites Training diffusion models with reinforcement learning, 2024.

IRL-VLA: Training an Vision-Language-Action Policy via Reward World Model Training diffusion models with reinforcement learning, 2024

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T23:41:32.349252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T23:41:31.731882Z digest=sha256:21e9e5b7693bee16396c86ef71cbcc9cb5d8600e0a6a34086f55af0cb1384768

Observation 10117937-04c5-4fdc-a9e3-327b56b8ff8a · outbound

This paper cites Pseudo- simulation for autonomous driving.

IRL-VLA: Training an Vision-Language-Action Policy via Reward World Model Pseudo- simulation for autonomous driving

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T23:41:31.736375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:41:31.736375Z digest=sha256:5a8c009f5a5692731f4d7109968e9a30b2e254ac7bb892d8cbfc93637f370d8c

Observation 12831453-a6e2-4209-8fce-de07c480d253 · outbound

This paper cites Transfuser: Imitation with transformer-based sensor fusion for au- tonomous driving.

IRL-VLA: Training an Vision-Language-Action Policy via Reward World Model Transfuser: Imitation with transformer-based sensor fusion for au- tonomous driving

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T23:41:32.339969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T23:41:31.740137Z digest=sha256:6ea2f8cec97900eb7d6ea298884512535eccbe7e50bcb37b39cb7cb63868dc5b

Observation 1c3578b4-b05c-4f19-97c9-5a75c2cb28ba · outbound

This paper cites Parting with misconceptions about learning-based vehicle motion planning.

IRL-VLA: Training an Vision-Language-Action Policy via Reward World Model Parting with misconceptions about learning-based vehicle motion planning

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T23:41:32.330769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T23:41:31.743912Z digest=sha256:2219a560e66cef5faf26799610739a2832cd8a266c2950b8d13f8ce1f9373fe1

Observation e3db77b6-c139-4555-9b31-0fb3986738a0 · outbound

This paper cites Navsim: Data-driven non-reactive autonomous vehicle simulation and benchmarking.

IRL-VLA: Training an Vision-Language-Action Policy via Reward World Model Navsim: Data-driven non-reactive autonomous vehicle simulation and benchmarking

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T23:41:32.321930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T23:41:31.747756Z digest=sha256:15f7d8ac931fc90b2b64ded898a609b0a8fb258c632c302b3742d69505c4234b

Observation 6bc197cc-12e3-47c6-bc39-60c10520e2c1 · outbound

This paper cites ORION: A Holistic End-to-End Autonomous Driving Framework by Vision-Language Instructed Action Generation.

IRL-VLA: Training an Vision-Language-Action Policy via Reward World Model ORION: A Holistic End-to-End Autonomous Driving Framework by Vision-Language Instructed Action Generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T23:41:31.751498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:41:31.751498Z digest=sha256:a3ef06a4a2e922cdc2af4e8c2cebcf3b1ee3e5e4e3bc435dc8277662fa05003f

Observation be38b2d6-bcb0-4ed9-94cd-b5159061e992 · outbound

This paper cites Rad: Training an end-to-end driving policy via large-scale 3dgs-based reinforcement learning.

IRL-VLA: Training an Vision-Language-Action Policy via Reward World Model Rad: Training an end-to-end driving policy via large-scale 3dgs-based reinforcement learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T23:41:31.755666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:41:31.755666Z digest=sha256:621c7114514a40b959fa3d9a45417b821d18481ed1694696b40a8427ebeb8f06

Observation 48f5bac0-284e-4315-9f00-8a476dcdee68 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

IRL-VLA: Training an Vision-Language-Action Policy via Reward World Model DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T23:41:31.759015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:41:31.759015Z digest=sha256:b5463a42b0a096918c53c808bf2ed2449e38125e524bbb5bbceb5e1a0ab9c7da

Observation fb25938f-71dc-4de2-8767-58b5d5feba6a · outbound

This paper cites Planning-oriented autonomous driving.

IRL-VLA: Training an Vision-Language-Action Policy via Reward World Model Planning-oriented autonomous driving

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T23:41:32.312110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T23:41:31.762813Z digest=sha256:26185edfa5df435944cb85aa1f9f946b43ef801aaf776368fc062c60cdb257eb

Observation 9a9efb1e-292e-48fb-8693-635745dd3494 · outbound

This paper cites EMMA: End-to-End Multimodal Model for Autonomous Driving.

IRL-VLA: Training an Vision-Language-Action Policy via Reward World Model EMMA: End-to-End Multimodal Model for Autonomous Driving

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T23:41:31.766176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:41:31.766176Z digest=sha256:f7b5021ecea468b43be5beb7af6d07bc96afeb8970cbe2ee5cd483d8b7855f14

Observation e9227cc0-15a1-4ff1-b4f9-4abff732a5c1 · outbound

This paper cites Planning with Diffusion for Flexible Behavior Synthesis.

IRL-VLA: Training an Vision-Language-Action Policy via Reward World Model Planning with Diffusion for Flexible Behavior Synthesis

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T23:41:31.769794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:41:31.769794Z digest=sha256:1e7e341f39f36adcb6e7ae25f10c1ce0051b66c99202ff4c397bb7c0fda9d618

Observation 539c03c9-b0de-49b4-acd6-2035043bae97 · outbound

This paper cites DiffVLA: Vision-Language Guided Diffusion Planning for Autonomous Driving.

IRL-VLA: Training an Vision-Language-Action Policy via Reward World Model DiffVLA: Vision-Language Guided Diffusion Planning for Autonomous Driving

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T23:41:31.774080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:41:31.774080Z digest=sha256:19f739a58d1bae15b8274819b3bf4bc5a24550346f40ccc88c94dcab9d9e81bb

Observation 213b3c2c-5e5c-4467-bafe-94ad9546f0c9 · outbound

This paper cites Senna: Bridging Large Vision-Language Models and End-to-End Autonomous Driving.

IRL-VLA: Training an Vision-Language-Action Policy via Reward World Model Senna: Bridging Large Vision-Language Models and End-to-End Autonomous Driving

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T23:41:31.777980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:41:31.777980Z digest=sha256:e85a0ae433de885399a8f879f1d10319117e13b5b908722472df0844bc698717

Observation 1ff95cbe-0d21-43ac-b262-001535e5e840 · outbound

This paper cites Vad: Vectorized scene representation for efficient autonomous driving.

IRL-VLA: Training an Vision-Language-Action Policy via Reward World Model Vad: Vectorized scene representation for efficient autonomous driving

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T23:41:32.303150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T23:41:31.781549Z digest=sha256:25b4ec6e2b91a8467bfe4912ea903945e979b2ca520f33d68b947c4c7d329b4d

Observation 230a0b04-8927-41b3-9e69-ca7de315e0e4 · outbound

This paper cites End-to-End Driving with Online Trajectory Evaluation via BEV World Model.

IRL-VLA: Training an Vision-Language-Action Policy via Reward World Model End-to-End Driving with Online Trajectory Evaluation via BEV World Model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T23:41:31.785006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:41:31.785006Z digest=sha256:39e096494df48047270d023ed75c88ae98a2c27e0d4b324c5f242ec4eb7e9ad4

Observation b9d1c1f7-27d1-4edd-ae28-ea730f060092 · outbound

This paper cites ReCogDrive: A Reinforced Cognitive Framework for End-to-End Autonomous Driving.

IRL-VLA: Training an Vision-Language-Action Policy via Reward World Model ReCogDrive: A Reinforced Cognitive Framework for End-to-End Autonomous Driving

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T23:41:31.788746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:41:31.788746Z digest=sha256:73acd65ba918e7652d852ca8bdced9d7e3792155b81a75e3fdb8a34a5338ee9a

Observation acc7543c-7e1d-4d01-898f-e197281976e8 · outbound

This paper cites Hydra-MDP: End-to-end Multimodal Planning with Multi-target Hydra-Distillation.

IRL-VLA: Training an Vision-Language-Action Policy via Reward World Model Hydra-MDP: End-to-end Multimodal Planning with Multi-target Hydra-Distillation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T23:41:31.792633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:41:31.792633Z digest=sha256:754a1c85523f89d7d5cdf483bebcacda9cee0f8d3aaf709ee60171afd1309c15

Observation 71701eec-44e2-461e-9382-3e4e5b9274f0 · outbound

This paper cites Generalized Trajectory Scoring for End-to-end Multimodal Planning.

IRL-VLA: Training an Vision-Language-Action Policy via Reward World Model Generalized Trajectory Scoring for End-to-end Multimodal Planning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T23:41:31.796375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:41:31.796375Z digest=sha256:1aeae817483e6bcda55dd84ec0898c651ca21f8d84b938020a22e600841eb21f

Observation 2be68a7b-7bbd-49c4-9637-7eca3de7b726 · outbound

This paper cites DiffusionDrive: Truncated Diffusion Model for End-to-End Autonomous Driving.

IRL-VLA: Training an Vision-Language-Action Policy via Reward World Model DiffusionDrive: Truncated Diffusion Model for End-to-End Autonomous Driving

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T23:41:31.799889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:41:31.799889Z digest=sha256:2411c65e8a4374b150fead3be94d650cc092516a2cc03e5d5166000613c8de4f

Observation 46338617-f4df-429c-b157-1ed26c088973 · outbound

This paper cites Diffusiondrive: Trun- cated diffusion model for end-to-end autonomous driv- ing.

IRL-VLA: Training an Vision-Language-Action Policy via Reward World Model Diffusiondrive: Trun- cated diffusion model for end-to-end autonomous driv- ing

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T23:41:32.293240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T23:41:31.803399Z digest=sha256:d97179d9fbfa7ee7c927850b8096c41d1a818ebbca6843e891941edc5e4bee29

Observation c3cab5c8-2364-4ba4-b226-026b9df6c532 · outbound

This paper cites Diffusion Policy Policy Optimization.

IRL-VLA: Training an Vision-Language-Action Policy via Reward World Model Diffusion Policy Policy Optimization

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T23:41:31.807052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:41:31.807052Z digest=sha256:025eafc75c64adfab94ac2e6563220dad7bed683d9e95bd68cbb7d33f5bf9b5e

Observation a7fe2bb6-be6c-4f05-bd13-76e76d5c13a4 · outbound

This paper cites Simlingo: Vision-only closed-loop au- tonomous driving with language-action alignment.

IRL-VLA: Training an Vision-Language-Action Policy via Reward World Model Simlingo: Vision-only closed-loop au- tonomous driving with language-action alignment

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T23:41:32.283875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T23:41:31.810869Z digest=sha256:7b63936c0cacba6add4a0d761e0b0c4438ab9949d41a5738ad7b426e93aa2548

Observation bd531091-cd19-4492-8439-a8e50e69d93c · outbound

This paper cites High-dimensional continuous control using generalized advantage estima- tion, 2018.

IRL-VLA: Training an Vision-Language-Action Policy via Reward World Model High-dimensional continuous control using generalized advantage estima- tion, 2018

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T23:41:32.274253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T23:41:31.813938Z digest=sha256:fa713dd8d80a5838f4634509fe8045f8ac28e88e3a277f5a6db597bc3b53fa16

Observation 39ef19ad-c823-45e8-a107-2e31eae148bc · outbound

This paper cites an unresolved cited work.

IRL-VLA: Training an Vision-Language-Action Policy via Reward World Model Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T23:41:31.817547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:41:31.817547Z digest=sha256:d1aa140c8f963acd0e4c1d29813172624c4f375088e775e5d5a0425ee8508a3b

Observation 3bfb75c9-dc1a-4684-9f11-68baacb63df7 · outbound

This paper cites SparseDrive: End-to-End Autonomous Driving via Sparse Scene Representation.

IRL-VLA: Training an Vision-Language-Action Policy via Reward World Model SparseDrive: End-to-End Autonomous Driving via Sparse Scene Representation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T23:41:31.820894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:41:31.820894Z digest=sha256:7eb5e409f393f8a9493523cd054db03ead8410aaacc74efb4daa61b889426752

Observation be67406b-29b3-46b7-8ea6-1bf9730e4512 · outbound

This paper cites Diffsemanticfusion: Semantic raster bev fusion for autonomous driving via online hd map diffusion, 2025.

IRL-VLA: Training an Vision-Language-Action Policy via Reward World Model Diffsemanticfusion: Semantic raster bev fusion for autonomous driving via online hd map diffusion, 2025

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T23:41:32.256924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T23:41:31.824390Z digest=sha256:5bda2d964208f846211a46d8d0ac0f77a7cb56996c773509505535c7e542af75

Observation 64093467-0a56-4ce8-aa72-b159878d7a16 · outbound

This paper cites Efficient Reinforcement Learning for Autonomous Driving with Parameterized Skills and Priors.

IRL-VLA: Training an Vision-Language-Action Policy via Reward World Model Efficient Reinforcement Learning for Autonomous Driving with Parameterized Skills and Priors

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T23:41:31.827833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:41:31.827833Z digest=sha256:d7a3dd7da6c8df593eb3da99b14640d6082925df21eea43224e1e8c653f70b91

Observation 6a08dcfb-c46b-421c-951d-930a25083ecc · outbound

This paper cites Carplanner: Consistent auto-regressive trajectory plan- ning for large-scale reinforcement learning in au- tonomous driving.

IRL-VLA: Training an Vision-Language-Action Policy via Reward World Model Carplanner: Consistent auto-regressive trajectory plan- ning for large-scale reinforcement learning in au- tonomous driving

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T23:41:32.246657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T23:41:31.831351Z digest=sha256:a0f82b7883301cce1de585e01849c66bdd449ffd006b18c51682e7058e632eb4

Observation 915dc76c-a89d-4f05-af92-afa619e7bb84 · outbound

This paper cites Accelerating reinforcement learning for autonomous driving using task-agnostic and ego- centric motion skills.

IRL-VLA: Training an Vision-Language-Action Policy via Reward World Model Accelerating reinforcement learning for autonomous driving using task-agnostic and ego- centric motion skills

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T23:41:32.236597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T23:41:31.834652Z digest=sha256:14b18b4ec19cec44a78fc535b2bb77cd00248d1288a2ca5a08089b087e6da9d4

Observation 02553d1d-7218-4414-8240-7aa7760f6556 · outbound

This paper cites Opendrivevla: Towards end-to-end autonomous driving with large vision language action model.

IRL-VLA: Training an Vision-Language-Action Policy via Reward World Model Opendrivevla: Towards end-to-end autonomous driving with large vision language action model

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T23:41:31.838029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:41:31.838029Z digest=sha256:d888b730ecf16af9f7df3fc71eaed691057cc4882463326387d8b5b3b65aaf00

Pith citing papers

Observation 0bf917a1-ce2c-4b1c-9be6-5d5b4f46a41b · inbound

World-Env: Leveraging World Model as a Virtual Environment for VLA Post-Training cites this paper.

World-Env: Leveraging World Model as a Virtual Environment for VLA Post-Training IRL-VLA: Training an Vision-Language-Action Policy via Reward World Model

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-18T12:51:23.519299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T12:48:32.123998Z digest=sha256:173d11c30b11053fc681031c0c0222ebb7991d0e5915254152e7c55e14254e8b

Observation b8797e64-ab29-46c2-9fbb-9996a361131f · inbound

Alpamayo-R1: Bridging Reasoning and Action Prediction for Generalizable Autonomous Driving in the Long Tail cites this paper.

Alpamayo-R1: Bridging Reasoning and Action Prediction for Generalizable Autonomous Driving in the Long Tail IRL-VLA: Training an Vision-Language-Action Policy via Reward World Model

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-18T02:35:13.275712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T02:35:13.126171Z digest=sha256:8956c63f2717707ec795e1fdf5129d4b948db6394ec5b7cd95ba87a30707dd58

Observation b1fde534-10ca-400a-bfb6-388f62c45197 · inbound

AVA-VLA: Improving Vision-Language-Action models with Active Visual Attention cites this paper.

AVA-VLA: Improving Vision-Language-Action models with Active Visual Attention IRL-VLA: Training an Vision-Language-Action Policy via Reward World Model

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-17T06:29:09.833895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T06:28:22.652509Z digest=sha256:71eecf4915411696a47da51d08909efe7a447a1c3f471d112d4ae404bc6997d2

Observation bba8b92a-e0e6-496f-b302-b97ab1d846c4 · inbound

Latent Chain-of-Thought World Modeling for End-to-End Driving cites this paper.

Latent Chain-of-Thought World Modeling for End-to-End Driving IRL-VLA: Training an Vision-Language-Action Policy via Reward World Model

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-16T23:18:39.859691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T23:16:41.916869Z digest=sha256:63d57e6b169e777142fa4b811e5d50ab347064e1743b72aaed163a0f860b7801

Observation 78427e88-2f27-4152-ac07-e3e37463355d · inbound

MindDrive: A Vision-Language-Action Model for Autonomous Driving via Online Reinforcement Learning cites this paper.

MindDrive: A Vision-Language-Action Model for Autonomous Driving via Online Reinforcement Learning IRL-VLA: Training an Vision-Language-Action Policy via Reward World Model

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T16:27:24.257956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:27:24.257956Z digest=sha256:0c7f0728787855bab731aa8bf2385a014e655330d466aae6736c0fcbcc2da6d4

Observation cc3c1ee7-953f-4aaa-a642-ed20e715d474 · inbound

Vision-and-Language Navigation for UAVs: Progress, Challenges, and a Research Roadmap cites this paper.

Vision-and-Language Navigation for UAVs: Progress, Challenges, and a Research Roadmap IRL-VLA: Training an Vision-Language-Action Policy via Reward World Model

Reference 169

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:10:29.734642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T13:48:08.135538Z digest=sha256:ed578df80f47ab2a646b25f8aa19abbbf5cdb18a70bb3cee741530f76c99ecd6

Observation 726b3100-d940-45dd-bc8e-612e43ff0621 · inbound

Human Cognition in Machines: A Unified Perspective of World Models cites this paper.

Human Cognition in Machines: A Unified Perspective of World Models IRL-VLA: Training an Vision-Language-Action Policy via Reward World Model

Reference 80

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T08:12:26.263488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T08:12:15.663761Z digest=sha256:cc83060d930b05755f90b3e8972ac7ccb344cbc2c9c355a8a919cdb231b6a74b

Observation 906c2d25-f002-4967-b119-fad978051e4b · inbound

SpanVLA: Efficient Action Bridging and Learning from Negative-Recovery Samples for Vision-Language-Action Model cites this paper.

SpanVLA: Efficient Action Bridging and Learning from Negative-Recovery Samples for Vision-Language-Action Model IRL-VLA: Training an Vision-Language-Action Policy via Reward World Model

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T12:56:10.808364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:34:29.624029Z digest=sha256:c8f8e509da35a3b4ae8b1fd43e8914147a2574b2f71fd2810957f80afbef2ddb

Observation 6140c01f-3951-42f6-806c-f72b4e4d7bb9 · inbound

Latency Analysis and Optimization of Alpamayo 1 via Efficient Trajectory Generation cites this paper.

Latency Analysis and Optimization of Alpamayo 1 via Efficient Trajectory Generation IRL-VLA: Training an Vision-Language-Action Policy via Reward World Model

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:46:50.857406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:52:02.233177Z digest=sha256:ea7505d15df29f0c7c663c7ba40d76149a87a7df291ea95e98d6cd6f785b1b77

Observation a9519e4e-7fde-4811-a07e-b7089c25c52e · inbound

Distill to Think, Foresee to Act: Cognitive-Physical Reinforcement Learning for Autonomous Driving cites this paper.

Distill to Think, Foresee to Act: Cognitive-Physical Reinforcement Learning for Autonomous Driving IRL-VLA: Training an Vision-Language-Action Policy via Reward World Model

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-21T04:49:35.792564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T04:44:48.349384Z digest=sha256:1761a781ed3e64dd822390427f8c788df00aaf711d8a809d01e449d4fe38e797

Observation b49abf41-3de4-495c-973f-050d5cf5a7f8 · inbound

Distill to Think, Foresee to Act: Cognitive-Physical Reinforcement Learning for Autonomous Driving cites this paper.

Distill to Think, Foresee to Act: Cognitive-Physical Reinforcement Learning for Autonomous Driving IRL-VLA: Training an Vision-Language-Action Policy via Reward World Model

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:55:24.815699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-25T05:52:10.143179Z digest=sha256:572ea5dc093e525aa1c5f3206e0be606b48c33c2ceb8972b13bc7f71e9319af8

Observation 0e8534ad-08b3-4702-9e8b-cbba6dfdff07 · inbound

NTR: Neural Token Reconstruction for Scene Token Bottleneck in End-to-End Driving cites this paper.

NTR: Neural Token Reconstruction for Scene Token Bottleneck in End-to-End Driving IRL-VLA: Training an Vision-Language-Action Policy via Reward World Model

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-06-28T23:02:46.926168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T22:57:01.742536Z digest=sha256:453854e5f2d864f7cb17b6a66d512fdccfe8b391cd3f52ea2081c6dacccb58b1

Observation b00599c5-6f3f-430f-bbc2-6a4f40027e71 · inbound

IDOL: Inverse-Dynamics-Guided Future Prediction for End-to-End Autonomous Driving cites this paper.

IDOL: Inverse-Dynamics-Guided Future Prediction for End-to-End Autonomous Driving IRL-VLA: Training an Vision-Language-Action Policy via Reward World Model

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-01T19:36:08.476569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T22:23:07.223613Z digest=sha256:d2ade330b2cf14b081b488d994dc5bbce381b46c37d5c157b82c8c158ad38d53

Observation ceaf1399-6b15-4a6f-9991-67721a9a07c6 · inbound

World Models for Robotic Manipulation: A Survey cites this paper.

World Models for Robotic Manipulation: A Survey IRL-VLA: Training an Vision-Language-Action Policy via Reward World Model

Reference 106

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T12:33:25.078010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T12:24:18.025364Z digest=sha256:a2a512b7fcc334e536825d8c8b0559edd5cb1cdeb6ff28df444573eae8fb1b04

Observation 860069f4-8aa2-41d2-9a69-2b774bdb3957 · inbound

Discrete-WAM: Unified Discrete Vision-Action Token Editing for World-Policy Learning cites this paper.

Discrete-WAM: Unified Discrete Vision-Action Token Editing for World-Policy Learning IRL-VLA: Training an Vision-Language-Action Policy via Reward World Model

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:46:57.089281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T01:49:25.510681Z digest=sha256:e1a4bba68d333b03943ce93e8217315be113ec8b822057536273664f45016c75

Observation 673c745e-6bdb-452c-9c15-9afbabf66cdb · inbound

World Engine: Towards the Era of Post-Training for Autonomous Driving cites this paper.

World Engine: Towards the Era of Post-Training for Autonomous Driving IRL-VLA: Training an Vision-Language-Action Policy via Reward World Model

Reference 101

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:59:33.870818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T17:21:15.456982Z digest=sha256:83ad565e51677ea58290c1a35351c30f82484c8b871386e4970f3047162b3d7c

Observation 5a592479-c7e2-440b-bb5b-c89a8e3a29bf · inbound

Post-Training in End-to-End Autonomous Driving cites this paper.

Post-Training in End-to-End Autonomous Driving IRL-VLA: Training an Vision-Language-Action Policy via Reward World Model

Reference 24

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T00:46:40.298575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-10T00:42:28.701843Z digest=sha256:5b7ef2ba9af7aa6514fbe6c1d8fc0c6773fc28239ac8fd4b8ed0d799200976c3

Observation 29aca45e-edf5-4daa-8fdf-06915c8773ba · inbound

WCog-VLA: A Dual-Level World-Cognitive Vision-Language-Action Model for End-to-End Autonomous Driving cites this paper.

WCog-VLA: A Dual-Level World-Cognitive Vision-Language-Action Model for End-to-End Autonomous Driving IRL-VLA: Training an Vision-Language-Action Policy via Reward World Model

Reference 24

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T08:36:59.781035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-10T08:30:29.351159Z digest=sha256:0b15834377c8e0ba18410f73c158799e958643320f7c41510ad1dcfb5fe7b00d