Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T00:50:26.697393Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 0 inbound Pith citation observations for arXiv:2608.04396.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T00:50:26.697393Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
52 of 52 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation b6f71cab-e266-4bca-a4b2-a37a35ac6020 · outbound
CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention RT-1: Robotics Transformer for Real-World Control at Scale
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0d3629f-de88-4a33-afae-05b55bf2769b · outbound
CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention RT-2: Vision-language-action models transfer web knowledge to robotic control
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8054f6bc-7f7d-49b3-8c2c-7d38459445be · outbound
CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Vision-language foundation models as effective robot imitators
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e4d9d27-88a6-43eb-801b-a74f053256e0 · outbound
CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Unleashing large-scale video generative pre-training for visual robot manipulation
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation de901bc5-da2d-4a13-8a56-60ef2f258d72 · outbound
CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention GR00T N1: An Open Foundation Model for Generalist Humanoid Robots
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a84c605b-de37-4893-b914-3ab9d360c93b · outbound
CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Moto: Latent motion token as the bridging language for learning robot manipulation from videos
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation fa97f0b7-d4f5-4d50-a2eb-f89691d234d0 · outbound
CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention TraceVLA: Visual trace prompting enhances spatial-temporal awareness for generalist robotic policies
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3dccfc03-3048-4a58-aa66-18bc8943b220 · outbound
CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention GraspVLA: a grasping foundation model pre-trained on billion-scale synthetic action data
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 950d67d4-e34a-42e8-9020-189a221d7359 · outbound
CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Hi Robot: Open-ended instruction following with hierarchical vision- language-action models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b2e030c3-90e5-4512-a19b-bde21f40aaf4 · outbound
CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 334f9ecf-77de-4d50-b55e-cc5c419d6bb2 · outbound
CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention CAST: Counterfactual Labels Improve Instruction Following in Vision-Language-Action Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70fa75bf-ec1c-4514-ad6d-9a5b8400625d · outbound
CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Counterfactual vla: Self-reflective vision-language-action model with adaptive reasoning.arXiv preprint arXiv:2512.24426, 2025
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76cc0ade-143d-440d-93dd-9885211e7555 · outbound
CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Classifier-Free Diffusion Guidance
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0e05596-ebcc-4894-b6e8-0fd1a317c0e0 · outbound
CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention OpenVLA: An Open-Source Vision-Language-Action Model
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 328daaaf-493d-4209-ab76-dba55535505a · outbound
CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention SimpleVLA-RL: Scaling VLA training via reinforcement learning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa8e7fd5-2319-4cea-958d-92ed7357fc33 · outbound
CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96a12dda-f2d5-48a8-9880-5ec808568bd6 · outbound
CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e64a4e7-578d-4e02-b565-a1dc4f37b850 · outbound
CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 982ebd07-d75a-44f5-84ae-80905517c7cc · outbound
CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention LoRA: Low-rank adaptation of large language models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9da8812f-a214-4a2a-bc9b-7f517c603a13 · outbound
CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Vla-adapter: An effective paradigm for tiny-scale vision-language-action model
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8f0d7a59-8f2c-4ec3-89a7-146f29d6cbb3 · outbound
CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Modality-experts coordinated adaptation for large multimodal models.Science China Information Sciences, 67(12):220107, 2024
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 989d7759-df49-48b8-8d10-8d5dc02bf9fd · outbound
CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Reconvla: Reconstructive vision-language-action model as effective robot perceiver
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d027d5e6-f9fa-4b9e-b998-5674eb2b85ff · outbound
CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Ren, Homer Walke, Quan Vuong, Lucy Xiaoyang Shi, and Sergey Levine
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ef64f7c8-53f2-411b-bfa1-bf139dc3d026 · outbound
CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Mole-vla: Dynamic layer-skipping vision language action model via mixture-of-layers for efficient robot manipulation
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0bdd637b-2375-4ef6-ba94-f6dae6e81490 · outbound
CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Spec-vla: speculative decoding for vision-language-action models with relaxed acceptance
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 05c8d3a0-d08d-4427-87ed-eb25f436b9b0 · outbound
CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention CogVLA: Cognition-aligned vision-language- action models via instruction-driven routing & sparsification
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 74a9e8c4-7ed1-45d4-82aa-6729f3ee0575 · outbound
CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Denoising diffusion probabilistic models
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 9fbb998f-301e-4a7d-9d4d-c4c1c02f40df · outbound
CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Diffusion policy: Visuomotor policy learning via action diffusion.The International Journal of Robotics Research, 44(10-11):1684–1704, 2025
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7bc1185-a8ee-46ef-b8be-95de4db797dd · outbound
CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Scalable diffusion models with transformers
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 84406f65-ee5a-4195-a3f6-ab8de5a8e57c · outbound
CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Unresolved cited work
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db5fcb8d-4137-4673-9b66-8116ba002fb0 · outbound
CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Flow straight and fast: Learning to generate and transfer data with rectified flow
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73767142-335e-4cc5-ac9a-25f01b4b4263 · outbound
CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention RDT-1b: a diffusion foundation model for bimanual manipulation
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4434afed-374d-4c48-9a2a-b4f9a02ec9ad · outbound
CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention One-step diffusion policy: Fast visuomotor policies via diffusion distillation
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e42e281f-9351-4c78-942a-60da2f88e4e6 · outbound
CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3445cde-6445-4ff2-8205-2ae91651aa53 · outbound
CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c28f8b25-f58c-447a-98a8-e9e4117e7712 · outbound
CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Toward causal representation learning.Proceedings of the IEEE, 109(5):612– 634, 2021
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4f17e05d-003b-4982-8a36-037f59e12389 · outbound
CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Robust agents learn causal world models
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7cfb8aa3-cc4b-4258-b3d7-2da483407873 · outbound
CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Causalworld: A robotic manipulation benchmark for causal structure and transfer learning
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3f3d1f1c-b783-4c1d-8853-01868deeac0f · outbound
CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention CDP: Towards robust autoregressive visuomotor policy learning via causal diffusion
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a797afcb-a6b0-4019-a748-25670f9c45f8 · outbound
CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention When Vision Overrides Language: Evaluating and Mitigating Counterfactual Failures in VLAs
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66296234-8303-4ae9-bf87-c98c28877ba3 · outbound
CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Libero: Benchmarking knowledge transfer for lifelong robot learning
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 691f0a42-a22b-4acd-9884-a271c21c9b58 · outbound
CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention DreamVLA: A vision-language-action model dreamed with comprehensive world knowledge
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8ed0afc1-d33b-4a28-8a35-021510895d9a · outbound
CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention X-VLA: Soft-prompted transformer as scalable cross-embodiment vision-language-action model
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 08851ab8-5ee7-4209-ad8d-2bbaed0d8fd4 · outbound
CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention NORA: A Small Open-Sourced Generalist Vision Language Action Model for Embodied Tasks
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f043e5b6-9f45-40ab-87a8-37eee2e7b6d7 · outbound
CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention WorldVLA: Towards Autoregressive Action World Model
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 529042c5-7ed2-4491-881c-669a61b1ffc2 · outbound
CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention UniVLA: Learning to Act Anywhere with Task-centric Latent Actions
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8976f62e-7d14-4537-a903-38dfc0e62830 · outbound
CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention FAST: Efficient Action Tokenization for Vision-Language-Action Models
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45f4e954-eadb-44ce-b3d2-cce441c7b7ec · outbound
CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Chance- constrained flow matching for high-fidelity constraint-aware generation.arXiv preprint arXiv:2509.25157, 2025
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3744e82a-5315-4c4b-bc5a-918182fb1a6b · outbound
CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention DDAT: Diffusion Policies Enforcing Dynamically Admissible Robot Trajectories
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c11d84a-a8ba-4999-a779-48b54ea3248d · outbound
CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention LeRobot: An open- source library for end-to-end robot learning
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 485b8b7b-004a-4ad0-b326-07755b41e02b · outbound
CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5956c433-731f-4b2f-bdaa-dc692277a9d3 · outbound
CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention remove the cuboid from the blue plate
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
No inbound Pith citation observations are available.