Pith. sign in

Paper Citation Record · LEDGER

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention

As of 20 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 0 inbound Pith citation observations for arXiv:2608.04396.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.04396 v1

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:50:26.697393Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

52 of 52 outbound references displayed

  • verified exact0
  • verified fuzzy24
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b6f71cab-e266-4bca-a4b2-a37a35ac6020 · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention RT-1: Robotics Transformer for Real-World Control at Scale

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:22.779361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:22.779361Z digest=sha256:acad1fde26dfa1f1222e6bafc8f972f2b9047548a00698cfcb30f80a377ca356

Observation e0d3629f-de88-4a33-afae-05b55bf2769b · outbound

This paper cites RT-2: Vision-language-action models transfer web knowledge to robotic control.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention RT-2: Vision-language-action models transfer web knowledge to robotic control

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:30.659067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:50:22.845872Z digest=sha256:e617b5871fa3a93cd2c3ad4fc1eff49855cf0db57e854117e5c9d9e6697ebe69

Observation 8054f6bc-7f7d-49b3-8c2c-7d38459445be · outbound

This paper cites Vision-language foundation models as effective robot imitators.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Vision-language foundation models as effective robot imitators

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:22.923393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:22.923393Z digest=sha256:3332d4d9a0625e7434ffbe5d74781154f4a75d5061eb4d6fd649df8fcc40a011

Observation 4e4d9d27-88a6-43eb-801b-a74f053256e0 · outbound

This paper cites Unleashing large-scale video generative pre-training for visual robot manipulation.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Unleashing large-scale video generative pre-training for visual robot manipulation

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:30.634338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:50:23.010380Z digest=sha256:8c4a8fed556f36714d358d65833dbaa0f8acf76ef545b1e27b71f9ab433a2568

Observation de901bc5-da2d-4a13-8a56-60ef2f258d72 · outbound

This paper cites GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:23.096156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:23.096156Z digest=sha256:f4c93491ab5c24511323d2e4f73eb8e7a293bec7f26f89fdc4ea7cb4da8380c8

Observation a84c605b-de37-4893-b914-3ab9d360c93b · outbound

This paper cites Moto: Latent motion token as the bridging language for learning robot manipulation from videos.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Moto: Latent motion token as the bridging language for learning robot manipulation from videos

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:30.619984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:50:23.183105Z digest=sha256:5be2d638d34cf58817ecfb0f200906d6683dace50cf0130edb302fd5c0dbee53

Observation fa97f0b7-d4f5-4d50-a2eb-f89691d234d0 · outbound

This paper cites TraceVLA: Visual trace prompting enhances spatial-temporal awareness for generalist robotic policies.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention TraceVLA: Visual trace prompting enhances spatial-temporal awareness for generalist robotic policies

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:30.605687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:50:23.271170Z digest=sha256:02e6f5eecd14e41aa495f7e3fa6472b21c13be3c713d17dc932eb8ba9868ac5a

Observation 3dccfc03-3048-4a58-aa66-18bc8943b220 · outbound

This paper cites GraspVLA: a grasping foundation model pre-trained on billion-scale synthetic action data.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention GraspVLA: a grasping foundation model pre-trained on billion-scale synthetic action data

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:30.585764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:50:23.346135Z digest=sha256:c1d8e04c078adea4b26fee2e653587acc101673e3c0885f2a8c20d0e095ae93d

Observation 950d67d4-e34a-42e8-9020-189a221d7359 · outbound

This paper cites Hi Robot: Open-ended instruction following with hierarchical vision- language-action models.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Hi Robot: Open-ended instruction following with hierarchical vision- language-action models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:30.561099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:50:23.458815Z digest=sha256:d8b411d87d43ae145c8eb5feebc91549085869d986df346fb5e3440a4f9725b0

Observation b2e030c3-90e5-4512-a19b-bde21f40aaf4 · outbound

This paper cites LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:23.546418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:23.546418Z digest=sha256:c4e88b5e3d65d290d09b29676f6e2d6a94f7104b935f5302d5b0b7c4ce0efa39

Observation 334f9ecf-77de-4d50-b55e-cc5c419d6bb2 · outbound

This paper cites CAST: Counterfactual Labels Improve Instruction Following in Vision-Language-Action Models.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention CAST: Counterfactual Labels Improve Instruction Following in Vision-Language-Action Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:23.635635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:23.635635Z digest=sha256:d8c99537538605113d1d8fd9b5fdedd55a77e8dee69ecef7ef8ffdac87306097

Observation 70fa75bf-ec1c-4514-ad6d-9a5b8400625d · outbound

This paper cites Counterfactual vla: Self-reflective vision-language-action model with adaptive reasoning.arXiv preprint arXiv:2512.24426, 2025.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Counterfactual vla: Self-reflective vision-language-action model with adaptive reasoning.arXiv preprint arXiv:2512.24426, 2025

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:23.716866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:23.716866Z digest=sha256:1aa64d21f59e4ae10d423530e17d5d430a761ff2c86bae020fb09b3f197ad492

Observation 76cc0ade-143d-440d-93dd-9885211e7555 · outbound

This paper cites Classifier-Free Diffusion Guidance.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Classifier-Free Diffusion Guidance

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:23.784864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:23.784864Z digest=sha256:11efdcb5ef7116b2b16b500014e08fd2348084e064a4cd9f24339e3f4554fc7a

Observation a0e05596-ebcc-4894-b6e8-0fd1a317c0e0 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention OpenVLA: An Open-Source Vision-Language-Action Model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:23.872232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:23.872232Z digest=sha256:ec2bef8a88ae1753d2b91bced23b6d11b3ddabdc4af584e1a03b584ec4d6e4b9

Observation 328daaaf-493d-4209-ab76-dba55535505a · outbound

This paper cites SimpleVLA-RL: Scaling VLA training via reinforcement learning.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention SimpleVLA-RL: Scaling VLA training via reinforcement learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:23.973337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:23.973337Z digest=sha256:8008783d51d4bdd84bec4633fd4a702a35562b7097c91c008b3345e6f537ab03

Observation aa8e7fd5-2319-4cea-958d-92ed7357fc33 · outbound

This paper cites SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:24.060801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:24.060801Z digest=sha256:34d654d2ff62c847ff07bee2a0712a2ab013112339dc6006bfdad998a0727f13

Observation 96a12dda-f2d5-48a8-9880-5ec808568bd6 · outbound

This paper cites $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:24.175572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:24.175572Z digest=sha256:692778631ff0edf3bf99d7d647a5214444aec3321f345e2a3e61aec0c01bce2e

Observation 1e64a4e7-578d-4e02-b565-a1dc4f37b850 · outbound

This paper cites Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:24.270042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:24.270042Z digest=sha256:1fdc754efdf17778ac4607866094529956c08be00f5f9da55d3bae927b1ffb8e

Observation 982ebd07-d75a-44f5-84ae-80905517c7cc · outbound

This paper cites LoRA: Low-rank adaptation of large language models.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention LoRA: Low-rank adaptation of large language models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:24.353418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:24.353418Z digest=sha256:64ea1bdeb9f52650c292b32af9c54392f969163e01b5a87b85b438c5a82eb722

Observation 9da8812f-a214-4a2a-bc9b-7f517c603a13 · outbound

This paper cites Vla-adapter: An effective paradigm for tiny-scale vision-language-action model.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Vla-adapter: An effective paradigm for tiny-scale vision-language-action model

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:30.403862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:50:24.417889Z digest=sha256:3f652fb706a6e41fd4f0e26a391a85aed48b970647f6247e510b9186ac0e4fc9

Observation 8f0d7a59-8f2c-4ec3-89a7-146f29d6cbb3 · outbound

This paper cites Modality-experts coordinated adaptation for large multimodal models.Science China Information Sciences, 67(12):220107, 2024.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Modality-experts coordinated adaptation for large multimodal models.Science China Information Sciences, 67(12):220107, 2024

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:30.177614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:50:24.510018Z digest=sha256:5671ce3f18fd93d9e0a7792e54366a67e541f27d2ffd2bb150171f89a695d8cb

Observation 989d7759-df49-48b8-8d10-8d5dc02bf9fd · outbound

This paper cites Reconvla: Reconstructive vision-language-action model as effective robot perceiver.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Reconvla: Reconstructive vision-language-action model as effective robot perceiver

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:29.928195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:50:24.597474Z digest=sha256:a8b80aac2db02f37f646cebd697e91da5fdfc616f55023579155ea4dbeefbb52

Observation d027d5e6-f9fa-4b9e-b998-5674eb2b85ff · outbound

This paper cites Ren, Homer Walke, Quan Vuong, Lucy Xiaoyang Shi, and Sergey Levine.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Ren, Homer Walke, Quan Vuong, Lucy Xiaoyang Shi, and Sergey Levine

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:29.722818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:50:24.682916Z digest=sha256:1ed1b8779308386bbdc88e4661a8f0cd28befe05bd9aa956a95a91209de14169

Observation ef64f7c8-53f2-411b-bfa1-bf139dc3d026 · outbound

This paper cites Mole-vla: Dynamic layer-skipping vision language action model via mixture-of-layers for efficient robot manipulation.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Mole-vla: Dynamic layer-skipping vision language action model via mixture-of-layers for efficient robot manipulation

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:29.576787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:50:24.772533Z digest=sha256:d435518fc7cebb86006708504d87c3f657c72856a9348ee40c7813d2d4ac5899

Observation 0bdd637b-2375-4ef6-ba94-f6dae6e81490 · outbound

This paper cites Spec-vla: speculative decoding for vision-language-action models with relaxed acceptance.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Spec-vla: speculative decoding for vision-language-action models with relaxed acceptance

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:29.370742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:50:24.860165Z digest=sha256:ece915db2cc36f785c054462bd95aa98c5f726b5864770f27e3348f6ee5607a3

Observation 05c8d3a0-d08d-4427-87ed-eb25f436b9b0 · outbound

This paper cites CogVLA: Cognition-aligned vision-language- action models via instruction-driven routing & sparsification.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention CogVLA: Cognition-aligned vision-language- action models via instruction-driven routing & sparsification

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:29.232451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:50:24.986073Z digest=sha256:a559002b5899ac7e7388425aa4601151980810c29d33e7dc57494a2889b5590c

Observation 74a9e8c4-7ed1-45d4-82aa-6729f3ee0575 · outbound

This paper cites Denoising diffusion probabilistic models.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Denoising diffusion probabilistic models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:29.005648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:50:25.077882Z digest=sha256:08e7c4386f05b21caaf4dc03bc8d1e2ffa1be9de06dea385a32a1476f7d5cd37

Observation 9fbb998f-301e-4a7d-9d4d-c4c1c02f40df · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action diffusion.The International Journal of Robotics Research, 44(10-11):1684–1704, 2025.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Diffusion policy: Visuomotor policy learning via action diffusion.The International Journal of Robotics Research, 44(10-11):1684–1704, 2025

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:25.175066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:25.175066Z digest=sha256:c79c1340e0bb512b2dd83d3ad0db05ae736163913b36f1d7c404796c52da74b1

Observation f7bc1185-a8ee-46ef-b8be-95de4db797dd · outbound

This paper cites Scalable diffusion models with transformers.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Scalable diffusion models with transformers

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:28.771650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:50:25.241353Z digest=sha256:6542f73e70fce7ac8c169640c1f6b60f41a1cce6e00f9abf490023c69bc6edd5

Observation 84406f65-ee5a-4195-a3f6-ab8de5a8e57c · outbound

This paper cites an unresolved cited work.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:25.326690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:25.326690Z digest=sha256:142d9ff9c37c81fe0cfad35bf74b8b0be2126432480b5b26dd6b47e8cbf2991b

Observation db5fcb8d-4137-4673-9b66-8116ba002fb0 · outbound

This paper cites Flow straight and fast: Learning to generate and transfer data with rectified flow.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Flow straight and fast: Learning to generate and transfer data with rectified flow

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:25.380701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:25.380701Z digest=sha256:fb75ef4b37a6cfa052475daa87567d4c5a3168ab0b88b6e9fea55ac2696fd405

Observation 73767142-335e-4cc5-ac9a-25f01b4b4263 · outbound

This paper cites RDT-1b: a diffusion foundation model for bimanual manipulation.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention RDT-1b: a diffusion foundation model for bimanual manipulation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:25.431921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:25.431921Z digest=sha256:e94464fc0ae24527f2616c34280fdf099fd705f91124f7a250120e2215f59651

Observation 4434afed-374d-4c48-9a2a-b4f9a02ec9ad · outbound

This paper cites One-step diffusion policy: Fast visuomotor policies via diffusion distillation.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention One-step diffusion policy: Fast visuomotor policies via diffusion distillation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:25.499650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:25.499650Z digest=sha256:867aef3a40e0786a584fae3fb274c571a63e5ca8efd032b0f1d01a5d7088a2cf

Observation e42e281f-9351-4c78-942a-60da2f88e4e6 · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:25.563377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:25.563377Z digest=sha256:429183285931c0d7f7d5b2b35bbb1ff46f4b93432af431b789b1db1a5c07bced

Observation b3445cde-6445-4ff2-8205-2ae91651aa53 · outbound

This paper cites HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:25.631024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:25.631024Z digest=sha256:569ca75e298a9915211efd006d6981cda7b85beeeb4341aaefd4b03590bddee9

Observation c28f8b25-f58c-447a-98a8-e9e4117e7712 · outbound

This paper cites Toward causal representation learning.Proceedings of the IEEE, 109(5):612– 634, 2021.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Toward causal representation learning.Proceedings of the IEEE, 109(5):612– 634, 2021

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:28.480887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:50:25.697551Z digest=sha256:980748ec9ce49084b1810003e4874956f472a266a1961f4dd8a8f0f2b15e4619

Observation 4f17e05d-003b-4982-8a36-037f59e12389 · outbound

This paper cites Robust agents learn causal world models.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Robust agents learn causal world models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:28.267867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:50:25.744032Z digest=sha256:819afc1662794879b8f1bdf4037c82520aacf8be469371701d10a2185375eb99

Observation 7cfb8aa3-cc4b-4258-b3d7-2da483407873 · outbound

This paper cites Causalworld: A robotic manipulation benchmark for causal structure and transfer learning.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Causalworld: A robotic manipulation benchmark for causal structure and transfer learning

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:28.110818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:50:25.794703Z digest=sha256:5131ee6afab22d0c3ec25ac6f47929a849249aa8c546c11d27fd3dc947c450de

Observation 3f3d1f1c-b783-4c1d-8853-01868deeac0f · outbound

This paper cites CDP: Towards robust autoregressive visuomotor policy learning via causal diffusion.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention CDP: Towards robust autoregressive visuomotor policy learning via causal diffusion

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:28.007250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:50:25.861570Z digest=sha256:01307bf0b15a39b0fbc362f04a15fa3afa1439a662d67ddc813d5d03f023fce5

Observation a797afcb-a6b0-4019-a748-25670f9c45f8 · outbound

This paper cites When Vision Overrides Language: Evaluating and Mitigating Counterfactual Failures in VLAs.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention When Vision Overrides Language: Evaluating and Mitigating Counterfactual Failures in VLAs

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:25.944143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:25.944143Z digest=sha256:ea78acd0c31ac045e34a4506386db20ae02910c50f466a4952fd3eb5101ef106

Observation 66296234-8303-4ae9-bf87-c98c28877ba3 · outbound

This paper cites Libero: Benchmarking knowledge transfer for lifelong robot learning.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Libero: Benchmarking knowledge transfer for lifelong robot learning

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:27.941903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:50:26.007215Z digest=sha256:6134ef9626cfc77a3396ff6d1b9e75a5e3299a14a3a084ceb9a39e5f31705542

Observation 691f0a42-a22b-4acd-9884-a271c21c9b58 · outbound

This paper cites DreamVLA: A vision-language-action model dreamed with comprehensive world knowledge.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention DreamVLA: A vision-language-action model dreamed with comprehensive world knowledge

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:27.844071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:50:26.057062Z digest=sha256:c2702c4da2f6f19155127b9e5b889ecc75f61a3e61d0edbfb2c027a63f8d3612

Observation 8ed0afc1-d33b-4a28-8a35-021510895d9a · outbound

This paper cites X-VLA: Soft-prompted transformer as scalable cross-embodiment vision-language-action model.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention X-VLA: Soft-prompted transformer as scalable cross-embodiment vision-language-action model

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:27.648242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:50:26.128371Z digest=sha256:6efa7ea3595c2fcc93aff5049e7a35ee504036fcee6b8547c8a964680f5dda13

Observation 08851ab8-5ee7-4209-ad8d-2bbaed0d8fd4 · outbound

This paper cites NORA: A Small Open-Sourced Generalist Vision Language Action Model for Embodied Tasks.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention NORA: A Small Open-Sourced Generalist Vision Language Action Model for Embodied Tasks

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:26.194504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:26.194504Z digest=sha256:ce6cfc3cde5d5b177f62af21f07c6ca6258a5c0c7f26d5324b41dd7d9ce5e939

Observation f043e5b6-9f45-40ab-87a8-37eee2e7b6d7 · outbound

This paper cites WorldVLA: Towards Autoregressive Action World Model.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention WorldVLA: Towards Autoregressive Action World Model

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:26.254317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:26.254317Z digest=sha256:819994e9fd74b3ef03d76ba7a136af331e5a91cefdb09e7cc8643e88b3616446

Observation 529042c5-7ed2-4491-881c-669a61b1ffc2 · outbound

This paper cites UniVLA: Learning to Act Anywhere with Task-centric Latent Actions.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention UniVLA: Learning to Act Anywhere with Task-centric Latent Actions

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:26.320751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:26.320751Z digest=sha256:bbce91c6d7fc545bf460bacf12ccd22163ab69de28a6830ad421c3b6000c6416

Observation 8976f62e-7d14-4537-a903-38dfc0e62830 · outbound

This paper cites FAST: Efficient Action Tokenization for Vision-Language-Action Models.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:26.384984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:26.384984Z digest=sha256:fd445047f7e49904f4abc3a742329f5d1c4c3d17a8b8dca4393e3bf830734b35

Observation 45f4e954-eadb-44ce-b3d2-cce441c7b7ec · outbound

This paper cites Chance- constrained flow matching for high-fidelity constraint-aware generation.arXiv preprint arXiv:2509.25157, 2025.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Chance- constrained flow matching for high-fidelity constraint-aware generation.arXiv preprint arXiv:2509.25157, 2025

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:26.448351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:26.448351Z digest=sha256:f8d705993d6ce8da40990e96f43457c3e6bc143aaee7aab86d4b0d9b29615389

Observation 3744e82a-5315-4c4b-bc5a-918182fb1a6b · outbound

This paper cites DDAT: Diffusion Policies Enforcing Dynamically Admissible Robot Trajectories.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention DDAT: Diffusion Policies Enforcing Dynamically Admissible Robot Trajectories

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:26.507752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:26.507752Z digest=sha256:9f729bcdebad5c97c63260f15ceee6ccfc63f8c93f80a52e0798dc95aad8bad3

Observation 0c11d84a-a8ba-4999-a779-48b54ea3248d · outbound

This paper cites LeRobot: An open- source library for end-to-end robot learning.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention LeRobot: An open- source library for end-to-end robot learning

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:27.475474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:50:26.576375Z digest=sha256:b8256be8249854262010ea5185f08edaf9e6938e778d8978fa0e6b8c9f4908ca

Observation 485b8b7b-004a-4ad0-b326-07755b41e02b · outbound

This paper cites LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:26.636070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:26.636070Z digest=sha256:3a7c3629d990ad8c1b03413f01054e20f00d088324a7e705ef3148286e67745d

Observation 5956c433-731f-4b2f-bdaa-dc692277a9d3 · outbound

This paper cites remove the cuboid from the blue plate.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention remove the cuboid from the blue plate

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:27.296826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:50:26.697393Z digest=sha256:aff084ee208e7290520db4570bc47935dfd34e6f712dda0e808361d2ff58e77b

Pith citing papers

No inbound Pith citation observations are available.