Pith. sign in

Paper Citation Record · LEDGER

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention

As of 11 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 0 inbound Pith citation observations for arXiv:2608.04396.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.04396 v1

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:50:26.697393Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

52 of 52 outbound references displayed

  • verified exact0
  • verified fuzzy24
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b6f71cab-e266-4bca-a4b2-a37a35ac6020 · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention RT-1: Robotics Transformer for Real-World Control at Scale

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:22.779361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:22.779361Z digest=sha256:6356ec5e8093537e139c1da2d1165a447a5c4cd55b7e67d2ed682a1328c9bb56

Observation e0d3629f-de88-4a33-afae-05b55bf2769b · outbound

This paper cites RT-2: Vision-language-action models transfer web knowledge to robotic control.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention RT-2: Vision-language-action models transfer web knowledge to robotic control

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:30.659067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:50:22.845872Z digest=sha256:ce66302bd1ff30e2a04998961a4d1969975afefa3192a54d712e2c8ab853ad15

Observation 8054f6bc-7f7d-49b3-8c2c-7d38459445be · outbound

This paper cites Vision-language foundation models as effective robot imitators.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Vision-language foundation models as effective robot imitators

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:22.923393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:22.923393Z digest=sha256:3ec4319325d9e732045adf61ce752b056cec6f581ec63a4eb363b23bf5085045

Observation 4e4d9d27-88a6-43eb-801b-a74f053256e0 · outbound

This paper cites Unleashing large-scale video generative pre-training for visual robot manipulation.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Unleashing large-scale video generative pre-training for visual robot manipulation

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:30.634338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:50:23.010380Z digest=sha256:cac949f567dd088790de4749e660f8384554f0cbeff1ed8968baea15dcb161d0

Observation de901bc5-da2d-4a13-8a56-60ef2f258d72 · outbound

This paper cites GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:23.096156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:23.096156Z digest=sha256:d0787575ffb42f014377ade8fe921f2e567f3a61413b29cf8f773f3f601eff7d

Observation a84c605b-de37-4893-b914-3ab9d360c93b · outbound

This paper cites Moto: Latent motion token as the bridging language for learning robot manipulation from videos.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Moto: Latent motion token as the bridging language for learning robot manipulation from videos

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:30.619984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:50:23.183105Z digest=sha256:8d8c3b5f202d283cb98076875e6d575b0adc82f6fb6d3c9a4910a451d358acf1

Observation fa97f0b7-d4f5-4d50-a2eb-f89691d234d0 · outbound

This paper cites TraceVLA: Visual trace prompting enhances spatial-temporal awareness for generalist robotic policies.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention TraceVLA: Visual trace prompting enhances spatial-temporal awareness for generalist robotic policies

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:30.605687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:50:23.271170Z digest=sha256:0a26290f66c2dc65e4c431a2cc35c6fcf5cde69acebeed82d16208551d967db5

Observation 3dccfc03-3048-4a58-aa66-18bc8943b220 · outbound

This paper cites GraspVLA: a grasping foundation model pre-trained on billion-scale synthetic action data.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention GraspVLA: a grasping foundation model pre-trained on billion-scale synthetic action data

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:30.585764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:50:23.346135Z digest=sha256:136bf4d47107635da0629eba7d06e5dc55fd7883c908dbaf46c28a0d3820e90a

Observation 950d67d4-e34a-42e8-9020-189a221d7359 · outbound

This paper cites Hi Robot: Open-ended instruction following with hierarchical vision- language-action models.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Hi Robot: Open-ended instruction following with hierarchical vision- language-action models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:30.561099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:50:23.458815Z digest=sha256:925a225a5d52a64146f0c1dda2884914b2f7cc4525ca887382b86093b93716b7

Observation b2e030c3-90e5-4512-a19b-bde21f40aaf4 · outbound

This paper cites LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:23.546418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:23.546418Z digest=sha256:f543f36605fe60bc803793bc3f93dde70ada8fb3819289ef39aaf8dcc4db5b3e

Observation 334f9ecf-77de-4d50-b55e-cc5c419d6bb2 · outbound

This paper cites CAST: Counterfactual Labels Improve Instruction Following in Vision-Language-Action Models.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention CAST: Counterfactual Labels Improve Instruction Following in Vision-Language-Action Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:23.635635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:23.635635Z digest=sha256:0c1fcf59dcbe825474f4b5579ad7edfab007af11537de5265ee86534c39216da

Observation 70fa75bf-ec1c-4514-ad6d-9a5b8400625d · outbound

This paper cites Counterfactual vla: Self-reflective vision-language-action model with adaptive reasoning.arXiv preprint arXiv:2512.24426, 2025.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Counterfactual vla: Self-reflective vision-language-action model with adaptive reasoning.arXiv preprint arXiv:2512.24426, 2025

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:23.716866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:23.716866Z digest=sha256:d5f1ecf644579ea000270cd47d1d1d00ef07c752e0bb733435c22dfb39121316

Observation 76cc0ade-143d-440d-93dd-9885211e7555 · outbound

This paper cites Classifier-Free Diffusion Guidance.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Classifier-Free Diffusion Guidance

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:23.784864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:23.784864Z digest=sha256:3da183f4cc74b6be9ca1e369535e16ffb642ececab19177956675696a24d66c7

Observation a0e05596-ebcc-4894-b6e8-0fd1a317c0e0 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention OpenVLA: An Open-Source Vision-Language-Action Model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:23.872232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:23.872232Z digest=sha256:8983c4e90de8c41156973e69568ad4790fd1bd818c022cf3685418f8c271baaa

Observation 328daaaf-493d-4209-ab76-dba55535505a · outbound

This paper cites SimpleVLA-RL: Scaling VLA training via reinforcement learning.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention SimpleVLA-RL: Scaling VLA training via reinforcement learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:23.973337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:23.973337Z digest=sha256:9cf007f58c3a49d62541d37e76cc151e2d56249db79660a57c347809c3a1e03e

Observation aa8e7fd5-2319-4cea-958d-92ed7357fc33 · outbound

This paper cites SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:24.060801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:24.060801Z digest=sha256:ba56976e9c4286f7a380c44b7d88078dc63c7ab02216102bc40bb6fd5fff0d19

Observation 96a12dda-f2d5-48a8-9880-5ec808568bd6 · outbound

This paper cites $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:24.175572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:24.175572Z digest=sha256:bb4cb78f96dbc968f2c9c4194662a0e511325042804f68a239ce06e086540c2a

Observation 1e64a4e7-578d-4e02-b565-a1dc4f37b850 · outbound

This paper cites Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:24.270042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:24.270042Z digest=sha256:2b6032218c6b5922715eefb2389ead1488920b5f4c0fc3b691e871e10a96b343

Observation 982ebd07-d75a-44f5-84ae-80905517c7cc · outbound

This paper cites LoRA: Low-rank adaptation of large language models.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention LoRA: Low-rank adaptation of large language models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:24.353418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:24.353418Z digest=sha256:00402afa2b59ce903358b81d5f685b3ed7b49e0cd823591f1d2f4fa403c961ba

Observation 9da8812f-a214-4a2a-bc9b-7f517c603a13 · outbound

This paper cites Vla-adapter: An effective paradigm for tiny-scale vision-language-action model.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Vla-adapter: An effective paradigm for tiny-scale vision-language-action model

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:30.403862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:50:24.417889Z digest=sha256:c841afd13253f3b80d444318e75bafec99d3050efab44651d203c6475b712c87

Observation 8f0d7a59-8f2c-4ec3-89a7-146f29d6cbb3 · outbound

This paper cites Modality-experts coordinated adaptation for large multimodal models.Science China Information Sciences, 67(12):220107, 2024.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Modality-experts coordinated adaptation for large multimodal models.Science China Information Sciences, 67(12):220107, 2024

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:30.177614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:50:24.510018Z digest=sha256:b9fa2d8aa907bde934ab0ef78edef3f037d0c5231c4f75ee8f1b356adcf82161

Observation 989d7759-df49-48b8-8d10-8d5dc02bf9fd · outbound

This paper cites Reconvla: Reconstructive vision-language-action model as effective robot perceiver.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Reconvla: Reconstructive vision-language-action model as effective robot perceiver

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:29.928195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:50:24.597474Z digest=sha256:9c431d292699eff06c2fb86d875d24ce7e36f6466d43f47aa7bee0d41c80d032

Observation d027d5e6-f9fa-4b9e-b998-5674eb2b85ff · outbound

This paper cites Ren, Homer Walke, Quan Vuong, Lucy Xiaoyang Shi, and Sergey Levine.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Ren, Homer Walke, Quan Vuong, Lucy Xiaoyang Shi, and Sergey Levine

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:29.722818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:50:24.682916Z digest=sha256:d50381e107589d081225a9039e6939f3e57d2a71588aba46e1bd0cd126b1e203

Observation ef64f7c8-53f2-411b-bfa1-bf139dc3d026 · outbound

This paper cites Mole-vla: Dynamic layer-skipping vision language action model via mixture-of-layers for efficient robot manipulation.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Mole-vla: Dynamic layer-skipping vision language action model via mixture-of-layers for efficient robot manipulation

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:29.576787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:50:24.772533Z digest=sha256:9f86694b06b5b219f0c9338aeed63f305b0ce9a23f3d61819cec66c5937c8e07

Observation 0bdd637b-2375-4ef6-ba94-f6dae6e81490 · outbound

This paper cites Spec-vla: speculative decoding for vision-language-action models with relaxed acceptance.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Spec-vla: speculative decoding for vision-language-action models with relaxed acceptance

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:29.370742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:50:24.860165Z digest=sha256:c17adc7de176ab05e89d4c494b27f75a88f9671cf81d6d5dedf501d3651cafc3

Observation 05c8d3a0-d08d-4427-87ed-eb25f436b9b0 · outbound

This paper cites CogVLA: Cognition-aligned vision-language- action models via instruction-driven routing & sparsification.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention CogVLA: Cognition-aligned vision-language- action models via instruction-driven routing & sparsification

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:29.232451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:50:24.986073Z digest=sha256:2ca16c6a4960c098751b736981aa66466db61f3aa12764ae5e0655c1775b77e7

Observation 74a9e8c4-7ed1-45d4-82aa-6729f3ee0575 · outbound

This paper cites Denoising diffusion probabilistic models.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Denoising diffusion probabilistic models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:29.005648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:50:25.077882Z digest=sha256:8683c8262beeb66b7a9c5defa76b3479bb268de5c6b85e6751c06d1d1b558728

Observation 9fbb998f-301e-4a7d-9d4d-c4c1c02f40df · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action diffusion.The International Journal of Robotics Research, 44(10-11):1684–1704, 2025.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Diffusion policy: Visuomotor policy learning via action diffusion.The International Journal of Robotics Research, 44(10-11):1684–1704, 2025

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:25.175066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:25.175066Z digest=sha256:bc25132b01c2b56d7c2b562cd2f1b383f8b2f3e092c66542e1a40a2321230f50

Observation f7bc1185-a8ee-46ef-b8be-95de4db797dd · outbound

This paper cites Scalable diffusion models with transformers.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Scalable diffusion models with transformers

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:28.771650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:50:25.241353Z digest=sha256:2b2d9af14c87613075e3f16797f63783c72a446ca7a0f05667f585e543df38dd

Observation 84406f65-ee5a-4195-a3f6-ab8de5a8e57c · outbound

This paper cites an unresolved cited work.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:25.326690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:25.326690Z digest=sha256:661da981c106ef7f6764b7ebb33240c6246e2ee70b278a870d227589858184fe

Observation db5fcb8d-4137-4673-9b66-8116ba002fb0 · outbound

This paper cites Flow straight and fast: Learning to generate and transfer data with rectified flow.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Flow straight and fast: Learning to generate and transfer data with rectified flow

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:25.380701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:25.380701Z digest=sha256:9d6f58b5a999cab75ccd7508ea054b947d7e478e198b1fc173ae1dcfd01641ef

Observation 73767142-335e-4cc5-ac9a-25f01b4b4263 · outbound

This paper cites RDT-1b: a diffusion foundation model for bimanual manipulation.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention RDT-1b: a diffusion foundation model for bimanual manipulation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:25.431921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:25.431921Z digest=sha256:0916392521ea44ba57433b879bda1c88ce3b13bb2f4406d14637b13504bb71bb

Observation 4434afed-374d-4c48-9a2a-b4f9a02ec9ad · outbound

This paper cites One-step diffusion policy: Fast visuomotor policies via diffusion distillation.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention One-step diffusion policy: Fast visuomotor policies via diffusion distillation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:25.499650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:25.499650Z digest=sha256:c1998d3e9bf1435681a8622cda178aab967724ca600fb3ad5b9e70f582d559f6

Observation e42e281f-9351-4c78-942a-60da2f88e4e6 · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:25.563377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:25.563377Z digest=sha256:ee02b6bb8db981d3dd87a463c10e31ab1a1721d3b05a6cef2286bb40625b2924

Observation b3445cde-6445-4ff2-8205-2ae91651aa53 · outbound

This paper cites HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:25.631024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:25.631024Z digest=sha256:e920dc1a6aaa209e52031ef2b066e43e0cec319911b033b5183cc187a2d0dde1

Observation c28f8b25-f58c-447a-98a8-e9e4117e7712 · outbound

This paper cites Toward causal representation learning.Proceedings of the IEEE, 109(5):612– 634, 2021.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Toward causal representation learning.Proceedings of the IEEE, 109(5):612– 634, 2021

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:28.480887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:50:25.697551Z digest=sha256:5588e4bd5ed621e3a2f10eba0051a77a9c162afb04d5a5153560a08a42b1e8ab

Observation 4f17e05d-003b-4982-8a36-037f59e12389 · outbound

This paper cites Robust agents learn causal world models.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Robust agents learn causal world models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:28.267867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:50:25.744032Z digest=sha256:df2edc9ce82b58f4354bb85cd184b28b8fbaca319499503837cdd28b71977c27

Observation 7cfb8aa3-cc4b-4258-b3d7-2da483407873 · outbound

This paper cites Causalworld: A robotic manipulation benchmark for causal structure and transfer learning.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Causalworld: A robotic manipulation benchmark for causal structure and transfer learning

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:28.110818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:50:25.794703Z digest=sha256:9fb2f1a8016556f091b535ca559ada77601b2bc41ca619d319fadadca43dfe98

Observation 3f3d1f1c-b783-4c1d-8853-01868deeac0f · outbound

This paper cites CDP: Towards robust autoregressive visuomotor policy learning via causal diffusion.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention CDP: Towards robust autoregressive visuomotor policy learning via causal diffusion

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:28.007250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:50:25.861570Z digest=sha256:fe42bc6ace0942190cbfdbf64c17591ed3cd673994762d7521dfb4455c9e3aca

Observation a797afcb-a6b0-4019-a748-25670f9c45f8 · outbound

This paper cites When Vision Overrides Language: Evaluating and Mitigating Counterfactual Failures in VLAs.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention When Vision Overrides Language: Evaluating and Mitigating Counterfactual Failures in VLAs

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:25.944143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:25.944143Z digest=sha256:8935c9ee7b720221387aee3ecf345c04e393301eb27e2342abdca98ed51e3cce

Observation 66296234-8303-4ae9-bf87-c98c28877ba3 · outbound

This paper cites Libero: Benchmarking knowledge transfer for lifelong robot learning.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Libero: Benchmarking knowledge transfer for lifelong robot learning

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:27.941903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:50:26.007215Z digest=sha256:dc0322fb15d63afd0ad222628b63bc9bdb1243687eb402d7866118fa769b30b5

Observation 691f0a42-a22b-4acd-9884-a271c21c9b58 · outbound

This paper cites DreamVLA: A vision-language-action model dreamed with comprehensive world knowledge.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention DreamVLA: A vision-language-action model dreamed with comprehensive world knowledge

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:27.844071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:50:26.057062Z digest=sha256:57159716ab39c3005672ee7935273b1a324def356af6416c258b1d2f82107ea7

Observation 8ed0afc1-d33b-4a28-8a35-021510895d9a · outbound

This paper cites X-VLA: Soft-prompted transformer as scalable cross-embodiment vision-language-action model.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention X-VLA: Soft-prompted transformer as scalable cross-embodiment vision-language-action model

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:27.648242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:50:26.128371Z digest=sha256:fe310539187a126a1b52a69a2811ea06b882999f868d5141a6ca8ecdba7ad410

Observation 08851ab8-5ee7-4209-ad8d-2bbaed0d8fd4 · outbound

This paper cites NORA: A Small Open-Sourced Generalist Vision Language Action Model for Embodied Tasks.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention NORA: A Small Open-Sourced Generalist Vision Language Action Model for Embodied Tasks

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:26.194504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:26.194504Z digest=sha256:c3e45f2b9435927bf27e4957d7c5f97ae969d5437338607dfa9e68310412d2ee

Observation f043e5b6-9f45-40ab-87a8-37eee2e7b6d7 · outbound

This paper cites WorldVLA: Towards Autoregressive Action World Model.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention WorldVLA: Towards Autoregressive Action World Model

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:26.254317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:26.254317Z digest=sha256:ce4d61e774e29a964e0b56435cdbe4323f19f127edc7a7687dd2827b449f0240

Observation 529042c5-7ed2-4491-881c-669a61b1ffc2 · outbound

This paper cites UniVLA: Learning to Act Anywhere with Task-centric Latent Actions.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention UniVLA: Learning to Act Anywhere with Task-centric Latent Actions

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:26.320751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:26.320751Z digest=sha256:2f0ee01f4731f836993d3f72d4fd32b1396e0d8468d69b5c9cb545abc2b3f611

Observation 8976f62e-7d14-4537-a903-38dfc0e62830 · outbound

This paper cites FAST: Efficient Action Tokenization for Vision-Language-Action Models.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:26.384984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:26.384984Z digest=sha256:3327d6d254eb69ed243242c79f1c64f09a8de240841283be73adb0e6314bbdce

Observation 45f4e954-eadb-44ce-b3d2-cce441c7b7ec · outbound

This paper cites Chance- constrained flow matching for high-fidelity constraint-aware generation.arXiv preprint arXiv:2509.25157, 2025.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Chance- constrained flow matching for high-fidelity constraint-aware generation.arXiv preprint arXiv:2509.25157, 2025

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:26.448351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:26.448351Z digest=sha256:ed98208c76d908f7f3472109232298e277478bb951399f48e0772b69f2983386

Observation 3744e82a-5315-4c4b-bc5a-918182fb1a6b · outbound

This paper cites DDAT: Diffusion Policies Enforcing Dynamically Admissible Robot Trajectories.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention DDAT: Diffusion Policies Enforcing Dynamically Admissible Robot Trajectories

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:26.507752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:26.507752Z digest=sha256:f6f9b9520888cc45d4018547f7d6275f4f910a89644dcf7e3efcd64ba75613ee

Observation 0c11d84a-a8ba-4999-a779-48b54ea3248d · outbound

This paper cites LeRobot: An open- source library for end-to-end robot learning.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention LeRobot: An open- source library for end-to-end robot learning

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:27.475474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:50:26.576375Z digest=sha256:27bbc6ea734a150266fe14e1c20aba68472021d001e5d49b0759bde56d440779

Observation 485b8b7b-004a-4ad0-b326-07755b41e02b · outbound

This paper cites LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:26.636070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:26.636070Z digest=sha256:7c674184857221b35effdcb49b322e93aae92a37f5b291de05eafc58fc01b4e3

Observation 5956c433-731f-4b2f-bdaa-dc692277a9d3 · outbound

This paper cites remove the cuboid from the blue plate.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention remove the cuboid from the blue plate

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:27.296826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:50:26.697393Z digest=sha256:1556b2ab40ced68f7c2e995919db67d0822e8fb8566ec09e68b78b49a80404c1

Pith citing papers

No inbound Pith citation observations are available.