Pith. sign in

Paper Citation Record · LEDGER

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention

As of 12 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 0 inbound Pith citation observations for arXiv:2608.04396.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.04396 v1

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:50:26.697393Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

52 of 52 outbound references displayed

  • verified exact0
  • verified fuzzy24
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b6f71cab-e266-4bca-a4b2-a37a35ac6020 · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention RT-1: Robotics Transformer for Real-World Control at Scale

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:22.779361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:22.779361Z digest=sha256:465d6140e38c5acbddae9f9530229f7b0bd6f5c77df0c19f8d9d6dd876cf4c1b

Observation e0d3629f-de88-4a33-afae-05b55bf2769b · outbound

This paper cites RT-2: Vision-language-action models transfer web knowledge to robotic control.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention RT-2: Vision-language-action models transfer web knowledge to robotic control

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:30.659067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:50:22.845872Z digest=sha256:e713398e6936ac2aa9f3527031cf7d13302967848b60592ddb1a561be92397cc

Observation 8054f6bc-7f7d-49b3-8c2c-7d38459445be · outbound

This paper cites Vision-language foundation models as effective robot imitators.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Vision-language foundation models as effective robot imitators

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:22.923393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:22.923393Z digest=sha256:ae289314192a7c7237fedb61e4086f098336189e4eca61756e63a8d336ff9992

Observation 4e4d9d27-88a6-43eb-801b-a74f053256e0 · outbound

This paper cites Unleashing large-scale video generative pre-training for visual robot manipulation.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Unleashing large-scale video generative pre-training for visual robot manipulation

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:30.634338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:50:23.010380Z digest=sha256:6f6ab6d6540cb8e2b37e0efaf798fa106d6f7509546a32a8e9fd6086fe62d3f3

Observation de901bc5-da2d-4a13-8a56-60ef2f258d72 · outbound

This paper cites GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:23.096156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:23.096156Z digest=sha256:08824915a34b554640d08bce9330891abd4f0cb72febe17aa0734c602b818f97

Observation a84c605b-de37-4893-b914-3ab9d360c93b · outbound

This paper cites Moto: Latent motion token as the bridging language for learning robot manipulation from videos.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Moto: Latent motion token as the bridging language for learning robot manipulation from videos

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:30.619984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:50:23.183105Z digest=sha256:6a7cefd45e6fcb9e68e19b3454a845d8ef8ae1f2f0ce4eb091f75802c441cada

Observation fa97f0b7-d4f5-4d50-a2eb-f89691d234d0 · outbound

This paper cites TraceVLA: Visual trace prompting enhances spatial-temporal awareness for generalist robotic policies.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention TraceVLA: Visual trace prompting enhances spatial-temporal awareness for generalist robotic policies

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:30.605687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:50:23.271170Z digest=sha256:ce9bbbca3dceaeebac7ce20b9dd597af7b6233b43c21eafcc98748f445113e75

Observation 3dccfc03-3048-4a58-aa66-18bc8943b220 · outbound

This paper cites GraspVLA: a grasping foundation model pre-trained on billion-scale synthetic action data.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention GraspVLA: a grasping foundation model pre-trained on billion-scale synthetic action data

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:30.585764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:50:23.346135Z digest=sha256:efd56464a2468048b2fa98132d6424c419805b9326fe747a4f62a01982b53890

Observation 950d67d4-e34a-42e8-9020-189a221d7359 · outbound

This paper cites Hi Robot: Open-ended instruction following with hierarchical vision- language-action models.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Hi Robot: Open-ended instruction following with hierarchical vision- language-action models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:30.561099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:50:23.458815Z digest=sha256:a12a74ca89c6483a6e145e719ae4f9bc4a9485967d89209040ef9edab839fa5e

Observation b2e030c3-90e5-4512-a19b-bde21f40aaf4 · outbound

This paper cites LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:23.546418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:23.546418Z digest=sha256:0690c57d423b7b7350b26a83b20a24bcd80818df75d370cced3bb5ab4eff7b6c

Observation 334f9ecf-77de-4d50-b55e-cc5c419d6bb2 · outbound

This paper cites CAST: Counterfactual Labels Improve Instruction Following in Vision-Language-Action Models.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention CAST: Counterfactual Labels Improve Instruction Following in Vision-Language-Action Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:23.635635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:23.635635Z digest=sha256:2f812b9c54a7dda6b8370faff6178f7ff53e3f9dfaa3128dd5f3faf476a01533

Observation 70fa75bf-ec1c-4514-ad6d-9a5b8400625d · outbound

This paper cites Counterfactual vla: Self-reflective vision-language-action model with adaptive reasoning.arXiv preprint arXiv:2512.24426, 2025.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Counterfactual vla: Self-reflective vision-language-action model with adaptive reasoning.arXiv preprint arXiv:2512.24426, 2025

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:23.716866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:23.716866Z digest=sha256:97e7e66535597d667cedfa7acd19ff4bb6d2287beb79de7d648e550d8f23df1e

Observation 76cc0ade-143d-440d-93dd-9885211e7555 · outbound

This paper cites Classifier-Free Diffusion Guidance.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Classifier-Free Diffusion Guidance

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:23.784864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:23.784864Z digest=sha256:b782cbc8b1e99fcff26f81f5dc0d9446a0a0df442cc459bb831550dc8a0e8e35

Observation a0e05596-ebcc-4894-b6e8-0fd1a317c0e0 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention OpenVLA: An Open-Source Vision-Language-Action Model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:23.872232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:23.872232Z digest=sha256:1e38177f89c1aff63076e9d14a31aca1961ed08b5de003e12603fa79de9345d3

Observation 328daaaf-493d-4209-ab76-dba55535505a · outbound

This paper cites SimpleVLA-RL: Scaling VLA training via reinforcement learning.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention SimpleVLA-RL: Scaling VLA training via reinforcement learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:23.973337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:23.973337Z digest=sha256:3f0b478019aad6c44f7f033f9c6c3635bd86145c11406223103b9fee0ed33a1f

Observation aa8e7fd5-2319-4cea-958d-92ed7357fc33 · outbound

This paper cites SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:24.060801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:24.060801Z digest=sha256:a07f42dbb4272a8a025179e807b9b9087c65b085667a98c3bb1bf63ec8476e9b

Observation 96a12dda-f2d5-48a8-9880-5ec808568bd6 · outbound

This paper cites $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:24.175572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:24.175572Z digest=sha256:655880e42fd934e5221ad4f57a8ff3de9c9fe0b7cc28a2e5b222295c6547218c

Observation 1e64a4e7-578d-4e02-b565-a1dc4f37b850 · outbound

This paper cites Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:24.270042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:24.270042Z digest=sha256:15717d53055503a0dd1b57b20acc805182252e5f099d372792329eca5b2ad392

Observation 982ebd07-d75a-44f5-84ae-80905517c7cc · outbound

This paper cites LoRA: Low-rank adaptation of large language models.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention LoRA: Low-rank adaptation of large language models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:24.353418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:24.353418Z digest=sha256:22f7fa381e09102593ad4e090520a2b07717a0a8af05c031c2109744508a5d87

Observation 9da8812f-a214-4a2a-bc9b-7f517c603a13 · outbound

This paper cites Vla-adapter: An effective paradigm for tiny-scale vision-language-action model.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Vla-adapter: An effective paradigm for tiny-scale vision-language-action model

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:30.403862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:50:24.417889Z digest=sha256:b1e5c132aa0c677174c233220cff3d725c44c72e09719005c9c1edd20354a7a9

Observation 8f0d7a59-8f2c-4ec3-89a7-146f29d6cbb3 · outbound

This paper cites Modality-experts coordinated adaptation for large multimodal models.Science China Information Sciences, 67(12):220107, 2024.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Modality-experts coordinated adaptation for large multimodal models.Science China Information Sciences, 67(12):220107, 2024

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:30.177614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:50:24.510018Z digest=sha256:95e0b409d6b14333fb85afe0e3401f200a24ae3890f0bb6040f5398bf6fa13db

Observation 989d7759-df49-48b8-8d10-8d5dc02bf9fd · outbound

This paper cites Reconvla: Reconstructive vision-language-action model as effective robot perceiver.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Reconvla: Reconstructive vision-language-action model as effective robot perceiver

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:29.928195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:50:24.597474Z digest=sha256:afd46f657b70bede4944154ec808f9ebf2ec155492b1a77a5abf574c6c6f9b8a

Observation d027d5e6-f9fa-4b9e-b998-5674eb2b85ff · outbound

This paper cites Ren, Homer Walke, Quan Vuong, Lucy Xiaoyang Shi, and Sergey Levine.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Ren, Homer Walke, Quan Vuong, Lucy Xiaoyang Shi, and Sergey Levine

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:29.722818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:50:24.682916Z digest=sha256:207b9cbddd6e934efca6bba8b40017a13f0299b7b6bd314d82cece669a579279

Observation ef64f7c8-53f2-411b-bfa1-bf139dc3d026 · outbound

This paper cites Mole-vla: Dynamic layer-skipping vision language action model via mixture-of-layers for efficient robot manipulation.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Mole-vla: Dynamic layer-skipping vision language action model via mixture-of-layers for efficient robot manipulation

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:29.576787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:50:24.772533Z digest=sha256:dad250643371f48a4efc8392e84ec35920f654ef407958796661584214cece41

Observation 0bdd637b-2375-4ef6-ba94-f6dae6e81490 · outbound

This paper cites Spec-vla: speculative decoding for vision-language-action models with relaxed acceptance.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Spec-vla: speculative decoding for vision-language-action models with relaxed acceptance

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:29.370742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:50:24.860165Z digest=sha256:0fb33c66955b6547e3c19c24820e6d8ca854a305d0e292c2b8b14b506aa8a77d

Observation 05c8d3a0-d08d-4427-87ed-eb25f436b9b0 · outbound

This paper cites CogVLA: Cognition-aligned vision-language- action models via instruction-driven routing & sparsification.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention CogVLA: Cognition-aligned vision-language- action models via instruction-driven routing & sparsification

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:29.232451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:50:24.986073Z digest=sha256:e3329baf49f5a54d189d9713cdd18ce0a8b44d24f91d34a0efc705330bbfa82e

Observation 74a9e8c4-7ed1-45d4-82aa-6729f3ee0575 · outbound

This paper cites Denoising diffusion probabilistic models.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Denoising diffusion probabilistic models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:29.005648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:50:25.077882Z digest=sha256:d81f7b8d4c8744dcd32eb4e68761fb6ae299916876147b0c84b928e441289a0e

Observation 9fbb998f-301e-4a7d-9d4d-c4c1c02f40df · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action diffusion.The International Journal of Robotics Research, 44(10-11):1684–1704, 2025.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Diffusion policy: Visuomotor policy learning via action diffusion.The International Journal of Robotics Research, 44(10-11):1684–1704, 2025

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:25.175066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:25.175066Z digest=sha256:f1a9a93f403ca273e795176ed0182e735a0ec675fde7c5418f6ce3297481c02d

Observation f7bc1185-a8ee-46ef-b8be-95de4db797dd · outbound

This paper cites Scalable diffusion models with transformers.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Scalable diffusion models with transformers

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:28.771650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:50:25.241353Z digest=sha256:31bee5edf18025727f1a4f63f4007ca05d33cf3be71409efcfa25dc8abdf32a1

Observation 84406f65-ee5a-4195-a3f6-ab8de5a8e57c · outbound

This paper cites an unresolved cited work.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:25.326690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:25.326690Z digest=sha256:c086656565ca7ec8cf132f2b4d37b16233c2b5e20a10fb4ad816d3c8bd5f20ab

Observation db5fcb8d-4137-4673-9b66-8116ba002fb0 · outbound

This paper cites Flow straight and fast: Learning to generate and transfer data with rectified flow.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Flow straight and fast: Learning to generate and transfer data with rectified flow

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:25.380701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:25.380701Z digest=sha256:aa8b0dd5d081da92e27156c56eaa9470b58c10c677caa357368ead8ebd26091a

Observation 73767142-335e-4cc5-ac9a-25f01b4b4263 · outbound

This paper cites RDT-1b: a diffusion foundation model for bimanual manipulation.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention RDT-1b: a diffusion foundation model for bimanual manipulation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:25.431921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:25.431921Z digest=sha256:a91f39856eb916ff8e365e3c2eae3fbcdd5deb8524810b3ac9ed4d7a95db4de3

Observation 4434afed-374d-4c48-9a2a-b4f9a02ec9ad · outbound

This paper cites One-step diffusion policy: Fast visuomotor policies via diffusion distillation.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention One-step diffusion policy: Fast visuomotor policies via diffusion distillation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:25.499650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:25.499650Z digest=sha256:d00b2e0e3db7d3249eae351ed2f88332d6e4623a0d4c4f69dfdb5ddc2f3f12cd

Observation e42e281f-9351-4c78-942a-60da2f88e4e6 · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:25.563377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:25.563377Z digest=sha256:52b8c3dec52ca451b8e0043fb19ddcc014c9c500e26d24ba19937fdee640a8b9

Observation b3445cde-6445-4ff2-8205-2ae91651aa53 · outbound

This paper cites HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:25.631024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:25.631024Z digest=sha256:502cbf4c6b7628138bba813b44d4ad264d746f0508fef1227eec9ca495771f58

Observation c28f8b25-f58c-447a-98a8-e9e4117e7712 · outbound

This paper cites Toward causal representation learning.Proceedings of the IEEE, 109(5):612– 634, 2021.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Toward causal representation learning.Proceedings of the IEEE, 109(5):612– 634, 2021

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:28.480887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:50:25.697551Z digest=sha256:36739b341d2df0f3970f1a48646a4184110eda2686da734db635ca27ca8a4b6f

Observation 4f17e05d-003b-4982-8a36-037f59e12389 · outbound

This paper cites Robust agents learn causal world models.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Robust agents learn causal world models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:28.267867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:50:25.744032Z digest=sha256:58e5ecd209cdda236624d71c65fb49437ec284dc2433cd98f0472f171ee3a166

Observation 7cfb8aa3-cc4b-4258-b3d7-2da483407873 · outbound

This paper cites Causalworld: A robotic manipulation benchmark for causal structure and transfer learning.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Causalworld: A robotic manipulation benchmark for causal structure and transfer learning

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:28.110818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:50:25.794703Z digest=sha256:c6eea9e6789adc4d45be432dd3a6824f6311672173ed14609e1c17b3a520e71b

Observation 3f3d1f1c-b783-4c1d-8853-01868deeac0f · outbound

This paper cites CDP: Towards robust autoregressive visuomotor policy learning via causal diffusion.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention CDP: Towards robust autoregressive visuomotor policy learning via causal diffusion

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:28.007250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:50:25.861570Z digest=sha256:f78cbecf19fcfea830252173a680f38bbbe853fed3fd38f3dd113ecb84a73cb9

Observation a797afcb-a6b0-4019-a748-25670f9c45f8 · outbound

This paper cites When Vision Overrides Language: Evaluating and Mitigating Counterfactual Failures in VLAs.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention When Vision Overrides Language: Evaluating and Mitigating Counterfactual Failures in VLAs

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:25.944143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:25.944143Z digest=sha256:33ad5889c3e1bfc5114cc2048eda449543a382390178a15b82fcbaebc62cc48b

Observation 66296234-8303-4ae9-bf87-c98c28877ba3 · outbound

This paper cites Libero: Benchmarking knowledge transfer for lifelong robot learning.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Libero: Benchmarking knowledge transfer for lifelong robot learning

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:27.941903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:50:26.007215Z digest=sha256:8fd8b4fb4f0d2f09cce1655809745c48ae933074aec1ded06b29ffa85d614709

Observation 691f0a42-a22b-4acd-9884-a271c21c9b58 · outbound

This paper cites DreamVLA: A vision-language-action model dreamed with comprehensive world knowledge.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention DreamVLA: A vision-language-action model dreamed with comprehensive world knowledge

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:27.844071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:50:26.057062Z digest=sha256:d644ede8de4fc3714e30aded0f145c004a27b94ffefbaeab118a7b684727beae

Observation 8ed0afc1-d33b-4a28-8a35-021510895d9a · outbound

This paper cites X-VLA: Soft-prompted transformer as scalable cross-embodiment vision-language-action model.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention X-VLA: Soft-prompted transformer as scalable cross-embodiment vision-language-action model

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:27.648242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:50:26.128371Z digest=sha256:062d045925d9df606ced170a9d4df7d1e9eff221292154f00a235768bfeaf1ba

Observation 08851ab8-5ee7-4209-ad8d-2bbaed0d8fd4 · outbound

This paper cites NORA: A Small Open-Sourced Generalist Vision Language Action Model for Embodied Tasks.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention NORA: A Small Open-Sourced Generalist Vision Language Action Model for Embodied Tasks

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:26.194504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:26.194504Z digest=sha256:9e4219c9d4cfd28abead2c75fca4062e9431e62ac85ff859e8a91d7e31be50cf

Observation f043e5b6-9f45-40ab-87a8-37eee2e7b6d7 · outbound

This paper cites WorldVLA: Towards Autoregressive Action World Model.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention WorldVLA: Towards Autoregressive Action World Model

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:26.254317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:26.254317Z digest=sha256:502b409d287195dc5a2ff0bdc983fd58d61405ae8e89fb99daeb81cd88d7063c

Observation 529042c5-7ed2-4491-881c-669a61b1ffc2 · outbound

This paper cites UniVLA: Learning to Act Anywhere with Task-centric Latent Actions.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention UniVLA: Learning to Act Anywhere with Task-centric Latent Actions

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:26.320751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:26.320751Z digest=sha256:d06cdf14c5ca4756d2f9e60460269a7b2e483920ecc52f305a1ea280515a6c26

Observation 8976f62e-7d14-4537-a903-38dfc0e62830 · outbound

This paper cites FAST: Efficient Action Tokenization for Vision-Language-Action Models.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:26.384984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:26.384984Z digest=sha256:e83a08b8bfc708807d6c6db9e19d7869528302a54c69b68a572266d118aba911

Observation 45f4e954-eadb-44ce-b3d2-cce441c7b7ec · outbound

This paper cites Chance- constrained flow matching for high-fidelity constraint-aware generation.arXiv preprint arXiv:2509.25157, 2025.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention Chance- constrained flow matching for high-fidelity constraint-aware generation.arXiv preprint arXiv:2509.25157, 2025

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:26.448351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:26.448351Z digest=sha256:0e97c331e984419745fbf1d9e78ed03710f76f1bfede05627fe3081c56eb9991

Observation 3744e82a-5315-4c4b-bc5a-918182fb1a6b · outbound

This paper cites DDAT: Diffusion Policies Enforcing Dynamically Admissible Robot Trajectories.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention DDAT: Diffusion Policies Enforcing Dynamically Admissible Robot Trajectories

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:26.507752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:26.507752Z digest=sha256:e1a78984bb7914750c5c18102a195245f2b18da74359aec7c3a949398840d62a

Observation 0c11d84a-a8ba-4999-a779-48b54ea3248d · outbound

This paper cites LeRobot: An open- source library for end-to-end robot learning.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention LeRobot: An open- source library for end-to-end robot learning

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:27.475474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:50:26.576375Z digest=sha256:9b34c55264bddcf7e51222b943e66202a96700755f0b8a70b519e0ca3081bb8b

Observation 485b8b7b-004a-4ad0-b326-07755b41e02b · outbound

This paper cites LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:26.636070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:26.636070Z digest=sha256:ede3087d7140528c6cea5a26a12a4b182d2a2aa961cf5e10d924c9261c2aa642

Observation 5956c433-731f-4b2f-bdaa-dc692277a9d3 · outbound

This paper cites remove the cuboid from the blue plate.

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention remove the cuboid from the blue plate

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:27.296826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:50:26.697393Z digest=sha256:d599f1d840a0b3de6d2219b65d165de02b5215966b7d5d3f9cd371d4a771427b

Pith citing papers

No inbound Pith citation observations are available.