Pith. sign in

Paper Citation Record · LEDGER

Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction

As of 23 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 0 inbound Pith citation observations for arXiv:2505.24156.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.24156 v1

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:44:08.215410Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

59 of 59 outbound references displayed

  • verified exact0
  • verified fuzzy24
  • unresolved35
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2fac1c56-cc2f-4620-a7b0-064742dec1e4 · outbound

This paper cites Peract2: Benchmarking and learning for robotic bimanual manipulation tasks.

Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction Peract2: Benchmarking and learning for robotic bimanual manipulation tasks

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:03.876268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:03.876268Z digest=sha256:ad56b649d32f00518654ba386ec292b2a3b13655c51f8641cc86293a53611965

Observation bcaf170e-95de-4a17-86b4-7433c74a3e05 · outbound

This paper cites RoboTwin: Dual-Arm Robot Benchmark with Generative Digital Twins (early version).

Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction RoboTwin: Dual-Arm Robot Benchmark with Generative Digital Twins (early version)

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:03.940078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:03.940078Z digest=sha256:072556f8563785da998e971b7dd1f928c818765d47320830573a2d0b23d93ef3

Observation a31546b0-097f-4998-86f7-64d31e6aeabb · outbound

This paper cites Zhao, Vikash Kumar, Sergey Levine, and Chelsea Finn.

Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction Zhao, Vikash Kumar, Sergey Levine, and Chelsea Finn

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:44:12.780640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:44:04.004285Z digest=sha256:fe557a2a8130ddc1e9f5ad4529325de3d82e88ff98d71f53c7025bd1177a0ab9

Observation 54af4f5a-da69-45e5-b09f-8b5933de8208 · outbound

This paper cites Sukhatme.

Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction Sukhatme

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:44:12.551504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:44:04.057663Z digest=sha256:43269001e354a66a933397c94fc6bd6cb2b55b8cdf5453607474f50df2ca50ec

Observation 56c91c17-8459-4094-8113-bb94cca24cfb · outbound

This paper cites Stabilize to act: Learning to coordinate for bimanual manipulation.

Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction Stabilize to act: Learning to coordinate for bimanual manipulation

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:44:12.422940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:44:04.088995Z digest=sha256:61c14df5e3ad090e27e01a70fe233f42b1e812ed63b450a66d575c0926800ff8

Observation 0430fe08-44b4-4835-aae1-d6576339ab25 · outbound

This paper cites Taco: Benchmarking generalizable bimanual tool-action-object understanding.

Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction Taco: Benchmarking generalizable bimanual tool-action-object understanding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:04.247130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:04.247130Z digest=sha256:3ce40d9340d941171e61de816557c0cc12c4855a6046c07a148f18f5aa8e7316

Observation e137aa09-a528-4cbf-9254-a739d9d3dd01 · outbound

This paper cites Towards human-level bimanual dexterous manipulation with reinforcement learning.Advances in Neural Information Processing Systems, 35:5150–5163, 2022.

Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction Towards human-level bimanual dexterous manipulation with reinforcement learning.Advances in Neural Information Processing Systems, 35:5150–5163, 2022

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:44:11.938001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:44:04.288605Z digest=sha256:8b5f5384aefeb9bd6dd0a04f489cb483fbc8689ef8569d36e0c6aacffc91908a

Observation 1e851552-939d-4968-a1d5-99057804227b · outbound

This paper cites Bi-touch: Bimanual tactile manipulation with sim-to-real deep reinforcement learning.IEEE Robotics and Automation Letters, 8(9):5472–5479, 2023.

Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction Bi-touch: Bimanual tactile manipulation with sim-to-real deep reinforcement learning.IEEE Robotics and Automation Letters, 8(9):5472–5479, 2023

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:44:11.772299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:44:04.377233Z digest=sha256:92b83765b07e555fc893db154475e33d85f6bf999a87953ee7ecf74ddb1b9dba

Observation b00183a8-b25b-4217-a162-3b8d970974bf · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction OpenVLA: An Open-Source Vision-Language-Action Model

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:04.482160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:04.482160Z digest=sha256:e41d0502254c32f096a95f3cde7e9ac2f6b38d5925c2b3788bf4c11c47f29782

Observation d6a663d2-4eb8-4385-a608-6c619c808dff · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:04.565880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:04.565880Z digest=sha256:677b9b50a0c72011534e2e2c8eee1ad756a1c460ea1887e233f0e3a0abf8a162

Observation 0408ea97-8c40-43cb-9232-0a03ef86a20e · outbound

This paper cites RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation.

Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:04.614233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:04.614233Z digest=sha256:fd6037f8f4c717ccb22515ff2430d2e1073cb2ee00f3b22b1194ce8ada6ac11f

Observation 6a5143d7-b632-43db-b835-ab8acb693f40 · outbound

This paper cites GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.

Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:04.662193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:04.662193Z digest=sha256:ef27a9b2d769ea0ef954c1db0bc415ffabf7f02ce2541d83573d74556daece96

Observation 7d75ce22-ca01-4944-ac0b-53e1aca7ee89 · outbound

This paper cites AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems.

Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:04.720791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:04.720791Z digest=sha256:8480d497f3f257e46c26924a9620aaefe2c1f1d2c3d40dc9b09af1c895ff44f0

Observation c4286d83-0b49-482e-8780-22cbafe551a9 · outbound

This paper cites Cogvideox: Text-to-video diffusion models with an expert transformer.

Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction Cogvideox: Text-to-video diffusion models with an expert transformer

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:44:11.619306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:44:04.806817Z digest=sha256:e8cd2cdec55a2fc412baaf09479c593981efff060862e1e9cc6aa2087bcc036c

Observation 6073b6e8-ad1b-42a6-a8d4-eb0e1cf6aa34 · outbound

This paper cites Open-television: Teleoperation with immersive active visual feedback.

Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction Open-television: Teleoperation with immersive active visual feedback

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:44:11.489368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:44:04.848490Z digest=sha256:5cc20150b0a2717629e1dfdcf08eb496c2bfacd0068cd7a2c25df7a614b945c1

Observation 0bf92015-6fa3-4b81-869b-0b24459e896e · outbound

This paper cites RoboMIND: Benchmark on Multi-embodiment Intelligence Normative Data for Robot Manipulation.

Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction RoboMIND: Benchmark on Multi-embodiment Intelligence Normative Data for Robot Manipulation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:04.884308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:04.884308Z digest=sha256:328a9557df106d6bf228b2b4736279cfd851aff7b3c1783127c63ba3097a34e2

Observation eb0a3954-2c33-4201-87a8-179a45056f8c · outbound

This paper cites Motion-i2v: Consistent and controllable image-to-video generation with explicit motion modeling.

Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction Motion-i2v: Consistent and controllable image-to-video generation with explicit motion modeling

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:04.940244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:04.940244Z digest=sha256:bcd05c5426049a9e54cf07fb67898f420acfb092aebd985956011302fbddd95e

Observation 61677f15-766e-43c4-ad4d-81e8968df44c · outbound

This paper cites Raft: Recurrent all-pairs field transforms for optical flow.

Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction Raft: Recurrent all-pairs field transforms for optical flow

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:05.020864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:05.020864Z digest=sha256:2543e4a29bc8f5ce8940e74ab8c20efdef485f9183de34281878e6d05a1354ec

Observation cd00a0c8-fdf3-41fc-990d-0c67755ed727 · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action diffusion.The International Journal of Robotics Research, page 02783649241273668, 2023.

Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction Diffusion policy: Visuomotor policy learning via action diffusion.The International Journal of Robotics Research, page 02783649241273668, 2023

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:05.098999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:05.098999Z digest=sha256:8841772f351ea3c61c64224e1da2b545556cb4788eee767c451c54f3c8fe5b36

Observation 2eccbf07-e555-4e86-97bf-6d047b2b7b58 · outbound

This paper cites Learning manipulation by sequencing motor primitives with a two-armed robot.

Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction Learning manipulation by sequencing motor primitives with a two-armed robot

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:44:11.373458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:44:05.146131Z digest=sha256:9556530ff7b0594031951a4144644abc8ddab78321ae442be1d8803444e145a1

Observation f3228e6d-c88d-4ca3-8bda-a61f047d91c5 · outbound

This paper cites A system for imitation learning of contact-rich bimanual manipulation policies.

Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction A system for imitation learning of contact-rich bimanual manipulation policies

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:44:11.239931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:44:05.197668Z digest=sha256:21424f265490ecdf8c033e822a6923f7a577d4fafc0c3cf52766a57bc4c328b7

Observation 5d8336de-a3bf-4548-87df-43ea6ebf4ae6 · outbound

This paper cites Deep imitation learning for bimanual robotic manipulation.Advances in neural information processing systems, 33:2327–2337, 2020.

Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction Deep imitation learning for bimanual robotic manipulation.Advances in neural information processing systems, 33:2327–2337, 2020

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:44:11.068028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:44:05.285339Z digest=sha256:4849d74ae5ffb3b655e1738a7a38149838aecf9c05f1f133553d220d8cbbf211

Observation 15afa95b-f4d9-4901-969a-88adc9cd4342 · outbound

This paper cites HumanoidBench: Simulated Humanoid Benchmark for Whole-Body Locomotion and Manipulation.

Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction HumanoidBench: Simulated Humanoid Benchmark for Whole-Body Locomotion and Manipulation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:05.423644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:05.423644Z digest=sha256:6600b7942755ebd718d0acddb7607000102f4f5eaaf461814f6f235c22a30063

Observation a5e7fb93-c5d5-4fa0-b5da-c3a0372b47c6 · outbound

This paper cites Bigym: A demo-driven mobile bi-manual manipulation benchmark.

Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction Bigym: A demo-driven mobile bi-manual manipulation benchmark

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:44:10.942838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:44:05.507727Z digest=sha256:9e4b8f84c599498565f3453ff84ea6ecff81d3909eeccdfc28a9c7529beef1b8

Observation 1c11c2fe-234f-4b6c-8226-0a098b1da6b1 · outbound

This paper cites Zhao, and Chelsea Finn.

Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction Zhao, and Chelsea Finn

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:44:10.670542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:44:05.693363Z digest=sha256:e309c585a1674cab5a2a590199796f40e4cf861919aa212d728d8c681fb61ea0

Observation 167d60bd-69cf-4f17-82e8-9cd07240c600 · outbound

This paper cites Mimicgen: A data generation system for scalable robot learning using human demonstrations.

Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction Mimicgen: A data generation system for scalable robot learning using human demonstrations

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:44:10.514527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:44:05.753884Z digest=sha256:964b0b539f9698ea68946a4b8732761738f776039fa6a743d7b4e3b02769e4e8

Observation d272dfd6-bcb4-4c07-a694-0a90644aaa95 · outbound

This paper cites Dexmimicgen: Automated data generation for bimanual dexterous manipulation via imitation learning.

Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction Dexmimicgen: Automated data generation for bimanual dexterous manipulation via imitation learning

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:44:10.167669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:44:05.795213Z digest=sha256:1604616a0fdb69e3a8ae0e41cff763fc33317c41d36021f16d836c1bc41d2103

Observation 1eb90710-51fb-4167-9347-2ce984bc88d0 · outbound

This paper cites Bi-kvil: Keypoints-based visual imitation learning of bimanual manipulation tasks.

Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction Bi-kvil: Keypoints-based visual imitation learning of bimanual manipulation tasks

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:44:12.095500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:44:05.857117Z digest=sha256:628dd8f8b96147f575bb08154c3a25c3c9c5d0d22690be6cabebc39850b36670

Observation 1dd43b64-256b-4be9-9f06-9c35b259422b · outbound

This paper cites Interactive imitation learning of bimanual movement primitives.IEEE/ASME Transactions on Mechatronics, 2023.

Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction Interactive imitation learning of bimanual movement primitives.IEEE/ASME Transactions on Mechatronics, 2023

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:44:12.265045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:44:05.939426Z digest=sha256:843143ade84ec282d450864dd1a5c4ea11e4e7a5ec2ee2ebaa0dad887e9a92a9

Observation 60c6a14d-e412-4b26-9edf-55b1e1b1aa56 · outbound

This paper cites InterACT: Inter-dependency aware action chunking with hierarchical attention transformers for bimanual manipulation.

Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction InterACT: Inter-dependency aware action chunking with hierarchical attention transformers for bimanual manipulation

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:44:09.833950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:44:05.991148Z digest=sha256:8cc078cf6eaca04d9068e08ddf997591654d851e15ece23c06fbd6e0e93ba9e3

Observation c28b8204-994e-4025-a49b-910d8f7ab99d · outbound

This paper cites Octo: An Open-Source Generalist Robot Policy.

Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction Octo: An Open-Source Generalist Robot Policy

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:06.059408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:06.059408Z digest=sha256:12988bd775b129bb146a2ec8e1ba3d72d7b3e43e73e1701eefb5f5eb97ad3498

Observation 5b935397-c482-4cc6-9d71-910456882a75 · outbound

This paper cites SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model.

Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:06.124651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:06.124651Z digest=sha256:a13f5a83a641502d91c1473cecafaae377fe78c4e4f0d41d47f1cde80ab061b5

Observation 231118b3-8f9b-4594-a547-cde708953bfe · outbound

This paper cites CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation.

Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:06.193883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:06.193883Z digest=sha256:d7fe91a9dd7f6ee6813d40d6b819946c12d60dc68b775bea0ee8cdcfe0fe95c1

Observation 3e83782f-60aa-4ab7-a5c5-87baef0780f6 · outbound

This paper cites Dexgraspvla: A vision-language-action framework towards general dexterous grasping.arXiv preprint arXiv:2502.20900, 2025.

Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction Dexgraspvla: A vision-language-action framework towards general dexterous grasping.arXiv preprint arXiv:2502.20900, 2025

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:06.269803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:06.269803Z digest=sha256:45ec7aecac43d0639a0b6a6f4a24e16af0482a2e6d4aefe399fd06d1d58a3cf9

Observation 6f11c451-0ad3-4ec4-acac-2e7d5d3fa5d5 · outbound

This paper cites MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation.

Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:06.343233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:06.343233Z digest=sha256:9971523ae6c8e576e86bf365a1f7239deed0268ea5a9bd8a6add5da76e1dc563

Observation 18a086b1-d1e2-4298-95f8-2b4e561103c0 · outbound

This paper cites Learning an actionable dis- crete diffusion policy via large-scale actionless video pre-training.

Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction Learning an actionable dis- crete diffusion policy via large-scale actionless video pre-training

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:06.440163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:06.440163Z digest=sha256:e3dfbcaa12d0a45d3f2f954a10482b10bf9d16f8c6d0ba2f08e6d965fc957f14

Observation bee84a65-eff2-407f-86dd-eab3fd439225 · outbound

This paper cites Video diffusion models.Advances in Neural Information Processing Systems, 35:8633–8646, 2022.

Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction Video diffusion models.Advances in Neural Information Processing Systems, 35:8633–8646, 2022

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:06.539740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:06.539740Z digest=sha256:bfb0ab6d4e053aaa19f4641b04aa8d0952c286e098e5ff9334442500752e9fc1

Observation 5bcbadb6-88c3-414d-8bb9-ca315084392a · outbound

This paper cites CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers.

Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:06.609040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:06.609040Z digest=sha256:92fa717884a0b37709e7730b74d4b6473da251c475c2b2ed4ac4eeccfe5f3bea

Observation da57f813-7684-47be-ac5b-b5790e4c2d0f · outbound

This paper cites Learning universal policies via text-guided video generation.Advances in neural information processing systems, 36:9156–9172, 2023.

Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction Learning universal policies via text-guided video generation.Advances in neural information processing systems, 36:9156–9172, 2023

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:06.677872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:06.677872Z digest=sha256:f55701ee735a04f93475831be29a7d827f3bd2d541c00e2ae2ddb6d24c409391

Observation 5c371ed9-8661-4b74-bd73-2598e2f92313 · outbound

This paper cites Tenenbaum.

Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction Tenenbaum

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:44:09.653079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:44:06.744303Z digest=sha256:98c2787a055f4c807fbff5f5e6798a4a0a696a99f549c0557842dcbbbd22fd29

Observation d116d889-5c4c-4831-b474-abd130333a4f · outbound

This paper cites Grounding video models to actions through goal conditioned exploration.

Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction Grounding video models to actions through goal conditioned exploration

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:44:09.592222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:44:06.847418Z digest=sha256:4fc725e670373566b87015e200e6d83075623657b43c98682a9ea6fe0b40b409

Observation 34c42fec-0338-40a2-9cf2-cdec63d56755 · outbound

This paper cites Robodreamer: Learning compositional world models for robot imagination.

Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction Robodreamer: Learning compositional world models for robot imagination

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:44:09.299336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:44:06.922885Z digest=sha256:72ad3a5c57fde6ef7e6536e6558800210f2249d12979b2e03e695ced11213aef

Observation 3cefc23a-d791-41c4-a1b5-45c13c0bb7ff · outbound

This paper cites Vidman: Exploiting implicit dynamics from video diffusion model for effective robot manipulation.Advances in Neural Information Processing Systems, 37:41051–41075, 2024.

Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction Vidman: Exploiting implicit dynamics from video diffusion model for effective robot manipulation.Advances in Neural Information Processing Systems, 37:41051–41075, 2024

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:44:09.148044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:44:07.016998Z digest=sha256:016d6d5cd6ecefb88e0812480ddc811a58ca37aa7f91f42ccc5fe85811fb3e99

Observation 3d7bb084-2812-4bb6-9c60-6fdca5151224 · outbound

This paper cites Open-Sora: Democratizing Efficient Video Production for All.

Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction Open-Sora: Democratizing Efficient Video Production for All

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:07.109191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:07.109191Z digest=sha256:b62fa01ef5f3ca153317e6bbbcec112a48aee167609ba8edad949c224d0cfdd2

Observation 2768bf18-0f9c-4e5f-8f63-771cb17a482c · outbound

This paper cites Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0.

Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:07.210196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:07.210196Z digest=sha256:3d18d84ca6e20adc69b797c5275b9e9a037135bdceb0da82303b4c59fefccf3b

Observation d80fbe2f-7cd1-40ee-9807-32de246ad89f · outbound

This paper cites IRASim: A Fine-Grained World Model for Robot Manipulation.

Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction IRASim: A Fine-Grained World Model for Robot Manipulation

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:07.307769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:07.307769Z digest=sha256:036fec2c0a864246bd5afe28b00e44bc98d7a1a9c33f8aa3b9d691e07d5f5d06

Observation 34e5f4aa-5e0c-4837-8273-c8eea0c9a849 · outbound

This paper cites AVID: Adapting Video Diffusion Models to World Models.

Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction AVID: Adapting Video Diffusion Models to World Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:07.372300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:07.372300Z digest=sha256:1dae707fcc35d4b4118244a63659a39b5f0f648b89cb4808e4d67258541494ab

Observation e1c12c8f-e986-43f3-8d92-73890d0503f6 · outbound

This paper cites AdaWM: Adaptive world model based planning for autonomous driving.

Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction AdaWM: Adaptive world model based planning for autonomous driving

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:44:09.049168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:44:07.423391Z digest=sha256:602819e1a09de5a63c3a0cbb9fe9532d5934cf664bc84515eb433d515c8755b3

Observation 8d20f82c-8c48-4edd-a587-1ff4ce3d668d · outbound

This paper cites Flowformer++: Masked cost volume autoencoding for pretraining optical flow estimation.

Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction Flowformer++: Masked cost volume autoencoding for pretraining optical flow estimation

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:44:08.813360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:44:07.497484Z digest=sha256:9be6dee6148b72f41188e61f8731c5967530f2a9066ddaf0573314f2b51434bf

Observation 650addb1-45ad-418b-9f65-ca3a1138b024 · outbound

This paper cites 3D Diffusion Policy: Generalizable Visuomotor Policy Learning via Simple 3D Representations.

Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction 3D Diffusion Policy: Generalizable Visuomotor Policy Learning via Simple 3D Representations

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:07.559077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:07.559077Z digest=sha256:ced91f19dadc3000c925128bff9eb82b4ac3f952703aa27cd54a1acd7581a700

Observation 3a264020-d2e0-4d8d-a3a7-a8611b32084e · outbound

This paper cites Image quality metrics: Psnr vs.

Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction Image quality metrics: Psnr vs

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:07.632798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:07.632798Z digest=sha256:5e0dbc983873ca571f7f5b0c775c8ae50a705ed72986059a814df30032e6dbd0

Observation a590c559-8eca-4f32-a6a0-249a978f4b86 · outbound

This paper cites Image quality assessment: from error visibility to structural similarity.IEEE transactions on image processing, 13(4):600–612, 2004.

Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction Image quality assessment: from error visibility to structural similarity.IEEE transactions on image processing, 13(4):600–612, 2004

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:07.718410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:07.718410Z digest=sha256:926da6e5fb8a62c9d48808cc968dc26abba6530249f48e98e400d0b171b171f2

Observation 031ffeed-379c-4cf0-b6be-1a3a8a29b77e · outbound

This paper cites The unreasonable effectiveness of deep features as a perceptual metric.

Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction The unreasonable effectiveness of deep features as a perceptual metric

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:07.790593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:07.790593Z digest=sha256:5d10c2a57d337d4e92c5ad318a311ebb6c490ac94e7ab5265a2a6a97d1824eb8

Observation 7927f42b-564d-48ce-9124-096462d42db5 · outbound

This paper cites Towards Accurate Generative Models of Video: A New Metric & Challenges.

Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction Towards Accurate Generative Models of Video: A New Metric & Challenges

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:07.872644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:07.872644Z digest=sha256:281554d961f70d3b6121f7cc6223ef041f6afafe701933082f14c5f484929485

Observation 46889b61-55be-4ffa-b472-269bd1302727 · outbound

This paper cites Bunny-VisionPro: Real-Time Bimanual Dexterous Teleoperation for Imitation Learning.

Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction Bunny-VisionPro: Real-Time Bimanual Dexterous Teleoperation for Imitation Learning

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:07.953294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:07.953294Z digest=sha256:da46f3a472ca5e14b786914b7a20c01f6c32612572a66aeec990b05d794a0c87

Observation 1cd04788-5a2c-48be-ba50-222e4c3e00a0 · outbound

This paper cites OmniH2O: Universal and Dexterous Human-to-Humanoid Whole-Body Teleoperation and Learning.

Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction OmniH2O: Universal and Dexterous Human-to-Humanoid Whole-Body Teleoperation and Learning

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:08.033340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:08.033340Z digest=sha256:14cbdc497bc627922d45180bc9bced2c3dbc32cc8c2786adf432fd9bd9d992ee

Observation 179cb5ee-4512-475f-b9e9-989c0d8982f7 · outbound

This paper cites Using apple vision pro to train and control robots, 2024.

Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction Using apple vision pro to train and control robots, 2024

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:08.139200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:08.139200Z digest=sha256:f77393207afc24b50c5306dfc38c0b8da0ec8036f11abc4afc5480b62e9af3d2

Observation ac0e4fe9-bc07-4976-9e0c-bfabf194b869 · outbound

This paper cites Grasp the rope on the box and pull together to bring the box closer.

Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction Grasp the rope on the box and pull together to bring the box closer

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:44:08.641167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:44:08.215410Z digest=sha256:a1911c4083556e6ffbe426afad3eef0c5453cd3a8a72965f7b9539b906f1c024

Observation 2b3325ee-2e6d-4d9a-9643-4f982f124578 · outbound

This paper cites an unresolved cited work.

Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction Unresolved cited work

Reference 2024

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:44:10.796678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:44:05.600471Z digest=sha256:4a31c90fb2b148c8477883375a28d1f4e75fd908fb4b33d43aa9c3435ed685cc

Pith citing papers

No inbound Pith citation observations are available.