Pith. sign in

Paper Citation Record · LEDGER

Motion-Focused Latent Action Enables Cross-Embodiment VLA Training from Human EgoVideos

As of 6 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 0 inbound Pith citation observations for arXiv:2606.18955.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.18955 v2

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-03T23:44:35.238493Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

30 of 30 outbound references displayed

  • verified exact20
  • verified fuzzy10
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bb342532-56d7-42a3-b206-d119dd4d806c · outbound

This paper cites Rt-2: Vision-language-action models transfer web knowledge to robotic control.

Motion-Focused Latent Action Enables Cross-Embodiment VLA Training from Human EgoVideos Rt-2: Vision-language-action models transfer web knowledge to robotic control

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T23:40:13.230888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T23:44:35.238493Z digest=sha256:987abf11b0488596fbada4b0cdb9cdb546ae3c642881a75ae2cc5f11e46b42b4

Observation 5a28e0a3-7ee6-4467-8b14-3e60063ec3b3 · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

Motion-Focused Latent Action Enables Cross-Embodiment VLA Training from Human EgoVideos $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-03T23:49:01.889233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T23:44:35.238493Z digest=sha256:f9bc1d7913eb89af7c98fe37aef408b17c891080802665fd1b1f2ab3844b751f

Observation 8aca3682-2511-48d3-a305-abda9e933487 · outbound

This paper cites RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation.

Motion-Focused Latent Action Enables Cross-Embodiment VLA Training from Human EgoVideos RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-03T23:49:01.874547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T23:44:35.238493Z digest=sha256:a531d9a041171377c75ec999f9a3e2606c82f5d3c97315a565bc8086813b0819

Observation 96d87619-2dc6-41bf-8509-7a1e8fdfd306 · outbound

This paper cites Open x-embodiment: Robotic learning datasets and rt-x models: Open x- embodiment collaboration 0.

Motion-Focused Latent Action Enables Cross-Embodiment VLA Training from Human EgoVideos Open x-embodiment: Robotic learning datasets and rt-x models: Open x- embodiment collaboration 0

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T23:40:13.241648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T23:44:35.238493Z digest=sha256:5f431eb9134290a78e1c8cc06d5fb45458ea8b8533b7a796a457bc9d3590f52c

Observation 962c5af2-eb31-4847-a09f-fc5f0761e44e · outbound

This paper cites AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems.

Motion-Focused Latent Action Enables Cross-Embodiment VLA Training from Human EgoVideos AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-03T23:49:01.896471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T23:44:35.238493Z digest=sha256:6d976e8e077ba38021765f9508a31f273e84c1824e82d64a0a36bfd9a6424dfe

Observation 25881b6c-f5df-4391-9789-9d9294e7237a · outbound

This paper cites Egomimic: Scaling imitation learning via egocentric video.

Motion-Focused Latent Action Enables Cross-Embodiment VLA Training from Human EgoVideos Egomimic: Scaling imitation learning via egocentric video

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T23:40:13.239467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T23:44:35.238493Z digest=sha256:fa604c5a0cb6b51b772abcb8961240eb5d4db4ce5a1de77bb63b85e37fcbd0e6

Observation baec75bf-216e-4c7f-a54a-623992b13991 · outbound

This paper cites Motiontrans: Human vr data enable motion-level learning for robotic manipulation policies.

Motion-Focused Latent Action Enables Cross-Embodiment VLA Training from Human EgoVideos Motiontrans: Human vr data enable motion-level learning for robotic manipulation policies

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-03T23:49:01.894675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T23:44:35.238493Z digest=sha256:f5192bcfe004ba64d25a386c8ab3506a69a35011318d98a33f5980e0210ead00

Observation 68c83ee3-f061-42eb-b622-b81583aae9fa · outbound

This paper cites H-RDT: Human Manipulation Enhanced Bimanual Robotic Manipulation.

Motion-Focused Latent Action Enables Cross-Embodiment VLA Training from Human EgoVideos H-RDT: Human Manipulation Enhanced Bimanual Robotic Manipulation

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-03T23:49:01.877582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T23:44:35.238493Z digest=sha256:e8f29c893440f4006fde82e1f3109d9f0af5b144f904154c82b132b3e928c0be

Observation f84b4399-2964-4786-994f-8c75f4e7438c · outbound

This paper cites Latent Action Pretraining from Videos.

Motion-Focused Latent Action Enables Cross-Embodiment VLA Training from Human EgoVideos Latent Action Pretraining from Videos

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-03T23:49:01.879308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T23:44:35.238493Z digest=sha256:73cca7042ad1ac73eec8c1203d401010354e8350dc19ff706cd2e20bd0cec78b

Observation f2cded77-958a-4ab6-a065-0eeb889d2748 · outbound

This paper cites IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AI.

Motion-Focused Latent Action Enables Cross-Embodiment VLA Training from Human EgoVideos IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AI

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-03T23:49:01.900671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T23:44:35.238493Z digest=sha256:51ef12935127ed93cfa7d9b751dd686d7f1d0af1964aa03e0c67312c84b352ef

Observation 2423e44a-e6ba-4d25-8d5f-08ed83fbd2f2 · outbound

This paper cites What do latent action models actually learn?.

Motion-Focused Latent Action Enables Cross-Embodiment VLA Training from Human EgoVideos What do latent action models actually learn?

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-03T23:49:01.871582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T23:44:35.238493Z digest=sha256:8fd182729161b3f8251b798d4a68129645f6d02dcf297d181886d303208fb8b0

Observation f51869ce-30bc-4a65-a8e3-86b2453ddd58 · outbound

This paper cites UniVLA: Learning to Act Anywhere with Task-centric Latent Actions.

Motion-Focused Latent Action Enables Cross-Embodiment VLA Training from Human EgoVideos UniVLA: Learning to Act Anywhere with Task-centric Latent Actions

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-03T23:49:01.881847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T23:44:35.238493Z digest=sha256:d88710f8fc78ebf466799bf3fcca2bfbecab05531ff7eccfcdba79705eacb033

Observation 579d7492-c339-49a3-9486-d4cdcdfbede7 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

Motion-Focused Latent Action Enables Cross-Embodiment VLA Training from Human EgoVideos OpenVLA: An Open-Source Vision-Language-Action Model

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-03T23:49:01.899325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T23:44:35.238493Z digest=sha256:9f5ef8e09db244a38698aff9dc2a61d2797e9827ac94b6b64034c9900062c99c

Observation d047ecf8-978d-4991-9db0-dc9d176960d9 · outbound

This paper cites GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.

Motion-Focused Latent Action Enables Cross-Embodiment VLA Training from Human EgoVideos GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-03T23:49:01.903760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T23:44:35.238493Z digest=sha256:8a8519c107ac1ac78bc5419ec9a430212fcba781b6edae8963fda6ccd78138f8

Observation 6133d7c9-625a-43c0-b8e5-61266fa3b3fe · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action diffusion.

Motion-Focused Latent Action Enables Cross-Embodiment VLA Training from Human EgoVideos Diffusion policy: Visuomotor policy learning via action diffusion

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T23:40:13.246257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T23:44:35.238493Z digest=sha256:834b1f1cfce43d6d6b37a48f3a30e23dc393448d4386f40caa7b43d5daf548ba

Observation 824544dd-455b-4142-a12a-5a01ccb88480 · outbound

This paper cites Flow Matching for Generative Modeling.

Motion-Focused Latent Action Enables Cross-Embodiment VLA Training from Human EgoVideos Flow Matching for Generative Modeling

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-03T23:49:01.889027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T23:44:35.238493Z digest=sha256:81516e8d89d3207170825a13c62b76d9a2d617f9b3e20717ea22d35b11b0fcb1

Observation b374a28c-3af3-4724-8064-9ad02607aed8 · outbound

This paper cites villa-X: Enhancing Latent Action Modeling in Vision-Language-Action Models.

Motion-Focused Latent Action Enables Cross-Embodiment VLA Training from Human EgoVideos villa-X: Enhancing Latent Action Modeling in Vision-Language-Action Models

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-07-03T23:49:01.891429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T23:44:35.238493Z digest=sha256:1c53ecb0bd6df66df4dd8ff74a94ded05ccd7bb4faded44b4b9c4d0769683183

Observation e831ddd6-8a35-48ae-a903-dd2dfc103e04 · outbound

This paper cites Universal manipulation interface: In-the-wild robot teaching without in-the-wild robots.

Motion-Focused Latent Action Enables Cross-Embodiment VLA Training from Human EgoVideos Universal manipulation interface: In-the-wild robot teaching without in-the-wild robots

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T23:40:13.235889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T23:44:35.238493Z digest=sha256:aaae2b0d69ace50d1a5e5872942a5d9275919bc89a22e298c488a1d8fee1bc61

Observation f5ba9d0b-e4cb-49be-8792-dc6820f2e1db · outbound

This paper cites Segment Anything.

Motion-Focused Latent Action Enables Cross-Embodiment VLA Training from Human EgoVideos Segment Anything

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-07-03T23:49:01.886677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T23:44:35.238493Z digest=sha256:4a5487d4a4d591a8f42d8ce88568a8021cb0bbe9dc8b5fd0c97a6fbc39ea2921

Observation 229b6af9-98a7-4b08-a783-682d4c61b499 · outbound

This paper cites Prismatic vlms: Investigating the design space of visually- conditioned language models,.

Motion-Focused Latent Action Enables Cross-Embodiment VLA Training from Human EgoVideos Prismatic vlms: Investigating the design space of visually- conditioned language models,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T23:40:13.228346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T23:44:35.238493Z digest=sha256:937a398fcb7c1cc54f8073d7b245018d9a3a9707d10a535fef05fd35f489bdc8

Observation a65ac982-f3e9-490d-b482-7be63559fad7 · outbound

This paper cites SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model.

Motion-Focused Latent Action Enables Cross-Embodiment VLA Training from Human EgoVideos SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-07-03T23:49:01.897292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T23:44:35.238493Z digest=sha256:e20684268b778b83ea88097591f0940c30fc807ff66890dd13fd09aaaba5ada5

Observation 04ff4257-5693-487c-9e42-ed5017a8f3ea · outbound

This paper cites FAST: Efficient Action Tokenization for Vision-Language-Action Models.

Motion-Focused Latent Action Enables Cross-Embodiment VLA Training from Human EgoVideos FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-07-03T23:49:01.883254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T23:44:35.238493Z digest=sha256:a132f0d1d17f43d30618bc8ecb817492f02a8c11d8ec83c19e147f3f15a1f05c

Observation 975188d9-b63c-4a61-922e-67ceed4c0b18 · outbound

This paper cites Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware.

Motion-Focused Latent Action Enables Cross-Embodiment VLA Training from Human EgoVideos Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-07-03T23:49:01.893940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T23:44:35.238493Z digest=sha256:2f7245f342b7051b58f01fd79bbf76a076402e3181f3e3b5b9efa90a4a2c607d

Observation b4e3148a-1276-46aa-b67a-925586bfd33d · outbound

This paper cites Bridgedata v2: A dataset for robot learning at scale,.

Motion-Focused Latent Action Enables Cross-Embodiment VLA Training from Human EgoVideos Bridgedata v2: A dataset for robot learning at scale,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T23:40:13.237616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T23:44:35.238493Z digest=sha256:aa1dffc7369f62d738dbf9216ac9c6933ba1e8783528bf44f35aa47f87bb94fe

Observation ae6cbf74-cd35-43c2-aa91-0a3697f67341 · outbound

This paper cites Libero: Benchmarking knowledge transfer for lifelong robot learn- ing.

Motion-Focused Latent Action Enables Cross-Embodiment VLA Training from Human EgoVideos Libero: Benchmarking knowledge transfer for lifelong robot learn- ing

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T23:40:13.243732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T23:44:35.238493Z digest=sha256:b9dadb6e991593e660c74c87a5290eb786db29a26bf4377059b82fbc9b2672f3

Observation e7cd1eda-e2f0-4e71-a07b-d71711045c2c · outbound

This paper cites RoboEngine: Plug-and-Play Robot Data Augmentation with Semantic Robot Segmentation and Background Generation.

Motion-Focused Latent Action Enables Cross-Embodiment VLA Training from Human EgoVideos RoboEngine: Plug-and-Play Robot Data Augmentation with Semantic Robot Segmentation and Background Generation

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-03T23:49:01.869844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T23:44:35.238493Z digest=sha256:6f66c7b0225114bf5f4dc516fe3780c4814c189bb213adaf6bad9ae5a3ac3090

Observation 148a5ba7-9175-4e64-820b-de5327d18669 · outbound

This paper cites RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation.

Motion-Focused Latent Action Enables Cross-Embodiment VLA Training from Human EgoVideos RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-07-03T23:49:01.880331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T23:44:35.238493Z digest=sha256:e1eaa15833ecdcee6fdc6c3f381a5c24bc9b72c1bf304c2c64154746145b781c

Observation 551bc605-f877-4b8d-94e9-6ea46cb48e3c · outbound

This paper cites EgoDex: Learning Dexterous Manipulation from Large-Scale Egocentric Video.

Motion-Focused Latent Action Enables Cross-Embodiment VLA Training from Human EgoVideos EgoDex: Learning Dexterous Manipulation from Large-Scale Egocentric Video

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-07-03T23:49:01.891877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T23:44:35.238493Z digest=sha256:cff7b1b82143ecf372f40f8ea5c7e45b8a38f8b4eb989498afebe26635874843

Observation 04f6150f-4924-444c-a45e-214a3e1ead4a · outbound

This paper cites Furniturebench: Reproducible real-world benchmark for long-horizon complex manipulation,.

Motion-Focused Latent Action Enables Cross-Embodiment VLA Training from Human EgoVideos Furniturebench: Reproducible real-world benchmark for long-horizon complex manipulation,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T23:40:13.225763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T23:44:35.238493Z digest=sha256:8c9ce98f9b04d58022217930679566ccb31bd3bfa2f4245bf3fba51129f352ab

Observation b3b6e15c-f670-4753-8908-0678d65b0790 · outbound

This paper cites Similarity of neural network representations revisited.

Motion-Focused Latent Action Enables Cross-Embodiment VLA Training from Human EgoVideos Similarity of neural network representations revisited

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T23:40:13.233943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T23:44:35.238493Z digest=sha256:cf73a2df11ad7cd7c988a93f0721aa9d378f26eda153430f843f4b67e3a2f0d2

Pith citing papers

No inbound Pith citation observations are available.