Pith. sign in

Paper Citation Record · LEDGER

Latent Action Pretraining from Videos

As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 73 inbound Pith citation observations for arXiv:2410.11758.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.11758 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 73 of 73 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 73 of 73 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T18:43:23.688296Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T02:04:26.214499Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 4b663d66-8493-4372-a97a-1ac02bb8f69b · inbound

What Matters in Building Vision-Language-Action Models for Generalist Robots cites this paper.

What Matters in Building Vision-Language-Action Models for Generalist Robots Latent Action Pretraining from Videos

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-05-17T21:37:50.804809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T21:37:50.617813Z digest=sha256:c45087b6b4f2bcc7e9c8313e425cfc722d1c4792ac28a66971c5a0929809adb4

Observation 718d15f3-c66f-4a6d-a748-ef6c9ce207ec · inbound

Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations cites this paper.

Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations Latent Action Pretraining from Videos

Reference 132

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:38:11.375731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-12T18:38:11.110166Z digest=sha256:774ae6657037c70ba72359b02c856ae6cd2bac20254e4f997ab05b3035ab11d0

Observation 57cfbfcb-8c20-4ce7-95b9-c5e853cb1f70 · inbound

FAST: Efficient Action Tokenization for Vision-Language-Action Models cites this paper.

FAST: Efficient Action Tokenization for Vision-Language-Action Models Latent Action Pretraining from Videos

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:52:32.103559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:a048582f8e7563cd0d42fc57b2381a95152ebc64518402a94f43cf92ad64b676

Observation 74dea9bf-a527-4cd8-a62e-b4c62ecff04c · inbound

Unified World Models: Coupling Video and Action Diffusion for Pretraining on Large Robotic Datasets cites this paper.

Unified World Models: Coupling Video and Action Diffusion for Pretraining on Large Robotic Datasets Latent Action Pretraining from Videos

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-13T16:25:00.423827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T16:25:00.365534Z digest=sha256:07a839c79330ad91cd7c244956b62bc3e9513742377dc452f63ecc184e813f06

Observation 59f40fba-6a85-41f2-8871-64a387c25b0b · inbound

GraspVLA: a Grasping Foundation Model Pre-trained on Billion-scale Synthetic Action Data cites this paper.

GraspVLA: a Grasping Foundation Model Pre-trained on Billion-scale Synthetic Action Data Latent Action Pretraining from Videos

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:55:52.344513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T20:55:52.109166Z digest=sha256:0f7e6ac170521f6308587f0f9203e43d1d53543e3b1d23c5ff91a3719c8c4a59

Observation 08f125d6-899a-48d4-b885-4a3582659e75 · inbound

DexWild: Dexterous Human Interactions for In-the-Wild Robot Policies cites this paper.

DexWild: Dexterous Human Interactions for In-the-Wild Robot Policies Latent Action Pretraining from Videos

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-05-22T15:21:44.551116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T15:21:21.778285Z digest=sha256:6e90099b6bf5c8e655f68cb9c30dac543079e04339cfe76c7850233b4b58a7fa

Observation f7a95971-8812-4ce4-ad0a-c9ddd885922a · inbound

Robotic Manipulation by Imitating Generated Videos Without Physical Demonstrations cites this paper.

Robotic Manipulation by Imitating Generated Videos Without Physical Demonstrations Latent Action Pretraining from Videos

Reference 128

Resolution
verified exact
local_arxiv, observed 2026-05-19T06:37:07.309359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T06:36:13.144868Z digest=sha256:563dfa64ac1c868f7a234b9c8a701f698c0156361ee212bc8728cbe39696c4b1

Observation 162c26cd-a0a7-4313-83cb-fec5c051812f · inbound

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos cites this paper.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos Latent Action Pretraining from Videos

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-05-21T04:32:58.826620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:ed1bf2d2a10e18fc445d0ca36d46438a74977c162cac7a017d5f895c0d1725b6

Observation d871c39a-369c-4474-83b8-32bcff3a5751 · inbound

GR-3 Technical Report cites this paper.

GR-3 Technical Report Latent Action Pretraining from Videos

Reference 77

Resolution
verified exact
local_arxiv, observed 2026-05-17T08:04:12.611143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T08:04:12.433863Z digest=sha256:6b5e82e4cf81e32a9f78c98961b9541b3f25dd21526f4f2222a42fc9c890522e

Observation 422b08cf-3e20-458e-8975-68ca0fa933fd · inbound

villa-X: Enhancing Latent Action Modeling in Vision-Language-Action Models cites this paper.

villa-X: Enhancing Latent Action Modeling in Vision-Language-Action Models Latent Action Pretraining from Videos

Reference 69

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T21:52:03.172957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T21:52:02.893886Z digest=sha256:5243b4927916be2e65b75bfa807a0d00c7b596b8a26dfff6768a057402f28be6

Observation c892f0a5-ee12-4a77-862d-73d35dd75a6e · inbound

Search for a Heavy-philic W' Boson using Proton-Proton Collisions at Center-of-Mass Energy of 13 TeV Using the ATLAS Detector cites this paper.

Search for a Heavy-philic W' Boson using Proton-Proton Collisions at Center-of-Mass Energy of 13 TeV Using the ATLAS Detector Latent Action Pretraining from Videos

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T18:43:23.688296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:43:23.688296Z digest=sha256:fc9aa9304b970d661e1a68032f30f995fa5c393292eeef2d645d198853c5536c

Observation ebcfce30-e81e-449d-a0f6-1bf1ddd6b4d4 · inbound

Generative Visual Foresight Meets Task-Agnostic Pose Estimation in Robotic Table-Top Manipulation cites this paper.

Generative Visual Foresight Meets Task-Agnostic Pose Estimation in Robotic Table-Top Manipulation Latent Action Pretraining from Videos

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T13:46:41.391260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:46:41.391260Z digest=sha256:75f7cd71169ccf83652c9168cfbdab212ec0d0ce298640bc2c09a221881b7771

Observation ca132bc9-fcf8-499f-8499-0959806fb879 · inbound

RynnVLA-002: A Unified Vision-Language-Action and World Model cites this paper.

RynnVLA-002: A Unified Vision-Language-Action and World Model Latent Action Pretraining from Videos

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T20:59:53.272765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:59:53.272765Z digest=sha256:2f114cb598027760a23b5639df4cc55433a209d839bc09dead4a91e82a2dc7b1

Observation 783b014c-5ef7-4a13-8c2a-707be2221463 · inbound

mimic-video: Video-Action Models for Generalizable Robot Control Beyond VLAs cites this paper.

mimic-video: Video-Action Models for Generalizable Robot Control Beyond VLAs Latent Action Pretraining from Videos

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-05-15T10:41:00.246660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T10:41:00.142543Z digest=sha256:f7c1bb4e410d43a2d1427778170ece9b04cef4c1b05194c22a833740728f296f

Observation 4cd1a96b-5b5c-4cad-8be9-8ed27698a72d · inbound

What Drives Success in Physical Planning with Joint-Embedding Predictive World Models? cites this paper.

What Drives Success in Physical Planning with Joint-Embedding Predictive World Models? Latent Action Pretraining from Videos

Reference 74

Resolution
verified exact
local_arxiv, observed 2026-05-21T15:34:15.028858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-21T15:33:24.616338Z digest=sha256:e19f01f68bd8ca984502439adae32f67ac317d8ba7aee6856ec1f4e21a7c6520

Observation 557db13a-c2fd-4172-80ca-9f1dacb3b113 · inbound

On the Sample Efficiency of Inverse Dynamics Models for Semi-Supervised Imitation Learning cites this paper.

On the Sample Efficiency of Inverse Dynamics Models for Semi-Supervised Imitation Learning Latent Action Pretraining from Videos

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-03T05:21:39.528682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T05:21:39.528682Z digest=sha256:3726699f8fae369b19bb1868b62cc430597027e55d9814c87eebbcce0b00c4f9

Observation 2cc0af82-5331-4aff-8017-07f87a689b9d · inbound

RISE: Self-Improving Robot Policy with Compositional World Model cites this paper.

RISE: Self-Improving Robot Policy with Compositional World Model Latent Action Pretraining from Videos

Reference 89

Resolution
verified exact
local_arxiv, observed 2026-05-16T02:30:32.079195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T02:28:37.997148Z digest=sha256:fa2fe59fb7f3e03adaa4edefe6940a610d32da0e07aec687f52348747b7a68d7

Observation e9cbe408-a726-4645-a5ac-a24067eae156 · inbound

Factored Latent Action World Models cites this paper.

Factored Latent Action World Models Latent Action Pretraining from Videos

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T22:41:10.969247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:41:10.969247Z digest=sha256:bed7348f8a110c1d89353192b928839b6755ba347415bede590710a444672209

Observation 4bcec5a9-7bee-4993-b2ab-4b2af87ba99f · inbound

UniLACT: Depth-Aware RGB Latent Action Learning for Vision-Language-Action Models cites this paper.

UniLACT: Depth-Aware RGB Latent Action Learning for Vision-Language-Action Models Latent Action Pretraining from Videos

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-15T20:20:17.696163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:31.988002Z digest=sha256:d9f22209fd2683100f036b458185cf1dd80713d850b6d97cc4c51e2319ab925f

Observation 281fab71-0799-48f0-ab05-7e4219056cad · inbound

The Latent Space: Foundation, Evolution, Mechanism, Ability, and Outlook cites this paper.

The Latent Space: Foundation, Evolution, Mechanism, Ability, and Outlook Latent Action Pretraining from Videos

Reference 272

Resolution
unresolved
no resolver link, observed 2026-07-13T14:03:01.974171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T14:03:01.974171Z digest=sha256:bdf09646587e62705c3ccdc6b45d076c23a1b98fa298ad0e9c56a449c0094b10

Observation c6dce25b-abb2-410d-9f86-122ab52e8761 · inbound

From Video to Control: A Survey of Learning Manipulation Interfaces from Temporal Visual Data cites this paper.

From Video to Control: A Survey of Learning Manipulation Interfaces from Temporal Visual Data Latent Action Pretraining from Videos

Reference 104

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T17:03:01.026422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T17:02:18.358675Z digest=sha256:074baa3cd2a351bd25e6573feb881fef24c4432107786031ec64f090699781f9

Observation 95c70fb4-f443-4072-8c9d-4748730cf6d0 · inbound

EgoVerse: An Egocentric Human Dataset for Robot Learning from Around the World cites this paper.

EgoVerse: An Egocentric Human Dataset for Robot Learning from Around the World Latent Action Pretraining from Videos

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:35:56.946457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T17:07:41.489995Z digest=sha256:40af006519683ec18712cb329f4a584a2230204d3f65c694d7b02ed5632144d0

Observation bfe5ab6b-e649-4b48-b065-68b3c60dd5e9 · inbound

EgoVerse: An Egocentric Human Dataset for Robot Learning from Around the World cites this paper.

EgoVerse: An Egocentric Human Dataset for Robot Learning from Around the World Latent Action Pretraining from Videos

Reference 57

Resolution
unresolved
no resolver link, observed 2026-07-13T08:25:22.011013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T08:25:22.011013Z digest=sha256:925353260697fdd734499afeb2178f82412d17e6e2c518a4d6352b1472bfc89c

Observation af5bd6a6-0072-47f5-9d62-9bd55e0a02cd · inbound

A Mechanistic Analysis of Sim-and-Real Co-Training in Generative Robot Policies cites this paper.

A Mechanistic Analysis of Sim-and-Real Co-Training in Generative Robot Policies Latent Action Pretraining from Videos

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:30:22.742733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T12:29:38.306670Z digest=sha256:f9d1c6df58a6da9ef0f404c3d6a1b9a1eab44be383cb731abdad56606d56c5f6

Observation ed6fc1aa-5a18-4638-9451-f87cd753eeb2 · inbound

${\pi}_{0.7}$: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities cites this paper.

${\pi}_{0.7}$: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities Latent Action Pretraining from Videos

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:45:21.748452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T11:42:34.409651Z digest=sha256:fdd79d09df5ec3a457966bd620db9a8ac8b5576951e8c113e7447cce79d66686

Observation 5feb4b7b-0335-40b3-8c38-b354eac17ea1 · inbound

RoboWM-Bench: A Benchmark for Evaluating World Models in Robotic Manipulation cites this paper.

RoboWM-Bench: A Benchmark for Evaluating World Models in Robotic Manipulation Latent Action Pretraining from Videos

Reference 58

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T03:03:37.436932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T03:02:26.084185Z digest=sha256:fdace5627de85dba38c04b171c86ffcbce0493457dcd4a6c42c031a38c519597

Observation d2e47202-b52d-4623-ab11-734bc07b01f5 · inbound

RoboWM-Bench: A Benchmark for Evaluating World Models in Robotic Manipulation cites this paper.

RoboWM-Bench: A Benchmark for Evaluating World Models in Robotic Manipulation Latent Action Pretraining from Videos

Reference 58

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T06:19:50.224025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T06:15:32.881140Z digest=sha256:29346d12e57a58a0ff2582fc53213b530e9b77128c061ec04c3bc1d7b34e23fa

Observation 95635ea6-2ab6-4671-aa50-56262d2d12c1 · inbound

UniT: Toward a Unified Physical Language for Human-to-Humanoid Policy Learning and World Modeling cites this paper.

UniT: Toward a Unified Physical Language for Human-to-Humanoid Policy Learning and World Modeling Latent Action Pretraining from Videos

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:16:06.498105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T02:04:25.672157Z digest=sha256:6f523da01fbc9c07f7b241a3afe3df4cc641ba73e0e851571e450601cd79bbba

Observation a2f3427d-62f4-44b1-814c-c769af1995ca · inbound

Hi-WM: Human-in-the-World-Model for Scalable Robot Post-Training cites this paper.

Hi-WM: Human-in-the-World-Model for Scalable Robot Post-Training Latent Action Pretraining from Videos

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:36:07.485520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-09T21:26:26.540403Z digest=sha256:8a60f92d6b5fbaed42d84546d963b8dcc939723a7c836e3f0b1d42d460d43e29

Observation 4f55cb2a-6e1f-4117-8e63-e78b2d38aac7 · inbound

GazeVLA: Learning Human Intention for Robotic Manipulation cites this paper.

GazeVLA: Learning Human Intention for Robotic Manipulation Latent Action Pretraining from Videos

Reference 73

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T19:36:12.842343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T11:37:30.784513Z digest=sha256:10ed935b2c0da90b3ca775bfdbc4f1a451fd9d7c18525a3c0580457aea349931

Observation b659df27-96a0-4d02-a1ce-c5190e908742 · inbound

LA-Pose: Latent Action Pretraining Meets Pose Estimation cites this paper.

LA-Pose: Latent Action Pretraining Meets Pose Estimation Latent Action Pretraining from Videos

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:51:28.799831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T08:57:27.503349Z digest=sha256:084753efdc1b622fd9e9a09705b25902abd365ede38e6cf5b19b588cc0aa3adf

Observation ab75a691-8c4e-4a04-b1b0-9be9c5be6d59 · inbound

Being-H0.7: A Latent World-Action Model from Egocentric Videos cites this paper.

Being-H0.7: A Latent World-Action Model from Egocentric Videos Latent Action Pretraining from Videos

Reference 86

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:56:08.337677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-09T20:48:01.461993Z digest=sha256:c8a9b2a3032cd1cac9a66115bed49e6303e7091da3bf3e46c72f116d7b42a7a2

Observation 53e5e099-6103-47b3-973e-8486b028db50 · inbound

Latent State Design for World Models under Sufficiency Constraints cites this paper.

Latent State Design for World Models under Sufficiency Constraints Latent Action Pretraining from Videos

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:36:04.583192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:55:31.825583Z digest=sha256:8099e447f51744247ba3a360666463f6305efda07f95c37e8bf08287552e5fb0

Observation fc4ada06-80e7-4344-95ab-7ce18c2cf58b · inbound

Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing cites this paper.

Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing Latent Action Pretraining from Videos

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-12T11:01:31.461201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T15:44:25.021995Z digest=sha256:46f829d28ff846f2f7289c29df9a7d72cd9ca40921d11b394a50235cf9c04b37

Observation a64c4529-209d-4c34-ace1-8a7ef737fec8 · inbound

From Pixels to Tokens: A Systematic Study of Latent Action Supervision for Vision-Language-Action Models cites this paper.

From Pixels to Tokens: A Systematic Study of Latent Action Supervision for Vision-Language-Action Models Latent Action Pretraining from Videos

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:16:09.847697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-08T17:43:00.453164Z digest=sha256:cf0cb4869f0c14cb6e9881e570b29e87185d88146ec78f23085d2a5f40a65322

Observation 7bab4f9d-daa8-4fb9-a197-cd5319671327 · inbound

When to Trust Imagination: Adaptive Action Execution for World Action Models cites this paper.

When to Trust Imagination: Adaptive Action Execution for World Action Models Latent Action Pretraining from Videos

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:21:13.052348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T09:17:18.017938Z digest=sha256:6edfd7e03e8d50807edca987dd3b197bd81a9fd69f55f35fc2ac55c2e369db9c

Observation be29d844-acc0-4621-9048-b9d923f4e127 · inbound

When to Trust Imagination: Adaptive Action Execution for World Action Models cites this paper.

When to Trust Imagination: Adaptive Action Execution for World Action Models Latent Action Pretraining from Videos

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:01:31.404350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:22:49.557507Z digest=sha256:e6cab5188d8a0424e86cb1b1cec63e6808f29f917591d26913429e7d835eb41d

Observation ee50c076-a6d5-4c52-bf2b-18b7207e9672 · inbound

Render, Don't Decode: Weight-Space World Models with Latent Structural Disentanglement cites this paper.

Render, Don't Decode: Weight-Space World Models with Latent Structural Disentanglement Latent Action Pretraining from Videos

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:51:07.838477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T13:39:05.969414Z digest=sha256:4726052acc6ddf4340bdda80e019429279949cbb7e9e932b78731a0bc05324d4

Observation d942fe46-0322-41ba-9368-dba678739730 · inbound

Render, Don't Decode: Weight-Space World Models with Latent Structural Disentanglement cites this paper.

Render, Don't Decode: Weight-Space World Models with Latent Structural Disentanglement Latent Action Pretraining from Videos

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:20:58.598501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:46:31.004068Z digest=sha256:92a0a9e448445b03e058c637e0eceaabc62a5d3c645d48f6129c123f3ef4b673

Observation 94b943df-dd28-46ef-a51b-ab29e5546782 · inbound

SABER: A Scalable Action-Based Embodied Dataset for Real-World VLA Adaptation cites this paper.

SABER: A Scalable Action-Based Embodied Dataset for Real-World VLA Adaptation Latent Action Pretraining from Videos

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:36:26.268876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T04:08:29.524972Z digest=sha256:8c6a2264af61b17b72591fed4c93e9fd246d77619785c5b5db3c815320defb6c

Observation 345855cf-2400-471f-9695-6726e08eff02 · inbound

ALAM: Algebraically Consistent Latent Action Model for Vision-Language-Action Models cites this paper.

ALAM: Algebraically Consistent Latent Action Model for Vision-Language-Action Models Latent Action Pretraining from Videos

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:26:26.843140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T04:14:54.885244Z digest=sha256:91ac8226204f7a6dea6a1689f791ec7aa3a85d0aca7015952ef7930d5e3721f8

Observation 9e2b3200-77ed-44d5-8a4c-74ace4c3b1fb · inbound

ALAM: Algebraically Consistent Latent Action Model for Vision-Language-Action Models cites this paper.

ALAM: Algebraically Consistent Latent Action Model for Vision-Language-Action Models Latent Action Pretraining from Videos

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-05-14T21:17:59.574580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T21:14:56.501485Z digest=sha256:4bbec55643727fc2956e4d311ed5a3160084ff9e358ff6ca6ded78f0c2e12271

Observation 50fb0860-61ce-4703-b3c6-16998b91ec14 · inbound

RotVLA: Rotational Latent Action for Vision-Language-Action Model cites this paper.

RotVLA: Rotational Latent Action for Vision-Language-Action Model Latent Action Pretraining from Videos

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-14T17:49:23.117129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T17:48:06.734816Z digest=sha256:c7ba959a52fbe47af85d0e9a63f355c441e7f37ff54b0abe6c4879f1b1ff9672

Observation 34f3fdd2-3ef4-4666-9f7e-78083bfac972 · inbound

CUBic: Coordinated Unified Bimanual Perception and Control Framework cites this paper.

CUBic: Coordinated Unified Bimanual Perception and Control Framework Latent Action Pretraining from Videos

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-05-14T19:07:51.118404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T19:06:13.250054Z digest=sha256:9f2bce08522a86e5c5e389163c50209dc729ee59b5a191b52932f0610b81415b

Observation 9d077384-3d62-4f57-a49a-0e2e233766a6 · inbound

DiLA: Disentangled Latent Action World Models cites this paper.

DiLA: Disentangled Latent Action World Models Latent Action Pretraining from Videos

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-20T19:38:56.304993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T19:35:37.527479Z digest=sha256:7ac6f172d51cc1100990e256b82baa2d909d8af0f977e3314085fe07ffaf376c

Observation 8517292a-d71d-47b9-ae81-f425017c7507 · inbound

RoboWits: Unexpected Challenges for Robotic Creative Problem Solving cites this paper.

RoboWits: Unexpected Challenges for Robotic Creative Problem Solving Latent Action Pretraining from Videos

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-06-29T07:13:16.894631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T07:06:10.838225Z digest=sha256:76a9035cd7545b31343f0467e336a63fd3dfef79cfadf07fb8513323c7bf4d75

Observation 35ba7f4b-76ac-43b1-8817-eabe773a0b30 · inbound

HARP-VLA: Human-Robot Aligned Representation Learning for Vision-Language-Action Model cites this paper.

HARP-VLA: Human-Robot Aligned Representation Learning for Vision-Language-Action Model Latent Action Pretraining from Videos

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-01T19:46:10.862338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T22:04:09.296855Z digest=sha256:807a4402e407a12673ef8a7e8fe3b2c6c083b898eb517adaa57a67720e90a112

Observation f1738901-b194-4700-a8bf-8b503e193eaf · inbound

CLAW: Learning Continuous Latent Action World Models via Adversarial Latent Regularization cites this paper.

CLAW: Learning Continuous Latent Action World Models via Adversarial Latent Regularization Latent Action Pretraining from Videos

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-07-02T03:36:29.608591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T09:55:00.402411Z digest=sha256:b1abc6a96ddfdd240319aa03e0545c91cdc0f8344703070495a3aae3ac70feab

Observation 4391f14c-a5c0-42de-8237-3ce75bbd106e · inbound

Robots Need More than VLA and World Models cites this paper.

Robots Need More than VLA and World Models Latent Action Pretraining from Videos

Reference 104

Resolution
verified exact
local_arxiv, observed 2026-07-02T13:46:59.186305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T01:01:33.530167Z digest=sha256:4cd477fe085228fa8eeeab8773c8346c88ac87e8c42690d8f880779171ea331f

Observation 7b69703f-24b7-4b53-85b7-a1b029ff7cc9 · inbound

LARA: Latent Action Representation Alignment for Vision-Language-Action Models cites this paper.

LARA: Latent Action Representation Alignment for Vision-Language-Action Models Latent Action Pretraining from Videos

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-07-02T16:47:09.516655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-27T22:25:17.522240Z digest=sha256:e8428c9ff40633c670e8f8c07c42f116d119577e139e347ec549bf5f36845814

Observation 7884fe78-56fe-497d-9805-d85f4d4f8ced · inbound

LARA: Latent Action Representation Alignment for Vision-Language-Action Models cites this paper.

LARA: Latent Action Representation Alignment for Vision-Language-Action Models Latent Action Pretraining from Videos

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-07-01T08:35:34.475835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-01T07:17:42.045939Z digest=sha256:a98bbcf478a22a4a726949ce6dee8d443cdd1d5646689b3436042278461c3ecb

Observation 948bdecb-550b-4f81-a356-b599cd920636 · inbound

Latent Diffusion Policy: Shaping Latent Spaces for Diffusion-Based Robotic Manipulation cites this paper.

Latent Diffusion Policy: Shaping Latent Spaces for Diffusion-Based Robotic Manipulation Latent Action Pretraining from Videos

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-07-02T23:07:26.991234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T18:28:06.679879Z digest=sha256:ecc2f79f1d5b847c7aa827b0dc378f8d21ad6ea553585f5b62cd7ce0e535586d

Observation 90730065-0f09-4da3-80ab-3d3352ee0464 · inbound

GHOST: Hierarchical Sub-Goal Policies for Generalizing Robot Manipulation cites this paper.

GHOST: Hierarchical Sub-Goal Policies for Generalizing Robot Manipulation Latent Action Pretraining from Videos

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-07-03T02:07:33.565797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:10:40.130366Z digest=sha256:9fb4114ea943a413b0b386018e8c90b0210014a1abf4cc7d752f4967720af954

Observation ceacc4e4-9757-4a90-82c6-c6eea57beb09 · inbound

Contrastive Action-Image Pre-training for Visuomotor Control cites this paper.

Contrastive Action-Image Pre-training for Visuomotor Control Latent Action Pretraining from Videos

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-07-03T18:08:46.891731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T03:16:33.702802Z digest=sha256:9f2b9c23a11f9b2d6dfd200bc241fcf7e89271d769ef6209056a79129f8e289b

Observation 8a17dfe7-1109-40ea-9c7e-46f3469a7fdc · inbound

PearlVLA: Progressive Embodied Action-Plan Refinement in Latent Space cites this paper.

PearlVLA: Progressive Embodied Action-Plan Refinement in Latent Space Latent Action Pretraining from Videos

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:58:58.523547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T00:57:58.312567Z digest=sha256:2313b4f41cd8a52bb89506f53f16a6659f7b6e823a66d977660160a79b7c1163

Observation a2d622b1-7f9d-4fb1-b214-fdf4b3103f76 · inbound

Motion-Focused Latent Action Enables Cross-Embodiment VLA Training from Human EgoVideos cites this paper.

Motion-Focused Latent Action Enables Cross-Embodiment VLA Training from Human EgoVideos Latent Action Pretraining from Videos

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-03T23:39:04.279675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T21:53:27.323112Z digest=sha256:31715aede4634be3dc0870fe2d12df00bf1665a62f67bd3926d2573efed48ac3

Observation f84b4399-2964-4786-994f-8c75f4e7438c · inbound

Motion-Focused Latent Action Enables Cross-Embodiment VLA Training from Human EgoVideos cites this paper.

Motion-Focused Latent Action Enables Cross-Embodiment VLA Training from Human EgoVideos Latent Action Pretraining from Videos

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-03T23:49:01.879308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-03T23:44:35.238493Z digest=sha256:5d13bd3b64ccdd59f9217afeeac0acf59c28cf9d6e5bc82973200d30626f6e17

Observation f5e9397d-83af-4da1-a90b-00efc767fa36 · inbound

FOCA: Future-Oriented Conditioning for Data-Efficient Vision-Language-Action Adaptation cites this paper.

FOCA: Future-Oriented Conditioning for Data-Efficient Vision-Language-Action Adaptation Latent Action Pretraining from Videos

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-07-04T03:29:30.436606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-26T18:01:08.616612Z digest=sha256:06a66991eb2ca12430407aa7cd0235390e6286ed51573a895604e6b77c65d254

Observation 6a3c773c-5550-4fa4-ae24-c45bc48b93a9 · inbound

PoLAR: Factorizing Extent and Mode in Latent Actions for Robot Policy Learning cites this paper.

PoLAR: Factorizing Extent and Mode in Latent Actions for Robot Policy Learning Latent Action Pretraining from Videos

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-04T06:29:37.214516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T14:28:36.505775Z digest=sha256:b879fe20854025ce52485e7d97caa92dcfaac742933f7cd4ba4ebd6f2767afa4

Observation bf443342-1d89-4983-89af-84090af3c5a4 · inbound

Imitation from Heterogeneous Demonstrations using Grounded Latent-Action World Models cites this paper.

Imitation from Heterogeneous Demonstrations using Grounded Latent-Action World Models Latent Action Pretraining from Videos

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-07-04T06:49:38.608997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T14:07:31.139419Z digest=sha256:ed96a6e2f9e52e9ae72220d36127cdb8a99ea65c06dd992ebeda19092ab96dd0

Observation 46779f88-38c3-41d5-bc8c-b5003f920236 · inbound

LaST-HD: Learning Latent Physical Reasoning from Scalable Human Data for Robot Manipulation cites this paper.

LaST-HD: Learning Latent Physical Reasoning from Scalable Human Data for Robot Manipulation Latent Action Pretraining from Videos

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-07-04T11:29:50.787907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T07:58:17.225491Z digest=sha256:47b58125e225edcbe62d8306cf2675eb0d60a681200c1bedf010acdb5cd3d488

Observation ff049e46-0992-4cf3-95e3-0ce3bb58cbcb · inbound

Learning Action Priors for Cross-embodiment Robot Manipulation cites this paper.

Learning Action Priors for Cross-embodiment Robot Manipulation Latent Action Pretraining from Videos

Reference 59

Resolution
verified exact
local_arxiv, observed 2026-07-04T21:00:09.755926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-25T19:09:56.409766Z digest=sha256:154b64872eff95879bcdc3167deab4ba88b8283d3683e076a4299e7ccae27e29

Observation c4a7442a-468c-41c6-b12c-3c8dc5c95e9a · inbound

Translation as a Bridging Action: Transferring Manipulation Skills from Humans to Robots cites this paper.

Translation as a Bridging Action: Transferring Manipulation Skills from Humans to Robots Latent Action Pretraining from Videos

Reference 63

Resolution
verified exact
local_arxiv, observed 2026-07-01T16:55:51.442636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T04:23:04.622902Z digest=sha256:1a34aa33cc3796ff4f0aa6493b488133d821a74a2b3615cc7778f5590b50c68d

Observation 039623ae-f676-429e-99f9-17c7b5ba3862 · inbound

Latent Actions from Factorized Transition Effects under Agent Ambiguity cites this paper.

Latent Actions from Factorized Transition Effects under Agent Ambiguity Latent Action Pretraining from Videos

Reference 57

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T06:04:21.449589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-30T06:00:35.407561Z digest=sha256:3860a6300c6bb31402b19a3cd43cadccf19769873c0d8100b4754a42f39dbeb1

Observation 089d6f47-89a4-4eb8-84bf-89146babc735 · inbound

ABot-M0.5: Unified Mobility-and-Manipulation World Action Model cites this paper.

ABot-M0.5: Unified Mobility-and-Manipulation World Action Model Latent Action Pretraining from Videos

Reference 72

Resolution
verified exact
local_arxiv, observed 2026-07-02T14:27:03.155846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-02T14:24:23.187164Z digest=sha256:77bada24230c57db76f99a60bd119163c945dece1928f9677abbe1b0c967e5dd

Observation ec100e60-b7b2-4e42-99bb-c4a79f5d3ece · inbound

ABot-M0.5: Unified Mobility-and-Manipulation World Action Model cites this paper.

ABot-M0.5: Unified Mobility-and-Manipulation World Action Model Latent Action Pretraining from Videos

Reference 72

Resolution
unresolved
no resolver link, observed 2026-07-12T09:22:08.000379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:22:08.000379Z digest=sha256:ec457c7194575f6987c8052c888cfaae1be2be6558838403b81a7c8d3da9179e

Observation 1f49447e-5467-4902-921e-85dd27b3df2c · inbound

Geometry-Aware Motion Latents for Learning Robust Manipulation Policies cites this paper.

Geometry-Aware Motion Latents for Learning Robust Manipulation Policies Latent Action Pretraining from Videos

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-11T14:43:12.450611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:43:12.450611Z digest=sha256:89da366c236b83c1b66d92ad419081eef94d2ba89546f56614e32d06f6f56a02

Observation 43e43fbf-4d9f-4984-bf03-74544447af01 · inbound

CAC-VLA: Context-Gated Action Conditioning for Vision-Language-Action Models cites this paper.

CAC-VLA: Context-Gated Action Conditioning for Vision-Language-Action Models Latent Action Pretraining from Videos

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-11T13:05:32.374206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T13:05:32.374206Z digest=sha256:6861ef598dc721a6db05e73d82ccf1ea5bbdb9b256be251c7361fed73188258b

Observation 88f73141-048e-4cc6-9507-421e773002ce · inbound

RynnWorld-4D: 4D Embodied World Models for Robotic Manipulation cites this paper.

RynnWorld-4D: 4D Embodied World Models for Robotic Manipulation Latent Action Pretraining from Videos

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-07-08T02:04:26.215924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-08T02:00:30.551638Z digest=sha256:077002382388465c6fef12e466d7ec008f97104f72cd5526401fdf2f0cbb125a

Observation 59bb24ef-efbe-433b-a9b7-d23c44443329 · inbound

WALA Learning Executable Latent Actions from Action-Labeled Demonstrations and Action-Free Videos cites this paper.

WALA Learning Executable Latent Actions from Action-Labeled Demonstrations and Action-Free Videos Latent Action Pretraining from Videos

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-14T05:52:34.589171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T05:52:34.589171Z digest=sha256:79cf228c633d36c7f86ba255474f2758a1df8a00aec13557f9db9b5af569432a

Observation 0d1d8b2e-1fda-45dc-b515-a1acf0c57c53 · inbound

RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation cites this paper.

RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation Latent Action Pretraining from Videos

Reference 268

Resolution
unresolved
no resolver link, observed 2026-08-01T14:39:52.535448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:39:52.535448Z digest=sha256:2ab5679bcc61450ca0bc250760c547480436b10afb0c620c166d72368a22f95c

Observation 1508b492-828f-45b8-883f-52994a9080f6 · inbound

DLAM: Distributional Latent Actions with Temporal Constraints cites this paper.

DLAM: Distributional Latent Actions with Temporal Constraints Latent Action Pretraining from Videos

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-30T11:05:16.022725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T11:05:16.022725Z digest=sha256:eb52d44cfa7936a9a4b371469b89ecaea4302c33879d27218ad24283e0ea7492

Observation 05fe152a-2587-4325-b096-f1af5b1967d0 · inbound

Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills cites this paper.

Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills Latent Action Pretraining from Videos

Reference 282

Resolution
unresolved
no resolver link, observed 2026-08-04T19:45:35.397349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:45:35.397349Z digest=sha256:16ab02cb5945e458920668dc3240051a742fd9ddccb054c14b065287dcb0ba8c