Pith. sign in

Paper Citation Record · LEDGER

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning

As of 11 August 2026, this Paper Citation Record lists 86 of 86 outbound references and 15 inbound Pith citation observations for arXiv:2505.17006.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.17006 v3

Coverage vector

measured 86 of 86 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:56:41.618867Z

measured 101 of 101 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T06:18:04.994485Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-03T04:27:36.264850Z

Reference resolution

86 of 86 outbound references displayed

  • verified exact0
  • verified fuzzy33
  • unresolved53
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 22386f00-d9c6-4ada-a704-6328c1c6624a · outbound

This paper cites GPT-4 Technical Report.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:31.783573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:56:31.783573Z digest=sha256:7db33b593e3e7b8a879c5ce77ea1decc227a1bf9e58132a79b6838006f0f7542

Observation 99f1f8fa-9ed1-4a25-945d-2f7e359b8c0b · outbound

This paper cites Learning dexterous in-hand manipulation.The Inter- national Journal of Robotics Research, 39(1):3–20, 2020.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning Learning dexterous in-hand manipulation.The Inter- national Journal of Robotics Research, 39(1):3–20, 2020

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:31.830695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:56:31.830695Z digest=sha256:a3455e6a70f5370372768b4c90dd3ac27b08cbc8c826d1586f4a85faf477b30f

Observation 76ea3cf0-984a-4faa-b96e-7d13b059ff97 · outbound

This paper cites Affordances from human videos as a versatile representation for robotics.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning Affordances from human videos as a versatile representation for robotics

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:31.950338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:56:31.950338Z digest=sha256:e65c934a138e7ddbad44f0cbdf85f9e424076d35eaed82f251c7c7c4c8db7ff6

Observation 1d7c1656-b81d-498e-b6b0-42b23703f075 · outbound

This paper cites Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:32.082618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:56:32.082618Z digest=sha256:92e590b4b920dec9f3c3a83d2570b9163565552b9a6c8620f7bbe2a576ab8070

Observation 386ece9e-0e21-4a07-9c8e-8e0413f31b48 · outbound

This paper cites Towards generalizable zero-shot manip- ulation via translating human interaction plans.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning Towards generalizable zero-shot manip- ulation via translating human interaction plans

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:32.233704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:56:32.233704Z digest=sha256:8e21cb683363cb41133b1df6ffdaceee50ed3dccff2c1e4830ae8d1c6b6fc1b3

Observation d3afaf54-fca5-4125-8d64-718bae8d046c · outbound

This paper cites Track2act: Predicting point tracks from internet videos enables diverse zero-shot robot manip- ulation.arXiv e-prints, pages arXiv–2405, 2024.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning Track2act: Predicting point tracks from internet videos enables diverse zero-shot robot manip- ulation.arXiv e-prints, pages arXiv–2405, 2024

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:32.384855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:56:32.384855Z digest=sha256:136bd68d782b7de777c3cea1b8d42f3d8438d635f1eca3d9a090ee1f7335037f

Observation 8f4c2a1b-f6b6-4ac7-94f4-5a4b5aee3c0d · outbound

This paper cites GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:32.499878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:56:32.499878Z digest=sha256:ec76788e9b55a533d17e8e23ddb4c94855b456fc06108f897fbd23ab48724e8a

Observation 1f6a5cbe-014d-4cec-a6da-53216c3957cf · outbound

This paper cites an unresolved cited work.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:32.607048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:56:32.607048Z digest=sha256:8891b709004749861e78315b953536a67d36a3f8880c5e9347f8acc4c520ba3f

Observation fbd56a7c-1d17-4f38-96a9-c1913a060f50 · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning RT-1: Robotics Transformer for Real-World Control at Scale

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:32.612991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:56:32.612991Z digest=sha256:77452086858169f49cf782bf2d2023d59131d176c5d6997fb7ed06658bace096

Observation cc155f46-a02e-438c-9f55-e2a452569f54 · outbound

This paper cites AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:32.616203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:56:32.616203Z digest=sha256:0d7e83e581a1df84b37f2d65e9f1f34708b1f2d34e77db218644d347d121eafd

Observation 7877c53e-9388-40f5-8970-c782b2c0e3f9 · outbound

This paper cites UniVLA: Learning to Act Anywhere with Task-centric Latent Actions.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning UniVLA: Learning to Act Anywhere with Task-centric Latent Actions

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:32.640982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:56:32.640982Z digest=sha256:81a3534d9bf2b59a206e9f4e1ed3aae9cd9c108a05f26deb1e1ef6df82183fe5

Observation 1d476429-2f2e-4b3f-a7de-7f7bf0476c35 · outbound

This paper cites WorldVLA: Towards Autoregressive Action World Model.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning WorldVLA: Towards Autoregressive Action World Model

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:32.779945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:56:32.779945Z digest=sha256:dccd344814c483cf21799b6ab2b5ca4ac9c239eb697ed539f4d3708c7ab466ce

Observation 24bbbd95-4594-40e9-8c1b-be99106b09b0 · outbound

This paper cites GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:32.880004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:56:32.880004Z digest=sha256:ec0b7e1347ee992cd4ce8195d33a21203a580b831caa2c6eb7b556d2c3802186

Observation 7535f42c-d1de-40cb-bfbf-eac6ee1cc129 · outbound

This paper cites DeepVerse: 4D Autoregressive Video Generation as a World Model.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning DeepVerse: 4D Autoregressive Video Generation as a World Model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:33.022075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:56:33.022075Z digest=sha256:8f30d0f93dd2f66b4dd9a4ebfe39585ebb8d65e73204edd43bb0375ffe782e2c

Observation 87d28d03-a0ad-4bca-a374-5a7b333b7764 · outbound

This paper cites IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AI.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AI

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:33.139916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:56:33.139916Z digest=sha256:d50d2df07b99c9fb944692bf70c34333056251c1c55f78335b2b8678aff2f9fb

Observation 890bf866-5c44-4517-8078-2936bd53db35 · outbound

This paper cites villa-X: Enhancing Latent Action Modeling in Vision-Language-Action Models.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning villa-X: Enhancing Latent Action Modeling in Vision-Language-Action Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:33.228796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:56:33.228796Z digest=sha256:736a0c8587aa938a2e2194c164002e353954cf98a6a47bb0fbae414f148e1f25

Observation d8314d15-f6de-4d58-99b1-a3cc8eea7fce · outbound

This paper cites Moto: Latent mo- tion token as the bridging language for learning robot ma- nipulation from videos.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning Moto: Latent mo- tion token as the bridging language for learning robot ma- nipulation from videos

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:33.344464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:56:33.344464Z digest=sha256:b9bff028ffa1073bf71d58ce36177dfebd6f842dd9675255be9d778ccec15baa

Observation b25bcbbb-f9ed-4986-9757-a39fa3f72be2 · outbound

This paper cites Universal manipulation interface: In-the-wild robot teaching without in-the-wild robots.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning Universal manipulation interface: In-the-wild robot teaching without in-the-wild robots

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:56:49.850891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T14:56:33.467736Z digest=sha256:377628029cea9d61d2c215afc9da1f2900a388a1307122274810d21773f59434

Observation e57e62d6-09f1-4b6e-a4f5-1baea46075bb · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action dif- fusion.The International Journal of Robotics Research, 44 (10-11):1684–1704, 2025.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning Diffusion policy: Visuomotor policy learning via action dif- fusion.The International Journal of Robotics Research, 44 (10-11):1684–1704, 2025

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:56:49.659069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T14:56:33.594230Z digest=sha256:1ef4c0e49d0e363eb8df0e8d75bec20efdfb18193df91bcb2c1805d24e418c2a

Observation 8ec858f1-e571-4441-a762-6b4c3f6b9520 · outbound

This paper cites Dynamo: In-domain dynamics pretraining for visuo-motor control.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning Dynamo: In-domain dynamics pretraining for visuo-motor control

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:56:49.424023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T14:56:33.657199Z digest=sha256:91f61e5bc3a9884912be6c355a91b594da84e2d8a008e21f5a783db29e6a802a

Observation dc50e3be-39b8-4711-8ce8-e63e5b8da00d · outbound

This paper cites Bert: Pre-training of deep bidirectional trans- formers for language understanding.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning Bert: Pre-training of deep bidirectional trans- formers for language understanding

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:56:49.183135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T14:56:33.803862Z digest=sha256:d766ec996ddb984d1ef1c6919d81aa66ab3e533989c0c6fe41eaf8a15621b1fe

Observation 017c1337-9fed-4e00-869e-2ad12297b903 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning An image is worth 16x16 words: Transformers for image recognition at scale

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:33.961506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:56:33.961506Z digest=sha256:ea05666b145852cfdb8bd0160ba03986a704a87f42e33918e005ae6e10fb84ab

Observation b8b988ff-16c2-4c2b-81a9-b914a0bf0585 · outbound

This paper cites Knowledge Insulating Vision-Language-Action Models: Train Fast, Run Fast, Generalize Better.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning Knowledge Insulating Vision-Language-Action Models: Train Fast, Run Fast, Generalize Better

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:34.138568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:56:34.138568Z digest=sha256:7e8b6ecfa20d530deff73084d66f01abe8df3b6ab6e9cc4ab83518df5c12be65

Observation ae624b62-179a-41cd-83ef-ce0090126007 · outbound

This paper cites Learning universal policies via text-guided video genera- tion.Advances in neural information processing systems, 36:9156–9172, 2023.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning Learning universal policies via text-guided video genera- tion.Advances in neural information processing systems, 36:9156–9172, 2023

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:34.230249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:56:34.230249Z digest=sha256:baf743110685b2dbc179f597ed30ff5718ec1d439834e60b8044f280a35e7cfd

Observation 63b67dff-4f12-45d7-a84c-d065bb2cc811 · outbound

This paper cites Fluid: Scaling Autoregressive Text-to-image Generative Models with Continuous Tokens.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning Fluid: Scaling Autoregressive Text-to-image Generative Models with Continuous Tokens

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:34.349931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:56:34.349931Z digest=sha256:aa1db1cfbe0f6f75d99df8f31f542851d905d3152c23b3debdda361e3f2681e4

Observation f23f9dd9-7973-42a6-a4de-bf8c67503b65 · outbound

This paper cites RH20T: A comprehensive robotic dataset for learning di- verse skills in one-shot.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning RH20T: A comprehensive robotic dataset for learning di- verse skills in one-shot

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:56:48.737146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T14:56:34.499860Z digest=sha256:b24f4314c761dc852cebeb5b9a11c015fd51511bf8f5eedce162499b4a727fcb

Observation 71c825b4-9aa3-4f05-98b7-4c0bee54a48f · outbound

This paper cites Adaworld: Learning adaptable world models with latent actions.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning Adaworld: Learning adaptable world models with latent actions

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:56:48.544160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T14:56:34.636543Z digest=sha256:912a74445013e7a08f996b288118d4ecb37a0cdca76f28608217ac05e0effed2

Observation 87530bff-c61e-4481-80ef-73a7fec925b7 · outbound

This paper cites DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:34.733788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:56:34.733788Z digest=sha256:2b0dd85d1ccaaf9f2e45898426c84f710cb2183f64c60d7c4128f1713a64f0fe

Observation fe736a04-2a99-4177-9d95-354fab74729e · outbound

This paper cites RT-Trajectory: Robotic Task Generalization via Hindsight Trajectory Sketches.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning RT-Trajectory: Robotic Task Generalization via Hindsight Trajectory Sketches

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:34.869860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:56:34.869860Z digest=sha256:9dd72eaaa5134e8802e7265916ef8b5adbeb78dcc705aca98e7804eab621ccb4

Observation 7da0c5f5-ca0f-4650-b360-6dfc8ec8397c · outbound

This paper cites Masked autoencoders are scalable vision learners.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning Masked autoencoders are scalable vision learners

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:56:48.345966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T14:56:35.045813Z digest=sha256:31098262b26389837e5f5352338036799b0454e557234d0b68ce282e1d970ccc

Observation 6b7d3554-622b-4b2d-93f3-c35a46291c63 · outbound

This paper cites Denoising dif- fusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning Denoising dif- fusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:56:48.099286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T14:56:35.187942Z digest=sha256:11d4acb94d6fabb0c91c743c699f3374370255dbfdf597330ebde76284ea29e7

Observation 531f0619-ed63-4ac3-820a-5e320462fff8 · outbound

This paper cites an unresolved cited work.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:56:47.821498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T14:56:35.348426Z digest=sha256:e141d982c7a550aaca4696ff463f781527b600f23b5c318d00cedbad2ac57e70

Observation d6229945-444b-4ae4-90a0-d372e6214676 · outbound

This paper cites Bc-z: Zero-shot task generalization with robotic imitation learning.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning Bc-z: Zero-shot task generalization with robotic imitation learning

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:56:47.631415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T14:56:35.506709Z digest=sha256:3555ff5b2194c0b061370abbf29c411495b71b4d5547ab0f6ec68c4ffee8a4b0

Observation b628af0f-f625-4924-acc0-de9089218dee · outbound

This paper cites Robotic grasp- ing using deep reinforcement learning.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning Robotic grasp- ing using deep reinforcement learning

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:56:47.412144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T14:56:35.676679Z digest=sha256:735a56830e21ef6bdd0d677ea994cf4c5c19431767d493eb5175128f62e3dbfa

Observation aecd2f10-6cbe-4424-907c-b452be976e7c · outbound

This paper cites DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:35.821488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:56:35.821488Z digest=sha256:9e27a3ca6d55bb7e5dbeadd4873a62ab236ba29c118d271db378833463d61084

Observation 6c8673aa-da7b-4db7-9ce5-7368e1f1a93b · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning OpenVLA: An Open-Source Vision-Language-Action Model

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:36.012538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:56:36.012538Z digest=sha256:1eaea2f7c71a9e47d467bd6b98d1d763a0d36bcdcd8c6af7bf8d006b6af6d001

Observation 17b4b429-82bc-484c-81d9-fa312b1c1061 · outbound

This paper cites Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:36.108597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:56:36.108597Z digest=sha256:bc446763f06cf4d5595e72c772a46ee2275dc508d5a21db9efecfa11e809c249

Observation da172c2c-1e2f-4231-9fb0-4479243c419b · outbound

This paper cites Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:36.234503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:56:36.234503Z digest=sha256:70a094e21bdd503d8287afadeced9ec331e0f25f1c7a2a90a75db3ff2c43508a

Observation 5e549f1e-a20d-4ea7-bd54-3d9136d0b7d7 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning Adam: A Method for Stochastic Optimization

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:36.351989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:56:36.351989Z digest=sha256:f60bad60890506e7d8116238dbe8d4c44be5a2f111fc3da6b52e04e844f1676b

Observation 897fb361-f246-4069-8dda-9c61961e321c · outbound

This paper cites Segment any- thing.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning Segment any- thing

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:36.485659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:56:36.485659Z digest=sha256:a38d78e4999f05c0ed6263cc90b368d6b70ef4bc7ce6f4e5bb5d1bf814a44685

Observation b90f2d53-e0af-4fc1-b1d3-47ac69419831 · outbound

This paper cites Tenenbaum.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning Tenenbaum

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:56:47.173564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T14:56:36.637966Z digest=sha256:a3ddc305e18a336d99ab3f70e144f57a31384ccae8016bd5e1077f264aaa183e

Observation e4f746b1-ce77-4f22-86d7-f8a30e27c75c · outbound

This paper cites CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:36.734418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:56:36.734418Z digest=sha256:bf536386cb744d26f992778961ca3b386ae666750fa703c2cb319d05e910e2ac

Observation 766028de-d46f-4126-9775-ffe1bf942bad · outbound

This paper cites Autoregressive image generation without vec- tor quantization.Advances in Neural Information Processing Systems, 37:56424–56445, 2024.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning Autoregressive image generation without vec- tor quantization.Advances in Neural Information Processing Systems, 37:56424–56445, 2024

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:36.858251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:56:36.858251Z digest=sha256:0ac7e1a4cdd8b9c68bb4b84f105c63dd02e9ed52abed7b7b9f98fe270553b3ff

Observation e06d4802-c6bb-46fa-89b9-e32f33860c88 · outbound

This paper cites CLAM: Continuous Latent Action Models for Robot Learning from Unlabeled Demonstrations.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning CLAM: Continuous Latent Action Models for Robot Learning from Unlabeled Demonstrations

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:36.948592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:56:36.948592Z digest=sha256:c50ecb243a1127f74cc8f6ce8fad369b66e78b8756523c7070be6ede9000fdda

Observation 1e14667b-ea42-4bdb-bd57-940576f72a41 · outbound

This paper cites Flow Matching for Generative Modeling.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning Flow Matching for Generative Modeling

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:37.093066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:56:37.093066Z digest=sha256:44dc0eddeffa658ae239439527e8c0d2854aa90b4123562b13edf6a377182552

Observation c1d3d381-2afc-4eae-bf5f-f94e1466912e · outbound

This paper cites Libero: Benchmarking knowl- edge transfer for lifelong robot learning.Advances in Neural Information Processing Systems, 36:44776–44791, 2023.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning Libero: Benchmarking knowl- edge transfer for lifelong robot learning.Advances in Neural Information Processing Systems, 36:44776–44791, 2023

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:56:46.867371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T14:56:37.223855Z digest=sha256:d93013b7caeecbbcfdce4765abc798ed257b936ae1683de3ddae9971e42c5c43

Observation b3a16e09-5265-4fce-a8f6-d80b1e4528fb · outbound

This paper cites HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:37.346550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:56:37.346550Z digest=sha256:936df0521c974b81b750d1dab196947ec17d958f44073bc69f646dba944b0b0e

Observation 91454c3f-fe8c-4128-9e8e-488ef8d54ee9 · outbound

This paper cites Rectified Flow: A Marginal Preserving Approach to Optimal Transport.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning Rectified Flow: A Marginal Preserving Approach to Optimal Transport

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:37.435792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:56:37.435792Z digest=sha256:55f3d6fbc498a917ca2ba7ae1bf2be5839e53c900de9c6c473da3cc928b16e88

Observation c016224b-a916-47ad-9b5a-5fbde84adce8 · outbound

This paper cites RDT-1B: a diffusion foundation model for bimanual manipulation.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning RDT-1B: a diffusion foundation model for bimanual manipulation

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:56:46.639358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T14:56:37.514294Z digest=sha256:6e9212b24ac98861f2b68b0d8e161d46621d640e9e44d2fee10a6a615d292d55

Observation 11329dda-5725-4c35-8729-918d64d90a99 · outbound

This paper cites Unified Embodied VLM Reasoning with Robotic Action via Autoregressive Discretized Pre-training.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning Unified Embodied VLM Reasoning with Robotic Action via Autoregressive Discretized Pre-training

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:37.623412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:56:37.623412Z digest=sha256:97699112873d1cbb0204e1a3513d2ca4a5d5be3fb9ae0d258f288ac0bc8a598b

Observation 4395234a-2a81-4578-aa68-10a761eca08c · outbound

This paper cites DPM-Solver++: Fast Solver for Guided Sampling of Diffusion Probabilistic Models.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning DPM-Solver++: Fast Solver for Guided Sampling of Diffusion Probabilistic Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:37.727000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:56:37.727000Z digest=sha256:8cf26df6cb24f10e245c9f593ec6210523e5b418ff56cfdf20b7f1820fa92441

Observation 69206639-e4b1-4a57-8de2-3699a78412ce · outbound

This paper cites Being-h0.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning Being-h0

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:37.804009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:56:37.804009Z digest=sha256:d31a6409f947944e1db3cfc7457878498cffd05ac9ecba674323e3c848985e07

Observation 22292ca0-7802-4340-b494-99ad30100678 · outbound

This paper cites Towards generalist robot learning from in- ternet video: A survey.Journal of Artificial Intelligence Re- search, 83, 2025.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning Towards generalist robot learning from in- ternet video: A survey.Journal of Artificial Intelligence Re- search, 83, 2025

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:56:46.438453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T14:56:37.898574Z digest=sha256:183f945e01fb4c636d82f1657bed5f7ea3ef849fe2468350d6f25a78c3b06a1c

Observation dd0e4a18-9733-4229-b570-dc9ec7d1cc63 · outbound

This paper cites Calvin: A benchmark for language- conditioned policy learning for long-horizon robot manip- ulation tasks.IEEE Robotics and Automation Letters, 7(3): 7327–7334, 2022.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning Calvin: A benchmark for language- conditioned policy learning for long-horizon robot manip- ulation tasks.IEEE Robotics and Automation Letters, 7(3): 7327–7334, 2022

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:56:46.271265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T14:56:37.995839Z digest=sha256:3f64d4edf3b8af494f6bd90e7d4630003fc16352ac1034fb0e496b9f1fc93ca0

Observation 4291300b-7f8b-49c4-a2ec-2e7e5d431785 · outbound

This paper cites Struc- tured world models from human videos.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning Struc- tured world models from human videos

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:56:46.080304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T14:56:38.233877Z digest=sha256:24acb26ffc0bd3d1279b6124380ad032d5a5254452bf5677cb76eb827f4c57dc

Observation 7162bf4f-ba51-4eb7-b278-ec4863ab93ea · outbound

This paper cites Robotwin: Dual-arm robot benchmark with generative digi- tal twins (early version).

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning Robotwin: Dual-arm robot benchmark with generative digi- tal twins (early version)

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:56:45.919268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T14:56:38.388917Z digest=sha256:52b7a4574d0d2588454200b038738f64480d8fe86a1c819e2c121dc255a60bab

Observation e8016e64-0167-46c4-b1b3-4fb732abbfaa · outbound

This paper cites Rt-affordance: Affordances are versatile intermediate rep- resentations for robot manipulation.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning Rt-affordance: Affordances are versatile intermediate rep- resentations for robot manipulation

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:38.520931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:56:38.520931Z digest=sha256:09a4389477576035d749da23a5560c9edd0942480f147f41c8d25a04fdb2c932

Observation 5ca2bb9f-76df-4146-ae5d-e1b022300281 · outbound

This paper cites Latent action learning requires supervision in the presence of distractors.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning Latent action learning requires supervision in the presence of distractors

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:56:45.739469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T14:56:38.635605Z digest=sha256:f6d1e89c63f932a3e083412e80686f80d26154611e8a795256ae4c3033c62d6f

Observation fc1438ea-4827-47a5-8a27-a5bf2bb2572e · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning Representation Learning with Contrastive Predictive Coding

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:38.805706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:56:38.805706Z digest=sha256:eda97258ce36a7ead731f0db0549214baeb7cffb688c32251c935f4590876b86

Observation e49b0b75-e49c-426a-98da-b2e5372db460 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning DINOv2: Learning Robust Visual Features without Supervision

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:38.957643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:56:38.957643Z digest=sha256:b33394030821491ed89ae5a2a7a07a3133f245995dc928854b933228fb40561a

Observation 403e9781-c3ab-46e6-b8ac-a526f42e6337 · outbound

This paper cites Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:39.094879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:56:39.094879Z digest=sha256:8e616a1492b5c2d938f44c48c53e133bc5e244e9d7e083005e557c5dc12bf4d4

Observation 20f02a57-d5b0-434b-8332-99e9539a6eda · outbound

This paper cites Scalable diffusion models with transformers.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning Scalable diffusion models with transformers

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:39.260337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:56:39.260337Z digest=sha256:902388a53ef1ccf32d520a01fe75cbddacf879fa06d7b3313868006759f936e5

Observation 132fd4aa-6801-4824-b057-cf858f0030ee · outbound

This paper cites FAST: Efficient Action Tokenization for Vision-Language-Action Models.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:39.356562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:56:39.356562Z digest=sha256:d0d138212e5e442db40c2b10810bd086abbb007d4904542db22c4c8dc9e96c78

Observation 2150779b-75d6-473c-a549-582a5500ee0a · outbound

This paper cites Improving language understanding by gen- erative pre-training.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning Improving language understanding by gen- erative pre-training

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:56:45.548114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T14:56:39.414860Z digest=sha256:bc163f8b67bc19a8cecee32afd659c8014674a5838e40b3d57fbe0816e36c636

Observation 7e1295c6-1dfe-4e3a-a64b-7d6e3b861745 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.Journal of machine learning research, 21(140):1–67, 2020.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning Exploring the limits of transfer learning with a unified text-to-text transformer.Journal of machine learning research, 21(140):1–67, 2020

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:56:45.360652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T14:56:39.519643Z digest=sha256:2fa042022049c31d714c41748c1d4a7b07b82a1b26b2db088c5aa5a7f384ca57

Observation d429a206-3785-4eba-9fd5-4058e2f94591 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning SAM 2: Segment Anything in Images and Videos

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:39.602403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:56:39.602403Z digest=sha256:2666a7c961c7645c232e51ee7228dfee6620b7e8a1fea9c66c2798b9874f5069

Observation ffea3315-c38c-4bc2-b5ef-c2902e437bf7 · outbound

This paper cites Motion before action: Diffusing object mo- tion as manipulation condition.IEEE Robotics and Automa- tion Letters, 2025.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning Motion before action: Diffusing object mo- tion as manipulation condition.IEEE Robotics and Automa- tion Letters, 2025

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:56:45.152200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T14:56:39.692573Z digest=sha256:78eb620a9f44cf978a99f8c337ac5226ae5fd0bb27d043c67bb006dab9cc5c33

Observation da5bdcaf-fb0d-4d2f-bc55-8f2b6902ab6f · outbound

This paper cites Behavioral Cloning from Observation.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning Behavioral Cloning from Observation

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:39.807343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:56:39.807343Z digest=sha256:5284c0ac13e573b47c428879155376288c937c1e6da0ba5159726f2aec3d8940

Observation 212b802a-5dfc-419d-8d5f-70d5c9340082 · outbound

This paper cites Neural discrete representation learning.Advances in neural information pro- cessing systems, 30, 2017.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning Neural discrete representation learning.Advances in neural information pro- cessing systems, 30, 2017

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:56:44.948701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T14:56:39.942079Z digest=sha256:f3eadbc6d2ed1ca131adb9e09d811043d89cbab1be4ac25c9ac938cf472ea26e

Observation f3cb3cee-3432-4b29-9d12-c3803e41c8d9 · outbound

This paper cites MimicPlay: Long-Horizon Imitation Learning by Watching Human Play.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning MimicPlay: Long-Horizon Imitation Learning by Watching Human Play

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:40.019748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:56:40.019748Z digest=sha256:9429cbea3ba534d73e653d7c8e36ffc20d8f6c7accec292dd37bb3f9d2720502

Observation 0e1ebdec-0fea-4c11-b809-ce33c88a7903 · outbound

This paper cites Tdn: Temporal difference networks for efficient action recog- nition.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning Tdn: Temporal difference networks for efficient action recog- nition

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:56:44.715870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T14:56:40.119419Z digest=sha256:4dc96508937880a5390a56cb9b571e82829702a0315af0f89f8479ccd336249a

Observation eba3de72-db0f-477b-ac34-c9d316dc5fd1 · outbound

This paper cites EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:40.227298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:56:40.227298Z digest=sha256:d454db30d48e6b0910c4d27816f66d38f9e9da8dbd4957dd6825a30d1d81ea5f

Observation 93929ffa-5587-4763-b82f-b365ee5c494e · outbound

This paper cites Vq-vla: Improving vision- language-action models via scaling vector-quantized action tokenizers.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning Vq-vla: Improving vision- language-action models via scaling vector-quantized action tokenizers

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:56:44.384737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T14:56:40.322261Z digest=sha256:929d2cd1781a464e4d7d8356b7b000177b7217f615a9639b155219b1c7332849

Observation 809b10e5-0f56-49df-b2f3-4c54e5a820d1 · outbound

This paper cites Any-point trajectory modeling for policy learning.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning Any-point trajectory modeling for policy learning

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:56:44.125895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T14:56:40.493789Z digest=sha256:bebf05539b61ef76dbf322cba3cfb9b876a1b63af04fc17c55f69854a16a75f5

Observation 38ddd823-3fd4-46c9-9381-1891ed181075 · outbound

This paper cites Tinyvla: Towards fast, data-efficient vision- language-action models for robotic manipulation.RAL,.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning Tinyvla: Towards fast, data-efficient vision- language-action models for robotic manipulation.RAL,

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:56:43.749630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T14:56:40.593389Z digest=sha256:929786858205f17cc7c401c03e97a02947b4bcd08014b26d0e648905154396cf

Observation b2761b91-1784-4d90-a111-4e3178c92092 · outbound

This paper cites Flow as the Cross-Domain Manipulation Interface.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning Flow as the Cross-Domain Manipulation Interface

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:40.706078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:56:40.706078Z digest=sha256:47eccb5fa9af88217d32cc1982a5d44725cc7b9eb5a456eecf84233d381d7263

Observation b8f439d0-6dd7-44b8-bde8-42a8b0dc0497 · outbound

This paper cites Spatiotemporal Predictive Pre-training for Robotic Motor Control.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning Spatiotemporal Predictive Pre-training for Robotic Motor Control

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:40.813754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:56:40.813754Z digest=sha256:37888c74b588f0bb2ac39d0a665dcc8a7290a5d77bdfa4cd3556af217c66c7c7

Observation c0221f6b-e15c-432a-a62d-c6e879b1df20 · outbound

This paper cites Transferring foundation models for generalizable robotic manipulation.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning Transferring foundation models for generalizable robotic manipulation

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:56:43.492257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T14:56:40.920821Z digest=sha256:cfdb35233f3b41bacd1db3ab6e063c0208f8c3b231f98ea3f141765121b661d4

Observation d992746c-b945-4f62-992a-34970f393fc3 · outbound

This paper cites Tra-moe: Learning trajectory pre- diction model from multiple domains for adaptive policy conditioning.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning Tra-moe: Learning trajectory pre- diction model from multiple domains for adaptive policy conditioning

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:56:43.165878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T14:56:41.000514Z digest=sha256:2b33ffcbdda7aa0f0ef8f81e8a63aada28203b0724b7bf41a9e4f250a51263a1

Observation 29df59bc-d9fd-4b6d-a80c-b8269c34710f · outbound

This paper cites EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:41.108636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:56:41.108636Z digest=sha256:5b1621d91957b39cb08b2ad4338c5876eeadfb26b278eb8b8faa6b3fa79131af

Observation 1a45bc75-9667-4dc9-a453-5efb935d7ec5 · outbound

This paper cites Latent action pretraining from videos.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning Latent action pretraining from videos

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:56:42.917516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T14:56:41.199862Z digest=sha256:5542d0e96d587f1b4d88cdecb7ed0f35fa45f1480840ba372eba73bad3284f96

Observation cdba56e2-5111-4ee8-a30b-44de3eb6b082 · outbound

This paper cites DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:41.284332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:56:41.284332Z digest=sha256:66f93e944c9f8fb9daca1294b9e7f02fdb58e9ecc06d13dffd02c15277ab9cb9

Observation 2c8ad561-3481-4112-965a-8bd66644f316 · outbound

This paper cites Zhao, Vikash Kumar, Sergey Levine, and Chelsea Finn.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning Zhao, Vikash Kumar, Sergey Levine, and Chelsea Finn

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:56:42.668221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T14:56:41.422555Z digest=sha256:526ac7509a7efe0ec7a29ab57b6ccdfbcf555d3fe04d3ad62a523048b93776bb

Observation d6f515ff-3bb3-4c0f-8819-ccb3a23225c0 · outbound

This paper cites Aether: Geometric-aware unified world modeling.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning Aether: Geometric-aware unified world modeling

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:56:42.479526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T14:56:41.529797Z digest=sha256:89e0943b1db250b4674deb530d745bb06ee230303555499ce20124326627242c

Observation a90562a2-a51c-4ed5-b952-f6c52278bd2a · outbound

This paper cites Rt-2: Vision-language-action models transfer web knowledge to robotic control.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning Rt-2: Vision-language-action models transfer web knowledge to robotic control

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:56:42.262283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T14:56:41.618867Z digest=sha256:51d4049938964a2ed1b114d342d5f038040e9bb29207af01889a269bd5863a80

Observation b6588c65-327e-4d39-85c5-d9660e10c3ad · outbound

This paper cites 3, 9, 10, 12, 13.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning 3, 9, 10, 12, 13

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:56:48.952101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T14:56:34.058846Z digest=sha256:9227999a20621599b1fe6b22ecc4a49bb6dc9287603b91234e9b36023342b01f

Pith citing papers

Observation 35f452a7-6957-459a-b4e1-943cec727dd7 · inbound

villa-X: Enhancing Latent Action Modeling in Vision-Language-Action Models cites this paper.

villa-X: Enhancing Latent Action Modeling in Vision-Language-Action Models CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-06-19T17:11:51.716122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T21:52:02.893886Z digest=sha256:cceb04e564a58f3c447fd73475d6d21dd30504a55c35bbd14b8e5211f39192a5

Observation 19f51061-b3a8-4216-a741-c974fe8f40c4 · inbound

Motus: A Unified Latent Action World Model cites this paper.

Motus: A Unified Latent Action World Model CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-06-19T17:11:51.716122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T18:44:36.636455Z digest=sha256:8846e1bcd8c532df3fd24bb90b5fa466b739605753b1dcf179a16f089f9b07e2

Observation a427a797-0c2d-43fc-9781-c1a7049a7834 · inbound

CoLA-Flow Policy: Temporally Coherent Imitation Learning via Continuous Latent Action Flow Matching for Robotic Manipulation cites this paper.

CoLA-Flow Policy: Temporally Coherent Imitation Learning via Continuous Latent Action Flow Matching for Robotic Manipulation CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-06-19T17:11:51.716122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T09:21:33.493584Z digest=sha256:ccfdc04a3be2793c274a98b9c7486e9f8e239dea3e66c6e6f5ca379027983c1b

Observation 549380b7-d68b-4b0a-8839-732c98d4436f · inbound

CoLA-Flow Policy: Temporally Coherent Imitation Learning via Continuous Latent Action Flow Matching for Robotic Manipulation cites this paper.

CoLA-Flow Policy: Temporally Coherent Imitation Learning via Continuous Latent Action Flow Matching for Robotic Manipulation CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-06-19T17:11:51.716122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-21T14:39:16.033010Z digest=sha256:103685174d723307724b5ecfbcd6dca4003608d4b1eccb9298b9dc2f455abea5

Observation 43183e4e-38a8-4b78-a3d5-e206e4bcabf8 · inbound

CoLA-Flow Policy: Temporally Coherent Imitation Learning via Continuous Latent Action Flow Matching for Robotic Manipulation cites this paper.

CoLA-Flow Policy: Temporally Coherent Imitation Learning via Continuous Latent Action Flow Matching for Robotic Manipulation CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T06:18:04.994485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:18:04.994485Z digest=sha256:41cdb1cd96212a64104c420cd8e1e08f4df1948e80b545a88508d51be0d47c0f

Observation 92e70c91-9fd9-4a23-ac3a-a09ce17febd5 · inbound

DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos cites this paper.

DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning

Reference 107

Resolution
verified exact
arxiv_id, observed 2026-06-19T17:11:51.716122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T17:02:33.997887Z digest=sha256:8ff4f6eca40d7d986526bb13bea78106f4d63a6aed7ebd22b94455533f6092f0

Observation 874a24f2-2312-4959-8be8-c15ee3bce7a9 · inbound

Lifting Unlabeled Internet-level Data for 3D Scene Understanding cites this paper.

Lifting Unlabeled Internet-level Data for 3D Scene Understanding CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning

Reference 113

Resolution
verified exact
arxiv_id, observed 2026-06-19T17:11:51.716122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T22:16:57.890955Z digest=sha256:d5381a5ab4df77876174f2fbaa6ea90e110c49f927d3dcd4e0715947c62521fa

Observation 7d4a54cc-9250-4d06-aab5-4d96e1867809 · inbound

GazeVLA: Learning Human Intention for Robotic Manipulation cites this paper.

GazeVLA: Learning Human Intention for Robotic Manipulation CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning

Reference 70

Resolution
metadata mismatch
arxiv_id, observed 2026-06-19T17:11:51.716122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T11:37:30.784513Z digest=sha256:f330f40760e18e157aaa6d801f6bb93144ce84de3abebddaa475197cc51a3776

Observation cce587a7-4eac-42e9-8db9-96e6f3b4f7d9 · inbound

Being-H0.7: A Latent World-Action Model from Egocentric Videos cites this paper.

Being-H0.7: A Latent World-Action Model from Egocentric Videos CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-06-19T17:11:51.716122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-09T20:48:01.461993Z digest=sha256:666b5a53faf0a7a46424feaa5a2f8390ae2fcebb3bdb227c6234d8cb9a99116d

Observation 93a0adec-ef9d-4b4c-b466-8e81a66d6128 · inbound

From Imagined Futures to Executable Actions: Mixture of Latent Actions for Robot Manipulation cites this paper.

From Imagined Futures to Executable Actions: Mixture of Latent Actions for Robot Manipulation CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-06-19T17:11:51.716122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T04:44:03.661688Z digest=sha256:da03bfd428805bea1013640aa5d3ea0cc7fa573f7e59f1dbeed983abc74a3728

Observation c073d9dd-2b6b-4df9-ac19-c5d2dfab073f · inbound

RotVLA: Rotational Latent Action for Vision-Language-Action Model cites this paper.

RotVLA: Rotational Latent Action for Vision-Language-Action Model CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-06-19T17:11:51.716122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-14T17:48:06.734816Z digest=sha256:b6fef6dcbfb77a837818bf207e0697f2c5f09ad4845d5478f1a2e4154effe116

Observation e664ebbc-fe12-4a04-9f07-670aef69ff51 · inbound

Mitigating State Aliasing in Vision-Language-Action Models via Inverse Dynamics Learning cites this paper.

Mitigating State Aliasing in Vision-Language-Action Models via Inverse Dynamics Learning CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-06-29T08:33:15.413551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-29T08:28:37.970450Z digest=sha256:44e86e59e86171f97847b07f0d7b25f15971672ddda2f419eb8bcf1386f1bafb

Observation 25fdcd7d-18ab-453b-b234-cf988e341c1f · inbound

CLAW: Learning Continuous Latent Action World Models via Adversarial Latent Regularization cites this paper.

CLAW: Learning Continuous Latent Action World Models via Adversarial Latent Regularization CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-07-02T03:36:29.609806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T09:55:00.402411Z digest=sha256:496241747064fb1d5ecaae180ff87c6f36b52790c19d2e9c5823ba74bf1d2633

Observation 8eeb06f6-8fdf-4ad8-b092-c9302b77841a · inbound

LAFP: Preserving Latent Action Structure in Latent Policy Learning via Flow Matching cites this paper.

LAFP: Preserving Latent Action Structure in Latent Policy Learning via Flow Matching CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-07-03T04:27:36.266068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-27T13:58:21.461547Z digest=sha256:6080b557f6328752be355d0a2d53a908ad7edd96ef7e47bcee5684e796bd006b

Observation 8baa9522-1109-4440-8095-57ac4bc1d5e1 · inbound

ShadowDancer: Teaching Video World Models Any Action by Learning Unified Dynamics Representations from a Video and Its Shadow cites this paper.

ShadowDancer: Teaching Video World Models Any Action by Learning Unified Dynamics Representations from a Video and Its Shadow CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning

Reference 60

Resolution
unresolved
no resolver link, observed 2026-07-31T09:42:45.914747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T09:42:45.914747Z digest=sha256:298d41c1731f4d01029f4e255eacba005bc19410d848e0e299050e387a7a24c3