Pith. sign in

Paper Citation Record · LEDGER

Masked Visual Actions for Unified World Modeling

As of 19 August 2026, this Paper Citation Record lists 94 of 94 outbound references and 1 inbound Pith citation observation for arXiv:2607.19343.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.19343 v1

Coverage vector

measured 94 of 94 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T12:49:09.834338Z

measured 95 of 95 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T23:21:53.430055Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T23:21:53.849642Z

Reference resolution

94 of 94 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved92
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e144fd38-614d-4d79-abd4-e49972c8b038 · outbound

This paper cites an unresolved cited work.

Masked Visual Actions for Unified World Modeling Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T12:48:57.894759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:48:57.894759Z digest=sha256:613bda2c5cf113ac6d6729b1c511907725e425aaae5c5daedb1c80cbae96911f

Observation e153cc02-ae9a-4090-93c7-4e9c44efb71b · outbound

This paper cites Unifying (Machine) Vision via Counterfactual World Modeling.

Masked Visual Actions for Unified World Modeling Unifying (Machine) Vision via Counterfactual World Modeling

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T12:48:58.164826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:48:58.164826Z digest=sha256:2ba49482024efc6e1179fae80fc5dba82949fd693d989e547988dcd74d841a37

Observation d43bb27a-7e9e-4a6c-8454-03d2bb8fd1d9 · outbound

This paper cites Track2act: Predicting point tracks from internet videos enables generalizable robot manipulation, 2024.

Masked Visual Actions for Unified World Modeling Track2act: Predicting point tracks from internet videos enables generalizable robot manipulation, 2024

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T12:48:58.324753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:48:58.324753Z digest=sha256:2c935f72308ca1feb4139d7e768f19e355781adc78789094ddc834e0a70b26c3

Observation c396244a-b33e-42f5-bf6f-64d27e5ee62d · outbound

This paper cites Video generation models as world simulators.

Masked Visual Actions for Unified World Modeling Video generation models as world simulators

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T12:48:58.544990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:48:58.544990Z digest=sha256:5c7b67041cbd33637d9a09c90741ffd573460a8a43176197196bba2e3c7f9b00

Observation 92e1cc1d-2170-4645-a044-ddedd9816a4f · outbound

This paper cites Go-with-the-flow: Motion-controllable video diffusion models using real-time warped noise.

Masked Visual Actions for Unified World Modeling Go-with-the-flow: Motion-controllable video diffusion models using real-time warped noise

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T12:48:58.718932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:48:58.718932Z digest=sha256:7182d7cdc4b664f4915a5eec587732c6fb3c252a3d87152d4c76dc53a369a165

Observation a06037db-a4bd-4f57-ba84-c9faefb0c543 · outbound

This paper cites SAM 3: Segment anything with concepts.

Masked Visual Actions for Unified World Modeling SAM 3: Segment anything with concepts

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T12:48:58.894834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:48:58.894834Z digest=sha256:588b0f6b446d6febbe0bca721499b055e9c6d39357fcedd6d8b7904851174509

Observation 7fe0aad0-dfcd-4a31-b296-27d8a4bf916c · outbound

This paper cites Freeman, Jitendra Malik, Russ Tedrake, Vincent Sitzmann, and Yilun Du.

Masked Visual Actions for Unified World Modeling Freeman, Jitendra Malik, Russ Tedrake, Vincent Sitzmann, and Yilun Du

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T12:48:59.064864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:48:59.064864Z digest=sha256:64ceb403f5f9c76b44af6e3a2b5702d53cf9b4bd9436a781414e43fbf15e63de

Observation 26bcfb5d-e03a-4987-8e76-c0ed0e00c3e9 · outbound

This paper cites Learning coordinated bimanual manipulation policies using state diffusion and inverse dynamics models.

Masked Visual Actions for Unified World Modeling Learning coordinated bimanual manipulation policies using state diffusion and inverse dynamics models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T12:48:59.254754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:48:59.254754Z digest=sha256:a02ad880c22a0a3830bfc0221414d1116dc06e5f59865a60e16e43e97d0c886d

Observation 43dbe581-b2c1-463d-9d58-6c9a81cf636a · outbound

This paper cites Tool-as-interface: Learning robot policies from observing human tool use.

Masked Visual Actions for Unified World Modeling Tool-as-interface: Learning robot policies from observing human tool use

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T12:48:59.497836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:48:59.497836Z digest=sha256:db2be7bac2dd50adbbae083c619d6f8e65ab869b2c9c1d0d0395238cf61aca5d

Observation bf9c6eef-c02e-458a-9ffa-3599ee0fdc41 · outbound

This paper cites Bridgev2w: Bridging video generation models to embodied world models via embodiment masks.arXiv preprint arXiv:2602.03793, 2026.

Masked Visual Actions for Unified World Modeling Bridgev2w: Bridging video generation models to embodied world models via embodiment masks.arXiv preprint arXiv:2602.03793, 2026

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T12:48:59.894801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:48:59.894801Z digest=sha256:53919bd9c081656fb389d24bd89f87d5163393b0032602566640ceee13955912

Observation 1702a4b4-9cc9-495d-96b9-1e3d7d5bd874 · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action diffusion.

Masked Visual Actions for Unified World Modeling Diffusion policy: Visuomotor policy learning via action diffusion

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:00.023922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:00.023922Z digest=sha256:b3b7181dd9eb710930ace235ee94da9608d6a3198a332a428e2d572491b54f6f

Observation ffcbb56b-0b93-4263-b34a-091724c31d4d · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action diffusion.

Masked Visual Actions for Unified World Modeling Diffusion policy: Visuomotor policy learning via action diffusion

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:00.102142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:00.102142Z digest=sha256:163f457c1eec9cf5edc332359c5efdb9e3184b3f20efb9c67820c65ec66c3a6d

Observation 918e5046-42ee-4c96-8962-40cdc95f6386 · outbound

This paper cites Wan-move: Motion-controllable video generation via latent trajectory guidance.arXiv preprint arXiv:2512.08765, 2025.

Masked Visual Actions for Unified World Modeling Wan-move: Motion-controllable video generation via latent trajectory guidance.arXiv preprint arXiv:2512.08765, 2025

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:00.184836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:00.184836Z digest=sha256:59090f32ba2e46ffefb27aa55bd35c071fdad42131eb930c47e9337ef57bd19d

Observation 5561a96c-a2c6-4294-87fa-ef81884764bc · outbound

This paper cites Embodis- wap for zero-shot robot imitation learning, 2025.

Masked Visual Actions for Unified World Modeling Embodis- wap for zero-shot robot imitation learning, 2025

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:00.324749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:00.324749Z digest=sha256:8906f0ecd640a273ca019e778e66823527195e6e0f383a057dc310fc9049d918

Observation 5747730c-af1d-45c0-8fff-4325a4508157 · outbound

This paper cites BERT: Pre-training of deep bidirectional transformers for language understanding.

Masked Visual Actions for Unified World Modeling BERT: Pre-training of deep bidirectional transformers for language understanding

Reference 15

Resolution
malformed identifier
no resolver link, observed 2026-08-01T12:49:00.475099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:00.475099Z digest=sha256:16665c7c1e32cfb3e5de3de108da7a88e7a553ab1d367d716dc0cc568cb200c7

Observation 8a7ed1bd-fb04-45ef-9469-912c526c3b19 · outbound

This paper cites Learning universal policies via text-guided video generation.Advances in neural information processing systems, 36:9156–9172, 2023.

Masked Visual Actions for Unified World Modeling Learning universal policies via text-guided video generation.Advances in neural information processing systems, 36:9156–9172, 2023

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:00.624831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:00.624831Z digest=sha256:9b0445ec6dfcf26152b89a64e51b45479751ff49b6e477ae57f259af9177cf89

Observation 49ffcd05-92a8-48b5-9401-c21ca9b33844 · outbound

This paper cites Video language planning.

Masked Visual Actions for Unified World Modeling Video language planning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:00.749493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:00.749493Z digest=sha256:99a5dfff72803f73a10309a8904477f3945792182bb09267644364d99ba8c73e

Observation f3f80707-501b-4a75-81ad-66c7d1ae9ecc · outbound

This paper cites Aim: Intent-aware unified world action modeling with spatial value maps, 2026.

Masked Visual Actions for Unified World Modeling Aim: Intent-aware unified world action modeling with spatial value maps, 2026

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:00.868780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:00.868780Z digest=sha256:10e71b97bb3ea9aa0f4fbd5132fa9bf9813569900f96f285573f2b4cd64c227c

Observation 218a1261-205c-49d1-b6d7-4c2ac1f53fab · outbound

This paper cites DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos.

Masked Visual Actions for Unified World Modeling DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:01.042267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:01.042267Z digest=sha256:6d01636a74d1555d5f3baef57dc2801a0467ec4c1a7c02926385a7eb3dbe60b8

Observation 9c26af47-6a88-4cbd-a16c-2d2cdd90d647 · outbound

This paper cites Motion prompting: Controlling video generation with motion trajectories.

Masked Visual Actions for Unified World Modeling Motion prompting: Controlling video generation with motion trajectories

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:01.159661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:01.159661Z digest=sha256:cd58e839560621ebdd3aba84adf1b1632dd1fefcfd5c7bbc674e3991628def92

Observation 36716550-3d81-459c-b6ac-995455b6fc26 · outbound

This paper cites Force prompting: Video generation models can learn and generalize physics-based control signals.

Masked Visual Actions for Unified World Modeling Force prompting: Video generation models can learn and generalize physics-based control signals

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:01.443907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:01.443907Z digest=sha256:80c97dd415295ca03bbbffb46d383bace97b30607728c167e0e3e63112e83a41

Observation f7e4cc6c-b7b1-4171-a1c9-16586381e5f6 · outbound

This paper cites Goal force: Teaching video models to accomplish physics-conditioned goals.

Masked Visual Actions for Unified World Modeling Goal force: Teaching video models to accomplish physics-conditioned goals

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:01.584937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:01.584937Z digest=sha256:d3f759853826e5230650196f9070f6a74687bd305aadc408dfe9599ac414f602

Observation 9cfc4af1-0d8f-4a72-934d-6001de896cfb · outbound

This paper cites World models for learning dexterous hand-object interactions from human videos.arXiv preprint arXiv:2512.13644, 2026.

Masked Visual Actions for Unified World Modeling World models for learning dexterous hand-object interactions from human videos.arXiv preprint arXiv:2512.13644, 2026

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:01.704765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:01.704765Z digest=sha256:2cbcbd5538362f71f597ced4c3bf401826cd306d2f84aca8d139c34d4f8f2ade

Observation c74a524d-0ea2-4d14-b1f0-f7a703a5db5b · outbound

This paper cites Unified 4d world action modeling from video priors with asynchronous denoising, 2026.

Masked Visual Actions for Unified World Modeling Unified 4d world action modeling from video priors with asynchronous denoising, 2026

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:01.844740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:01.844740Z digest=sha256:67e17141ff5b918a74a91a056d5c2b66254c342038807983e45b5dd3530c2726

Observation befed282-d539-46a9-8cfe-ff39cbb91804 · outbound

This paper cites Ctrl-world: A controllable generative world model for robot manipulation.

Masked Visual Actions for Unified World Modeling Ctrl-world: A controllable generative world model for robot manipulation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:02.035375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:02.035375Z digest=sha256:8ae107e307f47c32817f3f217c57bd94b6c9db87dadb9825fd042b551ce74c27

Observation d01eb595-33ff-422c-98e1-b41d4659e08c · outbound

This paper cites Video prediction policy: A generalist robot policy with predictive visual representations, 2024.

Masked Visual Actions for Unified World Modeling Video prediction policy: A generalist robot policy with predictive visual representations, 2024

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:02.110836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:02.110836Z digest=sha256:dbd62f4ee947a29c1926c8bf512ca9c4c3ba12ec62afaeca5be1512cab732699

Observation 7a64bf78-7d54-4a87-920f-a83b8a4d6bb0 · outbound

This paper cites Vid2world: Crafting video diffusion models to interactive world models, 2025.

Masked Visual Actions for Unified World Modeling Vid2world: Crafting video diffusion models to interactive world models, 2025

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:02.216476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:02.216476Z digest=sha256:f7e321594db419b31748b129c7740719111528ec1d03bc5bccde1b15352e8c2c

Observation 26522a78-c911-4297-b9e6-0d427f2dde09 · outbound

This paper cites Pointworld: Scaling 3d world models for in-the-wild robotic manipulation.

Masked Visual Actions for Unified World Modeling Pointworld: Scaling 3d world models for in-the-wild robotic manipulation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:02.362222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:02.362222Z digest=sha256:66853a2371ad1bdfc52c89fec6b605c9a2267f438d63bc31818dd6407680b2d1

Observation b4d59e25-62ab-47a1-82c2-9c5f5436eb9c · outbound

This paper cites Dreamgen: Unlocking generalization in robot learning through video world models, 2025.

Masked Visual Actions for Unified World Modeling Dreamgen: Unlocking generalization in robot learning through video world models, 2025

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:02.461021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:02.461021Z digest=sha256:45a99b137e9a35eb959160a84228a37ef2ea0acb51d32ecdfde70be607aa73c7

Observation 3929f513-5414-4dd7-a112-24134706778f · outbound

This paper cites Cotracker3: Simpler and better point tracking by pseudo-labelling real videos, 2024.

Masked Visual Actions for Unified World Modeling Cotracker3: Simpler and better point tracking by pseudo-labelling real videos, 2024

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:02.585802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:02.585802Z digest=sha256:6ee9946b5dcfa85d9096734c115c7150526590d947c54929a064f6fb701c6d91

Observation 9b1a1e0b-06b9-41dc-81a7-b2570c4306ce · outbound

This paper cites an unresolved cited work.

Masked Visual Actions for Unified World Modeling Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:02.742391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:02.742391Z digest=sha256:3a05ebfe1fb9ee212813a92e087451278e99d46c2ab87b0ceea1e687b97762ad

Observation 2f03d7b5-a577-4f6b-a1df-fc7892c172f2 · outbound

This paper cites Dexterous world models.

Masked Visual Actions for Unified World Modeling Dexterous world models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:02.843637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:02.843637Z digest=sha256:a5d8b2312505c9f7130e738cbd7800e7b0364d17575f63a63540bb51a63d6ac1

Observation 25ed980f-2583-4209-a8ce-bfa49fa23bc8 · outbound

This paper cites Cosmos policy: Fine-tuning video models for visuomotor control and planning, 2026.

Masked Visual Actions for Unified World Modeling Cosmos policy: Fine-tuning video models for visuomotor control and planning, 2026

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:02.952171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:02.952171Z digest=sha256:ad6535ac9a6fb7555f957a4ba70659baf3643920bf6a04d3d6067a43935ff40c

Observation d97b8c95-b61e-4c8d-854f-b13625b305e5 · outbound

This paper cites Learning to Act from Actionless Videos through Dense Correspondences.

Masked Visual Actions for Unified World Modeling Learning to Act from Actionless Videos through Dense Correspondences

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:03.044032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:03.044032Z digest=sha256:dae132137b62b3dd768e7a1785cb025b16c892639c106a08566c604340ed05f9

Observation 9fe16f9f-8f3c-4a56-86a0-ac38483916fd · outbound

This paper cites World Modeling with Probabilistic Structure Integration.

Masked Visual Actions for Unified World Modeling World Modeling with Probabilistic Structure Integration

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:03.136308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:03.136308Z digest=sha256:88b2f97a796fd15393f12f208ed9da1ae28faec225b661803c90b269d7426ac5

Observation 52ae5854-5dda-4b1d-8ecd-77042020424b · outbound

This paper cites Shadow: Leveraging segmentation masks for cross-embodiment policy transfer, 2025.

Masked Visual Actions for Unified World Modeling Shadow: Leveraging segmentation masks for cross-embodiment policy transfer, 2025

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:03.214635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:03.214635Z digest=sha256:ce7574cee4de38800d172b6186fc00fbe47a65abffa812cbff7faa8e017024ef

Observation a9afdf0e-3d47-4ea8-ac4a-73c7f631f9d4 · outbound

This paper cites Masquerade: Learning from in-the-wild human videos using data-editing, 2025.

Masked Visual Actions for Unified World Modeling Masquerade: Learning from in-the-wild human videos using data-editing, 2025

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:03.274841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:03.274841Z digest=sha256:801fe5077ecb9c3b112525a11c3e4fdb3184611cca3f32dc8a89b5f5701e267c

Observation a8ce6fbf-8a79-44b1-839b-0219aec38154 · outbound

This paper cites Phantom: Training robots without robots using only human videos, 2025.

Masked Visual Actions for Unified World Modeling Phantom: Training robots without robots using only human videos, 2025

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:03.364897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:03.364897Z digest=sha256:9329c2a6a254389dfecf8584c24f12f57770ffbafc539d78dbddb9904eef062e

Observation 14ccab1a-39b2-44ea-961e-92fbae02a8c0 · outbound

This paper cites BEHAVIOR-1K: A Human-Centered, Embodied AI Benchmark with 1,000 Everyday Activities and Realistic Simulation.

Masked Visual Actions for Unified World Modeling BEHAVIOR-1K: A Human-Centered, Embodied AI Benchmark with 1,000 Everyday Activities and Realistic Simulation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:03.472167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:03.472167Z digest=sha256:be86dac987bee4069606a20993c262b74db1c93204de134c7328a6cc7c8ee696

Observation 64a8cd3a-318b-48f5-a5bd-864f78f11963 · outbound

This paper cites Mask2iv: Interaction-centric video generation via mask trajectories, 2025.

Masked Visual Actions for Unified World Modeling Mask2iv: Interaction-centric video generation via mask trajectories, 2025

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:03.614750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:03.614750Z digest=sha256:feec75ca06e2f9f0ff465099d40a1b7c03918baa9cf191548cb8dde988f36565

Observation 00decb5e-3f9b-4070-bd75-a0d881ff15b5 · outbound

This paper cites Novaflow: Zero-shot manipulation via actionable flow from generated videos, 2025.

Masked Visual Actions for Unified World Modeling Novaflow: Zero-shot manipulation via actionable flow from generated videos, 2025

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:03.715409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:03.715409Z digest=sha256:01c34e7ba12acf234dafb7460797ed80f6ed0e0d7a7697dc7ba84c7175aa8e0f

Observation 2cd019cd-cb77-4e3b-8b63-0d3b591854f2 · outbound

This paper cites Unified video action model, 2025.

Masked Visual Actions for Unified World Modeling Unified video action model, 2025

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:03.844866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:03.844866Z digest=sha256:85c623cabad58c645a0cac6561eba9440350a99d9078aca8d0835cc65ebc366e

Observation e3a97ca1-f025-4352-8f6b-92a27f8b75c0 · outbound

This paper cites Genie envisioner: A unified world foundation platform for robotic manipulation, 2025.

Masked Visual Actions for Unified World Modeling Genie envisioner: A unified world foundation platform for robotic manipulation, 2025

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:03.969261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:03.969261Z digest=sha256:c90eeaff4fc63299d000960fb337a44cebd6b3a1d3a43474125fea04c08002e4

Observation 7074bee0-8bfb-406c-bb61-2d156196acaf · outbound

This paper cites Realwonder: Real-time physical action-conditioned video generation, 2026.

Masked Visual Actions for Unified World Modeling Realwonder: Real-time physical action-conditioned video generation, 2026

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:04.146958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:04.146958Z digest=sha256:2441693b567c2df0c9bddf2406f104bdff7284e74ac67020b70bb0a1cfef2adb

Observation aeda653a-28e4-49e1-8fb2-551c6dd612fb · outbound

This paper cites Zero-shot world models are developmentally efficient learners.arXiv e-prints, pages arXiv–2604, 2026.

Masked Visual Actions for Unified World Modeling Zero-shot world models are developmentally efficient learners.arXiv e-prints, pages arXiv–2604, 2026

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:04.313578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:04.313578Z digest=sha256:79028a2ccfa02ad2aac809ac6a20537a6b9cbf842c30953c06c60cef4167e87f

Observation 0a3626b9-18bd-4032-8342-7b75e430cd94 · outbound

This paper cites Decoupled Weight Decay Regularization.

Masked Visual Actions for Unified World Modeling Decoupled Weight Decay Regularization

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:04.440100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:04.440100Z digest=sha256:d448d628ca3aaeb49b86198bc4015faf06e8be85e618cd45fc66230e56615df0

Observation 53ca9292-a352-415c-867c-1d1d2f99feca · outbound

This paper cites Mask world model: Predicting what matters for robust robot policy learning, 2026.

Masked Visual Actions for Unified World Modeling Mask world model: Predicting what matters for robust robot policy learning, 2026

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:04.517419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:04.517419Z digest=sha256:fe4dd25a91a229f2d15ab18a7b9d25c343b83be1ad606d573e05497ebba3cb56

Observation 8d819bb7-73ec-48cb-a4fb-3a1113f2eabc · outbound

This paper cites Inference-time scaling for diffusion models beyond scaling denoising steps.

Masked Visual Actions for Unified World Modeling Inference-time scaling for diffusion models beyond scaling denoising steps

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:04.669230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:04.669230Z digest=sha256:e5c7695b967f9c35f6f62622cb710285ac2a80738f6cefb26837a93885dcba10

Observation 1797b1cc-201e-402b-903f-e5c8efb61ce4 · outbound

This paper cites MimicGen: A Data Generation System for Scalable Robot Learning using Human Demonstrations.

Masked Visual Actions for Unified World Modeling MimicGen: A Data Generation System for Scalable Robot Learning using Human Demonstrations

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:04.803449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:04.803449Z digest=sha256:9a0286bf9c1f9b6ec26f1a1b63b51cdd959867b1827fab204e0f9ff0c300a22a

Observation ba782044-7c09-43b3-ac79-7443413b524d · outbound

This paper cites Motubrain: An advanced world action model for robot control, 2026.

Masked Visual Actions for Unified World Modeling Motubrain: An advanced world action model for robot control, 2026

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:04.967163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:04.967163Z digest=sha256:dfe68986c9d344a2c10e71d7b4b9093b8b41bb73e4a5e79d7a07d6bee660dee9

Observation 9418a5d4-53e5-41d5-a965-49c5dfd9f8fc · outbound

This paper cites s1: Simple test-time scaling.

Masked Visual Actions for Unified World Modeling s1: Simple test-time scaling

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:05.141732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:05.141732Z digest=sha256:31e7da8251027ae0c19d94c1bc87069361853b6af8b249877e20afdc15ecc9a5

Observation b115d4ba-4cf2-4981-a1d7-77bdd298e8ce · outbound

This paper cites Robocasa365: A large-scale simulation framework for training and benchmarking generalist robots.

Masked Visual Actions for Unified World Modeling Robocasa365: A large-scale simulation framework for training and benchmarking generalist robots

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:05.213696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:05.213696Z digest=sha256:e12cd5df3e1a6bfeb75a87257d1996b0df48adbdea389c9e9cda4380868e2b06

Observation 66f9331b-073c-4c7f-b8e1-b45d435d366e · outbound

This paper cites Cosmos world foundation model platform for physical ai, 2025.

Masked Visual Actions for Unified World Modeling Cosmos world foundation model platform for physical ai, 2025

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:05.313907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:05.313907Z digest=sha256:9aa93cca01fc803da6b55b330dee0691806436b55651a40e5e66bb65912a612b

Observation 21d004a3-c5b7-49c1-8d44-8eee52cd6107 · outbound

This paper cites mimic-video: Video-action models for generalizable robot control beyond vlas, 2025.

Masked Visual Actions for Unified World Modeling mimic-video: Video-action models for generalizable robot control beyond vlas, 2025

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:05.491244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:05.491244Z digest=sha256:b2fc0c09c2761a1def50a2c10bce1302ee9f5cd17aecd681c026d1a33acac743

Observation dd72077b-bc3b-4629-93d2-a02b9ada1f64 · outbound

This paper cites Inference-time enhancement of generative robot policies via predictive world modeling.IEEE Robotics and Automation Letters, 2026.

Masked Visual Actions for Unified World Modeling Inference-time enhancement of generative robot policies via predictive world modeling.IEEE Robotics and Automation Letters, 2026

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:05.647880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:05.647880Z digest=sha256:3494a414c1eab2b9155d38c70e5d296292e17ec09cb2e236928edcd8cb90fac0

Observation 71b4d916-03dc-4871-b8a2-c00044d164e1 · outbound

This paper cites MotionStream: Real-Time Video Generation with Interactive Motion Controls.

Masked Visual Actions for Unified World Modeling MotionStream: Real-Time Video Generation with Interactive Motion Controls

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:05.770198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:05.770198Z digest=sha256:3449191fad2e34c8f8d31171b33cd8ba68c38d9ddbc9902a5978719c9f2275fd

Observation 1ec0c14c-6f66-4171-a489-5ad8b747535b · outbound

This paper cites SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics.

Masked Visual Actions for Unified World Modeling SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:05.914908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:05.914908Z digest=sha256:3017d34d39b091c42424587b68cacb2a1d225177f1306c4d161af1fdb0981d45

Observation 11d1fc99-ed4d-4540-8640-45f8b08d9725 · outbound

This paper cites Time-to-move: Training-free motion-controlled video generation via dual-clock denoising.

Masked Visual Actions for Unified World Modeling Time-to-move: Training-free motion-controlled video generation via dual-clock denoising

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:06.115706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:06.115706Z digest=sha256:65ab99cfb27dbaf3d9cc8c31ba4ca87b1c429491a943a86d809ab612599578e9

Observation 7af91829-ba8c-458f-a05f-24b492076649 · outbound

This paper cites Motion before action: Diffusing object motion as manipulation condition, 2024.

Masked Visual Actions for Unified World Modeling Motion before action: Diffusing object motion as manipulation condition, 2024

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:06.248811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:06.248811Z digest=sha256:78f15927f6833297acfdeaae53a6d20927efaca81bf74e7b2f7f1090d170a61a

Observation 5fcd5563-ee5d-4d2f-8558-01d98ef167a4 · outbound

This paper cites Evaluating gemini robotics policies in a veo world simulator, 2025.

Masked Visual Actions for Unified World Modeling Evaluating gemini robotics policies in a veo world simulator, 2025

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:06.418490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:06.418490Z digest=sha256:82adff80bba9c7005730a460155ea1e50883b89ab78dd1b1e0722a35cf887e3d

Observation 3bd46ac0-233f-4c9d-b740-f7251e10d84d · outbound

This paper cites Causal video models are data-efficient robot policy learners.Rhoda AI Blog, 2026.

Masked Visual Actions for Unified World Modeling Causal video models are data-efficient robot policy learners.Rhoda AI Blog, 2026

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:06.557565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:06.557565Z digest=sha256:b7cfe9fb9731008accabd294faf63cacc25d4f1265285558d53624778ecc632f

Observation eec87124-0133-4b11-8def-0642d0540e5a · outbound

This paper cites Understanding physical dynamics with counterfactual world modeling.

Masked Visual Actions for Unified World Modeling Understanding physical dynamics with counterfactual world modeling

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:06.659959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:06.659959Z digest=sha256:5223d2954e20aef423b0057be1e4861deda1e8d852a874eeb0bf244629317989

Observation 62940561-bdaf-4541-b98a-a626f1390d4b · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

Masked Visual Actions for Unified World Modeling Wan: Open and Advanced Large-Scale Video Generative Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:06.746750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:06.746750Z digest=sha256:899682d057f65548c3955f5de471aab2fde3ff7856af98b1cade0f5e16ced948

Observation 9b9c6d2d-e9a9-4417-8cb2-57c7a9b151d2 · outbound

This paper cites Eva: Aligning video world models with executable robot actions via inverse dynamics rewards, 2026.

Masked Visual Actions for Unified World Modeling Eva: Aligning video world models with executable robot actions via inverse dynamics rewards, 2026

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:06.915628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:06.915628Z digest=sha256:feb9735caf01fe94b9f4d0e8e14a03187d9c584fad849f4c881de3214a5ff853

Observation c3dcb25d-8b7a-455e-bf13-25ecb7dc3cbe · outbound

This paper cites Interactive world simulator for robot policy training and evaluation.arXiv preprint arXiv:2603.08546, 2026.

Masked Visual Actions for Unified World Modeling Interactive world simulator for robot policy training and evaluation.arXiv preprint arXiv:2603.08546, 2026

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:06.965342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:06.965342Z digest=sha256:4bc88a12ee0c3902d107428fb2e86ff765b5646a187c859a755ed9948da79395

Observation 741604b0-4a2d-4094-a57e-0a50123789e9 · outbound

This paper cites Precise action-to-video generation through visual action prompts.

Masked Visual Actions for Unified World Modeling Precise action-to-video generation through visual action prompts

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:07.047797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:07.047797Z digest=sha256:94899d224cebb0e5b07b85eba6b7356ac25bdfaa9aa8b5814965d2abffb1b76f

Observation f7cdd53a-fa44-48ef-960c-1a816e2d4875 · outbound

This paper cites Wolpert and J.

Masked Visual Actions for Unified World Modeling Wolpert and J

Reference 67

Resolution
verified exact
doi, observed 2026-08-01T12:53:42.293862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-01T12:49:07.189078Z digest=sha256:c7b837eed3d8958b547c02fc1e4100d6b9ee762dc47891e4c2d739793843cf1e

Observation 85aa13da-c5b9-4d46-801a-221179b3d2c5 · outbound

This paper cites Wolpert, Zoubin Ghahramani, and Michael I.

Masked Visual Actions for Unified World Modeling Wolpert, Zoubin Ghahramani, and Michael I

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:07.307228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:07.307228Z digest=sha256:3382aa26d455f29e0cc7e387279a395d0b5362c76caf2ed7bb07dd0cba3723a1

Observation cf8c8a53-9b1c-4fbf-9340-eae84e13fda8 · outbound

This paper cites Wolpert, R.

Masked Visual Actions for Unified World Modeling Wolpert, R

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:07.401627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:07.401627Z digest=sha256:74afe9fc7619ac8ef694ea61e40581363e8022bc2a4c8090f141437d2b6a425f

Observation 99715a3f-d73f-40a5-a1ab-86533f20e61c · outbound

This paper cites Sun, Ashley Neall, Tong Wu, Shengqu Cai, and Gordon Wetzstein.

Masked Visual Actions for Unified World Modeling Sun, Ashley Neall, Tong Wu, Shengqu Cai, and Gordon Wetzstein

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:07.439272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:07.439272Z digest=sha256:9cc0f13d5f8158af9b2f2d75ee5ab7c945a2918456bef899f02c996fbf303df0

Observation f5d9321b-9944-43d7-abc5-f5748404f8fb · outbound

This paper cites Kinema4d: Kinematic4d world modeling for spatiotemporal embodied simulation.arXiv preprint arXiv:2603.16669, 2026.

Masked Visual Actions for Unified World Modeling Kinema4d: Kinematic4d world modeling for spatiotemporal embodied simulation.arXiv preprint arXiv:2603.16669, 2026

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:07.518129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:07.518129Z digest=sha256:0312d23d0594b6d383162ffbbf85ace71257277c05d5a11f8ffdf1e8e6da9309

Observation 8cfdc1d7-e718-41b2-9505-ac8e1d8a3d17 · outbound

This paper cites RoboPanoptes: The All-Seeing Robot with Whole-body Dexterity.

Masked Visual Actions for Unified World Modeling RoboPanoptes: The All-Seeing Robot with Whole-body Dexterity

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:07.719051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:07.719051Z digest=sha256:93adf95bad114ece6725cd6ea0ee6c28566f02ef280c59363c5c6e3291da6797

Observation ab6ab918-592d-4345-85df-311e9e6c4ab7 · outbound

This paper cites Learning Interactive Real-World Simulators.

Masked Visual Actions for Unified World Modeling Learning Interactive Real-World Simulators

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:07.854932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:07.854932Z digest=sha256:70063efdc61fe718098a92344e5d57233c3333f05d28757acd51900e4bc62126

Observation 84c4a1d6-799e-470f-8c8a-c27e4f9caa3e · outbound

This paper cites Orv: 4d occupancy-centric robot video generation, 2025.

Masked Visual Actions for Unified World Modeling Orv: 4d occupancy-centric robot video generation, 2025

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:07.975262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:07.975262Z digest=sha256:a075713f9d074a34c57deb907f306308d4c396dc092be0ec8c1c7bda54e1d921

Observation a8e5f30f-c093-4a77-8085-dcaeba235004 · outbound

This paper cites Real-to-sim robot policy evaluation with gaussian splatting simulation of soft-body interactions.

Masked Visual Actions for Unified World Modeling Real-to-sim robot policy evaluation with gaussian splatting simulation of soft-body interactions

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:08.072132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:08.072132Z digest=sha256:23933c38fcff492f1f2bc978119d144ba5252187f99f44e4010e66b193c3bc6e

Observation 84f6297c-83a5-4b2f-ac76-2c38fd5fab04 · outbound

This paper cites Veo-act: How far can frontier video models advance generalizable robot manipulation?, 2026.

Masked Visual Actions for Unified World Modeling Veo-act: How far can frontier video models advance generalizable robot manipulation?, 2026

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:08.184918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:08.184918Z digest=sha256:51b58d1e37016e5a7bb595962b76972ba1611aba991361dfcc9961a4964ff1ce

Observation 7fccbe33-c8ae-4be6-a725-5274c7345b24 · outbound

This paper cites Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware.

Masked Visual Actions for Unified World Modeling Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:08.273965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:08.273965Z digest=sha256:26c1a2c120087ed88ce9fbc365b8e92302e7467a3dc7bbb6db2297b5d6e67840

Observation 6c349569-4e4e-44f4-a7b6-6acc18cc9ca9 · outbound

This paper cites TesserAct: Learning 4D Embodied World Models.

Masked Visual Actions for Unified World Modeling TesserAct: Learning 4D Embodied World Models

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:08.342623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:08.342623Z digest=sha256:7a0a30a710c20f47bbbe844300d19aa9a18c5633a9b1ef77610f02e86de32adc

Observation 8d9b496d-c710-48b7-8b9d-8339d9ea3ca8 · outbound

This paper cites Action images: End-to-end policy learning via multiview video generation, 2026.

Masked Visual Actions for Unified World Modeling Action images: End-to-end policy learning via multiview video generation, 2026

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:08.422235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:08.422235Z digest=sha256:1dabf1bf9fda16f00fcf3bdd92d32c260f5edcbe80a512f0fa2e0e1bf9030675

Observation d5264796-98d5-4805-88bd-ef0430366b8d · outbound

This paper cites 3dflowaction: Learning cross-embodiment manipulation from 3d flow world model, 2025.

Masked Visual Actions for Unified World Modeling 3dflowaction: Learning cross-embodiment manipulation from 3d flow world model, 2025

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:08.526653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:08.526653Z digest=sha256:e4607b6e377591470b8e17dfb6ec0af425e0c2d519e06261ca98c9ea2f8ff656

Observation 02ec8abe-6c1a-4de1-8892-851c0f196470 · outbound

This paper cites Unified world models: Coupling video and action diffusion for pretraining on large robotic datasets, 2025.

Masked Visual Actions for Unified World Modeling Unified world models: Coupling video and action diffusion for pretraining on large robotic datasets, 2025

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:08.630493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:08.630493Z digest=sha256:a720489113367ecf112678eba4196aa2412c5c4615d10071a498e7342329c5e0

Observation 89f044fc-98f3-4c8f-ace7-cb3e24366785 · outbound

This paper cites Irasim: Learning interactive real-robot action simulators.

Masked Visual Actions for Unified World Modeling Irasim: Learning interactive real-robot action simulators

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:08.710108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:08.710108Z digest=sha256:86ac1156c1cac8273042ec028bcf6bb573a03cd82549dd58e39af7076cc5e810

Observation 5ab4aa41-44c7-4018-a44f-70630b5e1320 · outbound

This paper cites an unresolved cited work.

Masked Visual Actions for Unified World Modeling Unresolved cited work

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:08.889648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:08.889648Z digest=sha256:29acde75d08115807c1a4808999745276455376082d31b1e572d71b3e2b35876

Observation bf97c2d8-397f-4f77-9d92-8dc248a27450 · outbound

This paper cites an unresolved cited work.

Masked Visual Actions for Unified World Modeling Unresolved cited work

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:09.033921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:09.033921Z digest=sha256:b6f9a38bea71fa2f45a4bec22dcd440b04dce4948642ff8445f3fdeae46bc8d9

Observation 9b682be1-bf24-422b-a688-e78410cff85f · outbound

This paper cites an unresolved cited work.

Masked Visual Actions for Unified World Modeling Unresolved cited work

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:09.150745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:09.150745Z digest=sha256:ef91fee63368b29af381b9e4bf29a1d4a9ef68eaeda153b7608d9b941dccabe6

Observation fe024360-1b51-40aa-a64c-a06c41f6aca3 · outbound

This paper cites an unresolved cited work.

Masked Visual Actions for Unified World Modeling Unresolved cited work

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:09.279685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:09.279685Z digest=sha256:d6881d8da4dd81c5b67a4d3051cfa31fda3564ff5d7e3b763040a48c5a1ba41c

Observation 95f7312e-19d5-42d6-94c0-7ad8072974d4 · outbound

This paper cites Covers Diffusion Policy, ACT, and SmolVLA baselines.

Masked Visual Actions for Unified World Modeling Covers Diffusion Policy, ACT, and SmolVLA baselines

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:09.427171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:09.427171Z digest=sha256:8b0982715ac4af8fdc652427c22796c289e3e4b976c2e7d91e342ea8e81365c4

Observation 21b2c77b-ec9e-4e58-ab39-bdd522ca28a7 · outbound

This paper cites an unresolved cited work.

Masked Visual Actions for Unified World Modeling Unresolved cited work

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:09.512352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:09.512352Z digest=sha256:4114dff0b150afe6256f9535fec9a03c2ab2cbb815ab53674784a0e168002193

Observation 40ca33fc-e5b8-44ee-b02f-2fe35df2e30d · outbound

This paper cites an unresolved cited work.

Masked Visual Actions for Unified World Modeling Unresolved cited work

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:09.557487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:09.557487Z digest=sha256:6ffb0a2da6538cab3d8039f639f9494e0412c0d0a2e48d1beaaf67234e036355

Observation 2d75b15b-efbb-4d84-bf59-8a6507ff0f7c · outbound

This paper cites the toaster door must be flush, latched, and closed by a push from below.

Masked Visual Actions for Unified World Modeling the toaster door must be flush, latched, and closed by a push from below

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:09.642077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:09.642077Z digest=sha256:912a660c67925ccd1c18fc3c5866409472ac2516c09804255eb8233b616987c2

Observation ddfca20e-1655-4c33-ac83-8b263a5b037e · outbound

This paper cites robot_pushed.

Masked Visual Actions for Unified World Modeling robot_pushed

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:09.706619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:09.706619Z digest=sha256:fc555669d58af41dda6a4f689a78714e122b24739f7dbe05d9564494d2ccea59

Observation fdc01ae2-3c49-4693-b456-1c166504bdcb · outbound

This paper cites Outcomes reached after disengagement, via ghost contact, teleport, or autonomous motion are FALSE.

Masked Visual Actions for Unified World Modeling Outcomes reached after disengagement, via ghost contact, teleport, or autonomous motion are FALSE

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:09.751858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:09.751858Z digest=sha256:eedeeac39a507757872bb96efda90bcb20a2396c3caaad8f71cbd43511c566e5

Observation a6af23b8-78ba-4f05-af0b-73a5d0d03040 · outbound

This paper cites Hovering near.

Masked Visual Actions for Unified World Modeling Hovering near

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:09.789865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:09.789865Z digest=sha256:293011d2ffb7b848c162fadc6070e8e8b28be695f76986be4dbc1f627cee2eac

Observation 7ab1c18a-7e90-414e-93a6-5b710927f094 · outbound

This paper cites Penalize teleporting, morphing, gripper passing through solids, vanishing or duplicated objects, frame-to-frame 18 jumps, ghost contact, post-disengagement coasting.

Masked Visual Actions for Unified World Modeling Penalize teleporting, morphing, gripper passing through solids, vanishing or duplicated objects, frame-to-frame 18 jumps, ghost contact, post-disengagement coasting

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:09.834338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:09.834338Z digest=sha256:06b406879475b2777d0e99bbe97447447106f93ff11cc9c6f45682849dc3383c

Pith citing papers

Observation 106f1d63-784b-4ced-a55e-ba457d3205a3 · inbound

XEWorld: Can Action-Conditioned World Models Generalize to Unseen Robot Embodiments? cites this paper.

XEWorld: Can Action-Conditioned World Models Generalize to Unseen Robot Embodiments? Masked Visual Actions for Unified World Modeling

Reference 29

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T23:21:53.853473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T23:21:53.430055Z digest=sha256:65bfe0236e026ae2f2a12e1a0c2b92bada325d910ed286b56096e8da01849968