Pith. sign in

Paper Citation Record · LEDGER

Masked Visual Actions for Unified World Modeling

As of 8 August 2026, this Paper Citation Record lists 94 of 94 outbound references and 1 inbound Pith citation observation for arXiv:2607.19343.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.19343 v1

Coverage vector

measured 94 of 94 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T12:49:09.834338Z

measured 95 of 95 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T23:21:53.430055Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T23:21:53.849642Z

Reference resolution

94 of 94 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved92
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e144fd38-614d-4d79-abd4-e49972c8b038 · outbound

This paper cites an unresolved cited work.

Masked Visual Actions for Unified World Modeling Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T12:48:57.894759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:48:57.894759Z digest=sha256:584c267d938929f85e3ee5c7da2b7317804fbea4c1f14103a9823ea46d6accd9

Observation e153cc02-ae9a-4090-93c7-4e9c44efb71b · outbound

This paper cites Unifying (Machine) Vision via Counterfactual World Modeling.

Masked Visual Actions for Unified World Modeling Unifying (Machine) Vision via Counterfactual World Modeling

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T12:48:58.164826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:48:58.164826Z digest=sha256:45661192733043cb7adfc931fb14edc7930894db82fbe226d80f7124a33a0e3e

Observation d43bb27a-7e9e-4a6c-8454-03d2bb8fd1d9 · outbound

This paper cites Track2act: Predicting point tracks from internet videos enables generalizable robot manipulation, 2024.

Masked Visual Actions for Unified World Modeling Track2act: Predicting point tracks from internet videos enables generalizable robot manipulation, 2024

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T12:48:58.324753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:48:58.324753Z digest=sha256:89fd923e3291961c10f93b7f06b352f7032a47b71767995b2f30be905de5e8cf

Observation c396244a-b33e-42f5-bf6f-64d27e5ee62d · outbound

This paper cites Video generation models as world simulators.

Masked Visual Actions for Unified World Modeling Video generation models as world simulators

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T12:48:58.544990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:48:58.544990Z digest=sha256:8a94f084ea2bd8885a030b6c138fd098c8d5467492b52855990fe44a6c07d020

Observation 92e1cc1d-2170-4645-a044-ddedd9816a4f · outbound

This paper cites Go-with-the-flow: Motion-controllable video diffusion models using real-time warped noise.

Masked Visual Actions for Unified World Modeling Go-with-the-flow: Motion-controllable video diffusion models using real-time warped noise

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T12:48:58.718932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:48:58.718932Z digest=sha256:70899ea02bd209356c74e14cd6c27aa9163285338aeecd207efbd97ea8720c55

Observation a06037db-a4bd-4f57-ba84-c9faefb0c543 · outbound

This paper cites SAM 3: Segment anything with concepts.

Masked Visual Actions for Unified World Modeling SAM 3: Segment anything with concepts

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T12:48:58.894834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:48:58.894834Z digest=sha256:8117172fa7fef586be2fa776adf4614809e349a00573cd87e602b14a02f5d847

Observation 7fe0aad0-dfcd-4a31-b296-27d8a4bf916c · outbound

This paper cites Freeman, Jitendra Malik, Russ Tedrake, Vincent Sitzmann, and Yilun Du.

Masked Visual Actions for Unified World Modeling Freeman, Jitendra Malik, Russ Tedrake, Vincent Sitzmann, and Yilun Du

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T12:48:59.064864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:48:59.064864Z digest=sha256:6f4d7e6d231602b544e32edbec75381c5d2d6924ec21c02d58be176577f81a5d

Observation 26bcfb5d-e03a-4987-8e76-c0ed0e00c3e9 · outbound

This paper cites Learning coordinated bimanual manipulation policies using state diffusion and inverse dynamics models.

Masked Visual Actions for Unified World Modeling Learning coordinated bimanual manipulation policies using state diffusion and inverse dynamics models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T12:48:59.254754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:48:59.254754Z digest=sha256:36fa52f8a6bab2dddc4d1ad5dea8e8c392250bbaf3e40409eeb8fbeb525ff66d

Observation 43dbe581-b2c1-463d-9d58-6c9a81cf636a · outbound

This paper cites Tool-as-interface: Learning robot policies from observing human tool use.

Masked Visual Actions for Unified World Modeling Tool-as-interface: Learning robot policies from observing human tool use

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T12:48:59.497836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:48:59.497836Z digest=sha256:8e0c823c82c463702b75a99f975d105f82a563c096d4d6a3c9c3df902c65200c

Observation bf9c6eef-c02e-458a-9ffa-3599ee0fdc41 · outbound

This paper cites Bridgev2w: Bridging video generation models to embodied world models via embodiment masks.arXiv preprint arXiv:2602.03793, 2026.

Masked Visual Actions for Unified World Modeling Bridgev2w: Bridging video generation models to embodied world models via embodiment masks.arXiv preprint arXiv:2602.03793, 2026

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T12:48:59.894801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:48:59.894801Z digest=sha256:09bd59e0580d7e3ca63152d7c039e7260abc787ae102c62953e3d2961aff9a7d

Observation 1702a4b4-9cc9-495d-96b9-1e3d7d5bd874 · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action diffusion.

Masked Visual Actions for Unified World Modeling Diffusion policy: Visuomotor policy learning via action diffusion

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:00.023922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:00.023922Z digest=sha256:8065aee98a771cbf1d4a808b9a94f793fc23b0297ae7c61842289a7ef642c1c8

Observation ffcbb56b-0b93-4263-b34a-091724c31d4d · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action diffusion.

Masked Visual Actions for Unified World Modeling Diffusion policy: Visuomotor policy learning via action diffusion

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:00.102142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:00.102142Z digest=sha256:8a6dd0498a9eb47c9b1c8acfe48c01aed8b6edc4a86c1f74278912de49cac195

Observation 918e5046-42ee-4c96-8962-40cdc95f6386 · outbound

This paper cites Wan-move: Motion-controllable video generation via latent trajectory guidance.arXiv preprint arXiv:2512.08765, 2025.

Masked Visual Actions for Unified World Modeling Wan-move: Motion-controllable video generation via latent trajectory guidance.arXiv preprint arXiv:2512.08765, 2025

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:00.184836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:00.184836Z digest=sha256:0ef2bd443f5705c66fffd096260b3a2c8465caed841771303905b18dfd8020ae

Observation 5561a96c-a2c6-4294-87fa-ef81884764bc · outbound

This paper cites Embodis- wap for zero-shot robot imitation learning, 2025.

Masked Visual Actions for Unified World Modeling Embodis- wap for zero-shot robot imitation learning, 2025

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:00.324749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:00.324749Z digest=sha256:e8da2606e8013629387d9b43c78d4292490dadb29780df76566e6f668a2dcdf9

Observation 5747730c-af1d-45c0-8fff-4325a4508157 · outbound

This paper cites BERT: Pre-training of deep bidirectional transformers for language understanding.

Masked Visual Actions for Unified World Modeling BERT: Pre-training of deep bidirectional transformers for language understanding

Reference 15

Resolution
malformed identifier
no resolver link, observed 2026-08-01T12:49:00.475099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:00.475099Z digest=sha256:4408634a30e565aa3742876c961a107620eab5e563c0cd23d5c8b05db66678c9

Observation 8a7ed1bd-fb04-45ef-9469-912c526c3b19 · outbound

This paper cites Learning universal policies via text-guided video generation.Advances in neural information processing systems, 36:9156–9172, 2023.

Masked Visual Actions for Unified World Modeling Learning universal policies via text-guided video generation.Advances in neural information processing systems, 36:9156–9172, 2023

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:00.624831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:00.624831Z digest=sha256:8fc342833243824527b212e36afd6baee8c76c7849eb8b1650a556fa90dc650a

Observation 49ffcd05-92a8-48b5-9401-c21ca9b33844 · outbound

This paper cites Video language planning.

Masked Visual Actions for Unified World Modeling Video language planning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:00.749493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:00.749493Z digest=sha256:be06fc41e451f4a0d72e58555e8bb1723cf4b7c725a75fbc1ab62bfad2a07443

Observation f3f80707-501b-4a75-81ad-66c7d1ae9ecc · outbound

This paper cites Aim: Intent-aware unified world action modeling with spatial value maps, 2026.

Masked Visual Actions for Unified World Modeling Aim: Intent-aware unified world action modeling with spatial value maps, 2026

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:00.868780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:00.868780Z digest=sha256:6b79b8d84a3f25cc47844b963f13d19618c94e378a9a0c15435a1d7c085aa343

Observation 218a1261-205c-49d1-b6d7-4c2ac1f53fab · outbound

This paper cites DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos.

Masked Visual Actions for Unified World Modeling DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:01.042267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:01.042267Z digest=sha256:56644fb29383b446a6bfd9165a1cb555d4b94b17b406b934d6146a195a0290af

Observation 9c26af47-6a88-4cbd-a16c-2d2cdd90d647 · outbound

This paper cites Motion prompting: Controlling video generation with motion trajectories.

Masked Visual Actions for Unified World Modeling Motion prompting: Controlling video generation with motion trajectories

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:01.159661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:01.159661Z digest=sha256:8f6d664286dac2ecd444d7b9a10a909ecd5f2c4a0b4ef26d766cebcec9cef57d

Observation 36716550-3d81-459c-b6ac-995455b6fc26 · outbound

This paper cites Force prompting: Video generation models can learn and generalize physics-based control signals.

Masked Visual Actions for Unified World Modeling Force prompting: Video generation models can learn and generalize physics-based control signals

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:01.443907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:01.443907Z digest=sha256:e9d18c0c3c81752276fc18a52399b5740f330dcac9d2b51c177d6830c922f079

Observation f7e4cc6c-b7b1-4171-a1c9-16586381e5f6 · outbound

This paper cites Goal force: Teaching video models to accomplish physics-conditioned goals.

Masked Visual Actions for Unified World Modeling Goal force: Teaching video models to accomplish physics-conditioned goals

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:01.584937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:01.584937Z digest=sha256:bc6782f0e669859053ab88d843f5ae21c22faa3efaf96314505373282b93de52

Observation 9cfc4af1-0d8f-4a72-934d-6001de896cfb · outbound

This paper cites World models for learning dexterous hand-object interactions from human videos.arXiv preprint arXiv:2512.13644, 2026.

Masked Visual Actions for Unified World Modeling World models for learning dexterous hand-object interactions from human videos.arXiv preprint arXiv:2512.13644, 2026

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:01.704765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:01.704765Z digest=sha256:5a3af3f42b98f418ea61be2d3dfa824578092243eb19f4b276fedd241cdc2826

Observation c74a524d-0ea2-4d14-b1f0-f7a703a5db5b · outbound

This paper cites Unified 4d world action modeling from video priors with asynchronous denoising, 2026.

Masked Visual Actions for Unified World Modeling Unified 4d world action modeling from video priors with asynchronous denoising, 2026

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:01.844740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:01.844740Z digest=sha256:60a653dffc4511f4aa5af439895643abe1cff61691931def6fd5b0c8d7435b2b

Observation befed282-d539-46a9-8cfe-ff39cbb91804 · outbound

This paper cites Ctrl-world: A controllable generative world model for robot manipulation.

Masked Visual Actions for Unified World Modeling Ctrl-world: A controllable generative world model for robot manipulation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:02.035375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:02.035375Z digest=sha256:fe80ef7e2161eebe7a6eb24425c4656b88b7aa5116221551820c36f1ed0e6f07

Observation d01eb595-33ff-422c-98e1-b41d4659e08c · outbound

This paper cites Video prediction policy: A generalist robot policy with predictive visual representations, 2024.

Masked Visual Actions for Unified World Modeling Video prediction policy: A generalist robot policy with predictive visual representations, 2024

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:02.110836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:02.110836Z digest=sha256:04caaedbb1d4c5620175102de04a154135a6bb340ba643c908bface8ce459209

Observation 7a64bf78-7d54-4a87-920f-a83b8a4d6bb0 · outbound

This paper cites Vid2world: Crafting video diffusion models to interactive world models, 2025.

Masked Visual Actions for Unified World Modeling Vid2world: Crafting video diffusion models to interactive world models, 2025

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:02.216476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:02.216476Z digest=sha256:42df5918e73f639ac2526a066688c890c3c063ecc32b13c7c1e1436840c69a20

Observation 26522a78-c911-4297-b9e6-0d427f2dde09 · outbound

This paper cites Pointworld: Scaling 3d world models for in-the-wild robotic manipulation.

Masked Visual Actions for Unified World Modeling Pointworld: Scaling 3d world models for in-the-wild robotic manipulation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:02.362222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:02.362222Z digest=sha256:3896413b1f041c6c750003eac9bee421dd76df12bd04aed5dacff6357db91fa4

Observation b4d59e25-62ab-47a1-82c2-9c5f5436eb9c · outbound

This paper cites Dreamgen: Unlocking generalization in robot learning through video world models, 2025.

Masked Visual Actions for Unified World Modeling Dreamgen: Unlocking generalization in robot learning through video world models, 2025

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:02.461021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:02.461021Z digest=sha256:9b2516574e895b1261ad26024964d60cf5ceea389bcd53182f6e485d67fdfa1c

Observation 3929f513-5414-4dd7-a112-24134706778f · outbound

This paper cites Cotracker3: Simpler and better point tracking by pseudo-labelling real videos, 2024.

Masked Visual Actions for Unified World Modeling Cotracker3: Simpler and better point tracking by pseudo-labelling real videos, 2024

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:02.585802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:02.585802Z digest=sha256:1607c2f32c1526bbad5d54e7c813080b3e13158dbe53cc90f618f27197ee46df

Observation 9b1a1e0b-06b9-41dc-81a7-b2570c4306ce · outbound

This paper cites an unresolved cited work.

Masked Visual Actions for Unified World Modeling Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:02.742391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:02.742391Z digest=sha256:8fec6ee42142c95f649a45f70e4bd133fc5c02546b93166aad319c38446ec8fa

Observation 2f03d7b5-a577-4f6b-a1df-fc7892c172f2 · outbound

This paper cites Dexterous world models.

Masked Visual Actions for Unified World Modeling Dexterous world models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:02.843637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:02.843637Z digest=sha256:7fc0284c136b07d34db845e5818782767ec2c67dcd7668c39e6f756ed42878ec

Observation 25ed980f-2583-4209-a8ce-bfa49fa23bc8 · outbound

This paper cites Cosmos policy: Fine-tuning video models for visuomotor control and planning, 2026.

Masked Visual Actions for Unified World Modeling Cosmos policy: Fine-tuning video models for visuomotor control and planning, 2026

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:02.952171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:02.952171Z digest=sha256:e0b14955dfc887fce7e5f1244eef019373001627ef0396c5da1f66497d2d48b8

Observation d97b8c95-b61e-4c8d-854f-b13625b305e5 · outbound

This paper cites Learning to Act from Actionless Videos through Dense Correspondences.

Masked Visual Actions for Unified World Modeling Learning to Act from Actionless Videos through Dense Correspondences

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:03.044032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:03.044032Z digest=sha256:a3ed1270ad08bba317dfbe9559d04eb9704988f44a90f127ae8d9b0d56b005bb

Observation 9fe16f9f-8f3c-4a56-86a0-ac38483916fd · outbound

This paper cites World Modeling with Probabilistic Structure Integration.

Masked Visual Actions for Unified World Modeling World Modeling with Probabilistic Structure Integration

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:03.136308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:03.136308Z digest=sha256:832fc40a637dee2f2bc78461b234617a3dbec549e347be82815881a455755804

Observation 52ae5854-5dda-4b1d-8ecd-77042020424b · outbound

This paper cites Shadow: Leveraging segmentation masks for cross-embodiment policy transfer, 2025.

Masked Visual Actions for Unified World Modeling Shadow: Leveraging segmentation masks for cross-embodiment policy transfer, 2025

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:03.214635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:03.214635Z digest=sha256:c19165ca0a9c310acef344acfd2186ff6ed653cb59b850d87fa11fe046a37d8b

Observation a9afdf0e-3d47-4ea8-ac4a-73c7f631f9d4 · outbound

This paper cites Masquerade: Learning from in-the-wild human videos using data-editing, 2025.

Masked Visual Actions for Unified World Modeling Masquerade: Learning from in-the-wild human videos using data-editing, 2025

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:03.274841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:03.274841Z digest=sha256:f1687e46c99c927e061fdf07cb1aba2a906b80f4ea329763bdccf6656157196a

Observation a8ce6fbf-8a79-44b1-839b-0219aec38154 · outbound

This paper cites Phantom: Training robots without robots using only human videos, 2025.

Masked Visual Actions for Unified World Modeling Phantom: Training robots without robots using only human videos, 2025

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:03.364897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:03.364897Z digest=sha256:66560feac685723fca3c4b8624e4347fd8f67e8eebc9fb50923dd73a894c9ae9

Observation 14ccab1a-39b2-44ea-961e-92fbae02a8c0 · outbound

This paper cites BEHAVIOR-1K: A Human-Centered, Embodied AI Benchmark with 1,000 Everyday Activities and Realistic Simulation.

Masked Visual Actions for Unified World Modeling BEHAVIOR-1K: A Human-Centered, Embodied AI Benchmark with 1,000 Everyday Activities and Realistic Simulation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:03.472167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:03.472167Z digest=sha256:fba188e49dba025ae40dbbb81dafc107cc3dff56ef7f3c16c4145850fbb347ee

Observation 64a8cd3a-318b-48f5-a5bd-864f78f11963 · outbound

This paper cites Mask2iv: Interaction-centric video generation via mask trajectories, 2025.

Masked Visual Actions for Unified World Modeling Mask2iv: Interaction-centric video generation via mask trajectories, 2025

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:03.614750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:03.614750Z digest=sha256:7b6bbba6d3fcc248ce59848957491583804e0baa236fccbd391cba6a7ecfd09c

Observation 00decb5e-3f9b-4070-bd75-a0d881ff15b5 · outbound

This paper cites Novaflow: Zero-shot manipulation via actionable flow from generated videos, 2025.

Masked Visual Actions for Unified World Modeling Novaflow: Zero-shot manipulation via actionable flow from generated videos, 2025

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:03.715409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:03.715409Z digest=sha256:13e198519a25ec23f560875cd95756861b24ef32100ef826af20d8041cc5c5e8

Observation 2cd019cd-cb77-4e3b-8b63-0d3b591854f2 · outbound

This paper cites Unified video action model, 2025.

Masked Visual Actions for Unified World Modeling Unified video action model, 2025

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:03.844866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:03.844866Z digest=sha256:c8a184a90a430566448052bbad400854d4f0a845a94ec5ce9c999f4aafe66311

Observation e3a97ca1-f025-4352-8f6b-92a27f8b75c0 · outbound

This paper cites Genie envisioner: A unified world foundation platform for robotic manipulation, 2025.

Masked Visual Actions for Unified World Modeling Genie envisioner: A unified world foundation platform for robotic manipulation, 2025

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:03.969261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:03.969261Z digest=sha256:39fd83e7203e9596219d346e802e145add066d4ad615b17ef345f4584c5c60d3

Observation 7074bee0-8bfb-406c-bb61-2d156196acaf · outbound

This paper cites Realwonder: Real-time physical action-conditioned video generation, 2026.

Masked Visual Actions for Unified World Modeling Realwonder: Real-time physical action-conditioned video generation, 2026

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:04.146958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:04.146958Z digest=sha256:98734c4605a5a218901a0248900a17f3c87d0ed7b4d12c91ff620350c9b9fcf7

Observation aeda653a-28e4-49e1-8fb2-551c6dd612fb · outbound

This paper cites Zero-shot world models are developmentally efficient learners.arXiv e-prints, pages arXiv–2604, 2026.

Masked Visual Actions for Unified World Modeling Zero-shot world models are developmentally efficient learners.arXiv e-prints, pages arXiv–2604, 2026

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:04.313578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:04.313578Z digest=sha256:ae92d18ec94ccb8c0a4c93ff9279f2d1611f0e00e74071dcf94adcc6148de343

Observation 0a3626b9-18bd-4032-8342-7b75e430cd94 · outbound

This paper cites Decoupled Weight Decay Regularization.

Masked Visual Actions for Unified World Modeling Decoupled Weight Decay Regularization

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:04.440100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:04.440100Z digest=sha256:a98a3a1335ff52f12d693d198cf49137d873f7b95f5a6c834015e900eba944a2

Observation 53ca9292-a352-415c-867c-1d1d2f99feca · outbound

This paper cites Mask world model: Predicting what matters for robust robot policy learning, 2026.

Masked Visual Actions for Unified World Modeling Mask world model: Predicting what matters for robust robot policy learning, 2026

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:04.517419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:04.517419Z digest=sha256:07231640187055297b13320ecfb9ec2684269edf722943433c5a362cecf0da0a

Observation 8d819bb7-73ec-48cb-a4fb-3a1113f2eabc · outbound

This paper cites Inference-time scaling for diffusion models beyond scaling denoising steps.

Masked Visual Actions for Unified World Modeling Inference-time scaling for diffusion models beyond scaling denoising steps

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:04.669230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:04.669230Z digest=sha256:e037ea19b192f99c9fa5199895a76adabf4e16455feac2e3916fd23f5c0b02aa

Observation 1797b1cc-201e-402b-903f-e5c8efb61ce4 · outbound

This paper cites MimicGen: A Data Generation System for Scalable Robot Learning using Human Demonstrations.

Masked Visual Actions for Unified World Modeling MimicGen: A Data Generation System for Scalable Robot Learning using Human Demonstrations

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:04.803449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:04.803449Z digest=sha256:6220631a04630cbd8541645116d769e5b16d8afaae256132752c1ba4102b280a

Observation ba782044-7c09-43b3-ac79-7443413b524d · outbound

This paper cites Motubrain: An advanced world action model for robot control, 2026.

Masked Visual Actions for Unified World Modeling Motubrain: An advanced world action model for robot control, 2026

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:04.967163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:04.967163Z digest=sha256:66dacd316a8e00b7b3f67efee5a5e0b68e86a9c3fd1cdfd56a76ca49a7e03546

Observation 9418a5d4-53e5-41d5-a965-49c5dfd9f8fc · outbound

This paper cites s1: Simple test-time scaling.

Masked Visual Actions for Unified World Modeling s1: Simple test-time scaling

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:05.141732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:05.141732Z digest=sha256:9eaed4e11299abe7a88c041057160b1f296cde1f6718f045f1954ea8f614e07c

Observation b115d4ba-4cf2-4981-a1d7-77bdd298e8ce · outbound

This paper cites Robocasa365: A large-scale simulation framework for training and benchmarking generalist robots.

Masked Visual Actions for Unified World Modeling Robocasa365: A large-scale simulation framework for training and benchmarking generalist robots

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:05.213696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:05.213696Z digest=sha256:8a0ebdcd2ead4eca1e76c12260f079926de7eb45024129a67e4a667cbdf917b7

Observation 66f9331b-073c-4c7f-b8e1-b45d435d366e · outbound

This paper cites Cosmos world foundation model platform for physical ai, 2025.

Masked Visual Actions for Unified World Modeling Cosmos world foundation model platform for physical ai, 2025

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:05.313907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:05.313907Z digest=sha256:a446ffdd9c9e13a7e67a0aacadc0a4be33b04b0f00679abcd0ba9b10c3316fb9

Observation 21d004a3-c5b7-49c1-8d44-8eee52cd6107 · outbound

This paper cites mimic-video: Video-action models for generalizable robot control beyond vlas, 2025.

Masked Visual Actions for Unified World Modeling mimic-video: Video-action models for generalizable robot control beyond vlas, 2025

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:05.491244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:05.491244Z digest=sha256:b8b7d1b3a6c8663fa6f81b2ad88db2d18b2377ae774afd904e7e0d16470e89ed

Observation dd72077b-bc3b-4629-93d2-a02b9ada1f64 · outbound

This paper cites Inference-time enhancement of generative robot policies via predictive world modeling.IEEE Robotics and Automation Letters, 2026.

Masked Visual Actions for Unified World Modeling Inference-time enhancement of generative robot policies via predictive world modeling.IEEE Robotics and Automation Letters, 2026

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:05.647880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:05.647880Z digest=sha256:8434095c359f60a9e4c07c288a617ce2442263a8c4bfe241bd1b50f3a99e10c4

Observation 71b4d916-03dc-4871-b8a2-c00044d164e1 · outbound

This paper cites MotionStream: Real-Time Video Generation with Interactive Motion Controls.

Masked Visual Actions for Unified World Modeling MotionStream: Real-Time Video Generation with Interactive Motion Controls

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:05.770198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:05.770198Z digest=sha256:4bf13e0c35a358d17e9a5d38398faedf9a2ea965fddad599fdc7e6ff891b66de

Observation 1ec0c14c-6f66-4171-a489-5ad8b747535b · outbound

This paper cites SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics.

Masked Visual Actions for Unified World Modeling SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:05.914908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:05.914908Z digest=sha256:afb877965e9e8a89f92b2b5f0a3002399031efc6a47486d473058add6f860dbc

Observation 11d1fc99-ed4d-4540-8640-45f8b08d9725 · outbound

This paper cites Time-to-move: Training-free motion-controlled video generation via dual-clock denoising.

Masked Visual Actions for Unified World Modeling Time-to-move: Training-free motion-controlled video generation via dual-clock denoising

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:06.115706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:06.115706Z digest=sha256:54a8b2a298ce395173bee87b898485e8592530f8635cd04cb904c8be4c419d46

Observation 7af91829-ba8c-458f-a05f-24b492076649 · outbound

This paper cites Motion before action: Diffusing object motion as manipulation condition, 2024.

Masked Visual Actions for Unified World Modeling Motion before action: Diffusing object motion as manipulation condition, 2024

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:06.248811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:06.248811Z digest=sha256:dce4fd14c14a36053ef3eeb1c56412dff155dc93d5b876e2803d3ed9014e1b4b

Observation 5fcd5563-ee5d-4d2f-8558-01d98ef167a4 · outbound

This paper cites Evaluating gemini robotics policies in a veo world simulator, 2025.

Masked Visual Actions for Unified World Modeling Evaluating gemini robotics policies in a veo world simulator, 2025

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:06.418490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:06.418490Z digest=sha256:d3dbfc7fa1a6cf89d448d9668b6f57f15a00058e0749da911d80e3b9fdb375dd

Observation 3bd46ac0-233f-4c9d-b740-f7251e10d84d · outbound

This paper cites Causal video models are data-efficient robot policy learners.Rhoda AI Blog, 2026.

Masked Visual Actions for Unified World Modeling Causal video models are data-efficient robot policy learners.Rhoda AI Blog, 2026

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:06.557565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:06.557565Z digest=sha256:c93e63a2d7a4b9d50dffde574d8863aca86cdaa142c22f7ea1ef39819364e7fa

Observation eec87124-0133-4b11-8def-0642d0540e5a · outbound

This paper cites Understanding physical dynamics with counterfactual world modeling.

Masked Visual Actions for Unified World Modeling Understanding physical dynamics with counterfactual world modeling

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:06.659959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:06.659959Z digest=sha256:c052fd18727b2cdd2bde3a98b90f06b55de6f34d2ba82dd8d6356426b2f4a121

Observation 62940561-bdaf-4541-b98a-a626f1390d4b · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

Masked Visual Actions for Unified World Modeling Wan: Open and Advanced Large-Scale Video Generative Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:06.746750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:06.746750Z digest=sha256:c37e13184c9900b7e02eef6fe0cd4dc5df535f9eac2ef7b95a824fda07958990

Observation 9b9c6d2d-e9a9-4417-8cb2-57c7a9b151d2 · outbound

This paper cites Eva: Aligning video world models with executable robot actions via inverse dynamics rewards, 2026.

Masked Visual Actions for Unified World Modeling Eva: Aligning video world models with executable robot actions via inverse dynamics rewards, 2026

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:06.915628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:06.915628Z digest=sha256:e6b406750a6ede63738d4b3e57f4a80a19d38674a50492592add41172d426b30

Observation c3dcb25d-8b7a-455e-bf13-25ecb7dc3cbe · outbound

This paper cites Interactive world simulator for robot policy training and evaluation.arXiv preprint arXiv:2603.08546, 2026.

Masked Visual Actions for Unified World Modeling Interactive world simulator for robot policy training and evaluation.arXiv preprint arXiv:2603.08546, 2026

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:06.965342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:06.965342Z digest=sha256:732a595da0b2ba2810bed5abb9957fddfbf42c2186fb047c7c5b91e695d5ca7e

Observation 741604b0-4a2d-4094-a57e-0a50123789e9 · outbound

This paper cites Precise action-to-video generation through visual action prompts.

Masked Visual Actions for Unified World Modeling Precise action-to-video generation through visual action prompts

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:07.047797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:07.047797Z digest=sha256:cbd93a0e91be60ad334f05d7caf5c995c25deb7f34076d6be42454786dad9572

Observation f7cdd53a-fa44-48ef-960c-1a816e2d4875 · outbound

This paper cites Wolpert and J.

Masked Visual Actions for Unified World Modeling Wolpert and J

Reference 67

Resolution
verified exact
doi, observed 2026-08-01T12:53:42.293862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-01T12:49:07.189078Z digest=sha256:a647524699fc102991c4a157b7547717c0c693c539e66560245df8bf7d6ccb9d

Observation 85aa13da-c5b9-4d46-801a-221179b3d2c5 · outbound

This paper cites Wolpert, Zoubin Ghahramani, and Michael I.

Masked Visual Actions for Unified World Modeling Wolpert, Zoubin Ghahramani, and Michael I

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:07.307228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:07.307228Z digest=sha256:ebaaa6fccfe707d2c85bf19a344abc6891f44d76d70c78f0bab15bcfc961eb93

Observation cf8c8a53-9b1c-4fbf-9340-eae84e13fda8 · outbound

This paper cites Wolpert, R.

Masked Visual Actions for Unified World Modeling Wolpert, R

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:07.401627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:07.401627Z digest=sha256:5721dcee38313ec0a993c4b025a104a9eccf5cb0f2bceeb3f6d76f186b6e7bd8

Observation 99715a3f-d73f-40a5-a1ab-86533f20e61c · outbound

This paper cites Sun, Ashley Neall, Tong Wu, Shengqu Cai, and Gordon Wetzstein.

Masked Visual Actions for Unified World Modeling Sun, Ashley Neall, Tong Wu, Shengqu Cai, and Gordon Wetzstein

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:07.439272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:07.439272Z digest=sha256:b622aa6a2883a64a62dd1859ba0822d2bc79de634bee98ddc9f54b24f00a29d7

Observation f5d9321b-9944-43d7-abc5-f5748404f8fb · outbound

This paper cites Kinema4d: Kinematic4d world modeling for spatiotemporal embodied simulation.arXiv preprint arXiv:2603.16669, 2026.

Masked Visual Actions for Unified World Modeling Kinema4d: Kinematic4d world modeling for spatiotemporal embodied simulation.arXiv preprint arXiv:2603.16669, 2026

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:07.518129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:07.518129Z digest=sha256:5369185707f319c9970a307458b7bb9fab4b7801c138e23d483c6927802b489a

Observation 8cfdc1d7-e718-41b2-9505-ac8e1d8a3d17 · outbound

This paper cites RoboPanoptes: The All-Seeing Robot with Whole-body Dexterity.

Masked Visual Actions for Unified World Modeling RoboPanoptes: The All-Seeing Robot with Whole-body Dexterity

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:07.719051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:07.719051Z digest=sha256:b9f97ead077b371723db105181872010639ef997684b5af445d38c078aa7ea5a

Observation ab6ab918-592d-4345-85df-311e9e6c4ab7 · outbound

This paper cites Learning Interactive Real-World Simulators.

Masked Visual Actions for Unified World Modeling Learning Interactive Real-World Simulators

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:07.854932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:07.854932Z digest=sha256:61e20f8fd41f13b80f2148077d4e159bef5d6cb54c10ef93491a67cc3ed263f3

Observation 84c4a1d6-799e-470f-8c8a-c27e4f9caa3e · outbound

This paper cites Orv: 4d occupancy-centric robot video generation, 2025.

Masked Visual Actions for Unified World Modeling Orv: 4d occupancy-centric robot video generation, 2025

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:07.975262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:07.975262Z digest=sha256:f458e4bbcfe57fcf05ed24986e19d9b4efb575ff0ac5397fcd95c97f03985974

Observation a8e5f30f-c093-4a77-8085-dcaeba235004 · outbound

This paper cites Real-to-sim robot policy evaluation with gaussian splatting simulation of soft-body interactions.

Masked Visual Actions for Unified World Modeling Real-to-sim robot policy evaluation with gaussian splatting simulation of soft-body interactions

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:08.072132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:08.072132Z digest=sha256:999ef8015ffb471352f0513370e2a23081578352fe3ff5ecd254adfed6d8bf37

Observation 84f6297c-83a5-4b2f-ac76-2c38fd5fab04 · outbound

This paper cites Veo-act: How far can frontier video models advance generalizable robot manipulation?, 2026.

Masked Visual Actions for Unified World Modeling Veo-act: How far can frontier video models advance generalizable robot manipulation?, 2026

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:08.184918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:08.184918Z digest=sha256:40cd1706438bd09f4a7ddfec64632aeef4437473f01d5d69ab7b814727aaabd2

Observation 7fccbe33-c8ae-4be6-a725-5274c7345b24 · outbound

This paper cites Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware.

Masked Visual Actions for Unified World Modeling Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:08.273965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:08.273965Z digest=sha256:fa00f7cb48a44a51e8140dce96cfd2b00f44c0a1c576401d5b40b8c8c030fa20

Observation 6c349569-4e4e-44f4-a7b6-6acc18cc9ca9 · outbound

This paper cites TesserAct: Learning 4D Embodied World Models.

Masked Visual Actions for Unified World Modeling TesserAct: Learning 4D Embodied World Models

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:08.342623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:08.342623Z digest=sha256:6042edda5e94b14e7b9b7cf147d6150a176518487e9d0d9bac5ee8b05b0cd3f1

Observation 8d9b496d-c710-48b7-8b9d-8339d9ea3ca8 · outbound

This paper cites Action images: End-to-end policy learning via multiview video generation, 2026.

Masked Visual Actions for Unified World Modeling Action images: End-to-end policy learning via multiview video generation, 2026

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:08.422235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:08.422235Z digest=sha256:338764bfba34e03d4bf978fb50e3a98142864bd39eaf53a467af84b15ef00deb

Observation d5264796-98d5-4805-88bd-ef0430366b8d · outbound

This paper cites 3dflowaction: Learning cross-embodiment manipulation from 3d flow world model, 2025.

Masked Visual Actions for Unified World Modeling 3dflowaction: Learning cross-embodiment manipulation from 3d flow world model, 2025

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:08.526653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:08.526653Z digest=sha256:847c875acf23c138d56627dd88ff3154e800fc51531f53cfa16fc90291454fa7

Observation 02ec8abe-6c1a-4de1-8892-851c0f196470 · outbound

This paper cites Unified world models: Coupling video and action diffusion for pretraining on large robotic datasets, 2025.

Masked Visual Actions for Unified World Modeling Unified world models: Coupling video and action diffusion for pretraining on large robotic datasets, 2025

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:08.630493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:08.630493Z digest=sha256:351cb087570b53854bf2f056a957ff666f260aa4f37c468baf8543b86cb5d6c8

Observation 89f044fc-98f3-4c8f-ace7-cb3e24366785 · outbound

This paper cites Irasim: Learning interactive real-robot action simulators.

Masked Visual Actions for Unified World Modeling Irasim: Learning interactive real-robot action simulators

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:08.710108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:08.710108Z digest=sha256:0d1d8da2277a31e5a2f9303ea9ed4de46b9ab07de6f6a5c1cb9f83f074c6c50f

Observation 5ab4aa41-44c7-4018-a44f-70630b5e1320 · outbound

This paper cites an unresolved cited work.

Masked Visual Actions for Unified World Modeling Unresolved cited work

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:08.889648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:08.889648Z digest=sha256:3244588efeaed811a4c048663179b29a5004a831ab026e16fd1a0af78f232d27

Observation bf97c2d8-397f-4f77-9d92-8dc248a27450 · outbound

This paper cites an unresolved cited work.

Masked Visual Actions for Unified World Modeling Unresolved cited work

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:09.033921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:09.033921Z digest=sha256:1e951fce95772b58e51c6c4e380bc3bf383a0315de87bdedf4f710e676a77972

Observation 9b682be1-bf24-422b-a688-e78410cff85f · outbound

This paper cites an unresolved cited work.

Masked Visual Actions for Unified World Modeling Unresolved cited work

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:09.150745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:09.150745Z digest=sha256:37e8982a1d551b56b60f881f5acc0ca5e4046578512bf6079f24fb2ab1734703

Observation fe024360-1b51-40aa-a64c-a06c41f6aca3 · outbound

This paper cites an unresolved cited work.

Masked Visual Actions for Unified World Modeling Unresolved cited work

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:09.279685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:09.279685Z digest=sha256:a932cd6eb6a1a8924bba3bc290c062f3310a5776771d74ce8fd615a36bffe045

Observation 95f7312e-19d5-42d6-94c0-7ad8072974d4 · outbound

This paper cites Covers Diffusion Policy, ACT, and SmolVLA baselines.

Masked Visual Actions for Unified World Modeling Covers Diffusion Policy, ACT, and SmolVLA baselines

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:09.427171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:09.427171Z digest=sha256:3e885b4735d0b81dc300b2b1d752ca9d1452b59957f43379d67340260da57723

Observation 21b2c77b-ec9e-4e58-ab39-bdd522ca28a7 · outbound

This paper cites an unresolved cited work.

Masked Visual Actions for Unified World Modeling Unresolved cited work

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:09.512352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:09.512352Z digest=sha256:7ec23b55a7534425c81f4ed3cdef9e6a965b0d584557dffd5add922cda32d0b7

Observation 40ca33fc-e5b8-44ee-b02f-2fe35df2e30d · outbound

This paper cites an unresolved cited work.

Masked Visual Actions for Unified World Modeling Unresolved cited work

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:09.557487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:09.557487Z digest=sha256:7176ea2c886f39bfc8739745341698e3028dce8daa613a902515175b09454e99

Observation 2d75b15b-efbb-4d84-bf59-8a6507ff0f7c · outbound

This paper cites the toaster door must be flush, latched, and closed by a push from below.

Masked Visual Actions for Unified World Modeling the toaster door must be flush, latched, and closed by a push from below

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:09.642077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:09.642077Z digest=sha256:0b27a9fab3e58784e288526459fa0f67a46a601d0cd6dab31ca483a7c6bead3b

Observation ddfca20e-1655-4c33-ac83-8b263a5b037e · outbound

This paper cites robot_pushed.

Masked Visual Actions for Unified World Modeling robot_pushed

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:09.706619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:09.706619Z digest=sha256:07ffd80ee65cff1728cba5ba9e66147c36dbf48604171ff2c5a040990a6f26d0

Observation fdc01ae2-3c49-4693-b456-1c166504bdcb · outbound

This paper cites Outcomes reached after disengagement, via ghost contact, teleport, or autonomous motion are FALSE.

Masked Visual Actions for Unified World Modeling Outcomes reached after disengagement, via ghost contact, teleport, or autonomous motion are FALSE

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:09.751858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:09.751858Z digest=sha256:155ed09e556c93292be96713679bd4386a0a52db6b2a6f25add057a9e65f2d94

Observation a6af23b8-78ba-4f05-af0b-73a5d0d03040 · outbound

This paper cites Hovering near.

Masked Visual Actions for Unified World Modeling Hovering near

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:09.789865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:09.789865Z digest=sha256:f8ef36bc944efbb10bafb5f62c62f0f3df18ad16f600a5974a131876d41fe5f7

Observation 7ab1c18a-7e90-414e-93a6-5b710927f094 · outbound

This paper cites Penalize teleporting, morphing, gripper passing through solids, vanishing or duplicated objects, frame-to-frame 18 jumps, ghost contact, post-disengagement coasting.

Masked Visual Actions for Unified World Modeling Penalize teleporting, morphing, gripper passing through solids, vanishing or duplicated objects, frame-to-frame 18 jumps, ghost contact, post-disengagement coasting

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:09.834338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:09.834338Z digest=sha256:3fefc7ec2535d695293ef7119ac1499fc4e94b609ca470f3f346dd1f495451d6

Pith citing papers

Observation 106f1d63-784b-4ced-a55e-ba457d3205a3 · inbound

XEWorld: Can Action-Conditioned World Models Generalize to Unseen Robot Embodiments? cites this paper.

XEWorld: Can Action-Conditioned World Models Generalize to Unseen Robot Embodiments? Masked Visual Actions for Unified World Modeling

Reference 29

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T23:21:53.853473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T23:21:53.430055Z digest=sha256:de1e829b07b6f4356137d431393a4604ae27c217757025de9a7ceb328b75ea5f