Pith. sign in

Paper Citation Record · LEDGER

MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models

As of 5 August 2026, this Paper Citation Record lists 67 of 67 outbound references and 7 inbound Pith citation observations for arXiv:2606.13515.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.13515 v1

Coverage vector

measured 67 of 67 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-27T06:55:53.760595Z

measured 74 of 74 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T03:23:47.997711Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-09T22:16:36.389750Z

Reference resolution

67 of 67 outbound references displayed

  • verified exact41
  • verified fuzzy0
  • unresolved22
  • parse uncertain1
  • malformed identifier1
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0ef878e4-4015-4b38-ab3b-8392410e3685 · outbound

This paper cites Affordances from human videos as a versatile representation for robotics.

MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models Affordances from human videos as a versatile representation for robotics

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-27T06:55:53.760595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T06:55:53.760595Z digest=sha256:ed9b2b8bf3b1fa8f9cd1269e84f81b7fc7c566f33053860e00dd9b5a2a165bdd

Observation 62632859-50c4-46c8-a9a4-a012ea57daeb · outbound

This paper cites Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation.

MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-03T14:48:32.776438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T06:55:53.760595Z digest=sha256:33589a6ecdec2fd9254e98208851c0916a8d7e0fef8e0deac286a5ea7d07a468

Observation 2baca82f-68ff-4221-aee3-b16c43be7bab · outbound

This paper cites Track2act: Predicting point tracks from internet videos enables generalizable robot manipulation.

MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models Track2act: Predicting point tracks from internet videos enables generalizable robot manipulation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-27T06:55:53.760595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T06:55:53.760595Z digest=sha256:5c2edc87cff95d32712330187f848c9f9b9c5b9b29f92b70a737aa86a938456e

Observation c6954900-03d1-4037-933a-6d348024c767 · outbound

This paper cites Motus: A unified latent action world model.

MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models Motus: A unified latent action world model

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-27T06:55:53.760595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T06:55:53.760595Z digest=sha256:43e71d3ecb124143c9f3e47f394f7067d6d885abc350f8d7850ff9ef93023e64

Observation b7c4cbfd-d64c-4699-bbbc-432fb21dfafa · outbound

This paper cites Gr00t n1: An open foundation model for generalist humanoid robots.arXiv preprint, 2025.

MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models Gr00t n1: An open foundation model for generalist humanoid robots.arXiv preprint, 2025

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-27T06:55:53.760595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T06:55:53.760595Z digest=sha256:b4a970a32ff049525da4809949ceecece44727a6c1ea4b60a61022275aa4e859

Observation a9671f4d-5362-4e68-86be-5c07bcb55fe6 · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T14:48:32.808722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T06:55:53.760595Z digest=sha256:eed7ae63ca70962b6d528d4ffc9fa06c83e7ac6cca8a8bc9af8968921cac2d14

Observation b1c01647-054c-45b3-b7b9-c7cd281b73bb · outbound

This paper cites $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization.

MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T14:48:32.835026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T06:55:53.760595Z digest=sha256:4dab3d9fd7231fad7fee8f7ad8a9fd2ea8a1e93c32f3f3976ad6d5f4d9e8ea09

Observation eda8f390-75e2-46c6-9e49-4bf52d52c1cb · outbound

This paper cites SAM 3: Segment Anything with Concepts.

MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models SAM 3: Segment Anything with Concepts

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-03T14:48:32.790570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T06:55:53.760595Z digest=sha256:720cb67748b63efdbe9a84f509d4ceb40ad9a3c3fc980ec5d8f7c2e477d1a613

Observation 35ea5501-a3d7-4205-b1a2-b5043ee776ff · outbound

This paper cites RynnVLA-002: A Unified Vision-Language-Action and World Model.

MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models RynnVLA-002: A Unified Vision-Language-Action and World Model

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-03T14:48:32.839921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T06:55:53.760595Z digest=sha256:510b98e4e5cda56255e89a4e7502e6a986bfe933abe6b4a332a88aeef13ab10f

Observation 06f20129-1aa3-4900-be49-7b25b6c79678 · outbound

This paper cites Worldvla: Towards autoregressive action model with world knowledge.arXiv preprint, 2025.

MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models Worldvla: Towards autoregressive action model with world knowledge.arXiv preprint, 2025

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-27T06:55:53.760595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T06:55:53.760595Z digest=sha256:72e168924cde55cd38788662e6e9bf5f3bbd35dbb053e160d676b76754f75c9c

Observation dcf9b0ac-d4eb-4b17-b971-811e024b1a21 · outbound

This paper cites GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation.

MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-03T14:48:32.853298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T06:55:53.760595Z digest=sha256:e9e118a5edad93b29a6b7d1aba3d91d67a35c516203e6b27f081e4407d8bbf86

Observation bde575de-2795-4e93-964b-3ff137efb37f · outbound

This paper cites RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation.

MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-03T14:48:32.837659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T06:55:53.760595Z digest=sha256:c00c6e19babfe58c08ddf88a91799c68cac81998cdd40fdcba6a969d84c37068

Observation 3e5861dc-ecca-40a2-a6b2-008152697a0f · outbound

This paper cites Learning Universal Policies via Text-Guided Video Generation.

MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models Learning Universal Policies via Text-Guided Video Generation

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:48:32.861585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T06:55:53.760595Z digest=sha256:3d13dbf75ee458a673b5df6c45da8f1e1e40666533580c8552e7059d4081e4ec

Observation 557305d0-2f96-4f16-b0eb-97e7eb89e3eb · outbound

This paper cites Vidar: Embodied Video Diffusion Model for Generalist Manipulation.

MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models Vidar: Embodied Video Diffusion Model for Generalist Manipulation

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-03T14:48:32.789995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T06:55:53.760595Z digest=sha256:d80f7e0c9cd1595a7d43cb413ca4bd8d6e9acbc04c02d5ce402169c3f1f57640

Observation 82781542-4b18-4e5c-bda1-af2d38814604 · outbound

This paper cites Rt-trajectory: Robotic task generalization via hindsight trajectory sketches, 2023.

MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models Rt-trajectory: Robotic task generalization via hindsight trajectory sketches, 2023

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-27T06:55:53.760595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T06:55:53.760595Z digest=sha256:8f6fd5aa3055b8daa5f8e378a6ea48d09f1abfda0f7770414b2c12fae928852e

Observation 88fe2297-96d9-4f80-abd5-7ad075120542 · outbound

This paper cites Point Policy: Unifying Observations and Actions with Key Points for Robot Manipulation.

MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models Point Policy: Unifying Observations and Actions with Key Points for Robot Manipulation

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:48:32.858643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T06:55:53.760595Z digest=sha256:a19baed8caac5caef16ea557c839f4d68745c58c015e9216ed25cef3e7a2e80d

Observation fd9bc1df-e626-4ad1-a0d9-abce5ed54223 · outbound

This paper cites Spot: Se (3) pose trajectory diffusion for object-centric manipulation.

MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models Spot: Se (3) pose trajectory diffusion for object-centric manipulation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-27T06:55:53.760595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T06:55:53.760595Z digest=sha256:78d5c2e67b654041efb7df3875ed455116b13d6341a117357337a85a9cf490c6

Observation 5ec2dbbc-d991-41ec-926e-3e8b5b533aed · outbound

This paper cites Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations.

MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-07-03T14:48:32.840658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T06:55:53.760595Z digest=sha256:be5dd1d19bd8e006da3af4f3124653ac251ed5cf89d6f0eecb9ba98cb3ac5e78

Observation 257cc779-9244-4790-acf0-fa65057e9ae4 · outbound

This paper cites Roboground: Robotic manipulation with grounded vision-language priors.

MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models Roboground: Robotic manipulation with grounded vision-language priors

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-27T06:55:53.760595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T06:55:53.760595Z digest=sha256:0d7ddc5c985011c4590e1affd683ecc9978b7dcc4794d74b503417f6bba4ff47

Observation 96c786c4-e0e6-471d-b014-af5e5eadc024 · outbound

This paper cites ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation.

MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-03T14:48:32.799999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T06:55:53.760595Z digest=sha256:bd5fdaad353588cb89d648f784c54ffce3edb5d833e963a82962b86b36b6d866

Observation 506b6d6b-7b09-4ab5-857c-93d06d2ea19b · outbound

This paper cites DreamGen: Unlocking Generalization in Robot Learning through Video World Models.

MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models DreamGen: Unlocking Generalization in Robot Learning through Video World Models

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-07-03T14:48:32.871031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T06:55:53.760595Z digest=sha256:ee6c8d59a7cf1f286ff9b84d09b4076cef5dd3ff66e6f8355c20ebad5641755e

Observation 677e8e10-4c45-4e43-b51e-e5d9627ea60c · outbound

This paper cites Robo-abc: Affordance generalization beyond categories via semantic correspondence for robot manipulation.

MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models Robo-abc: Affordance generalization beyond categories via semantic correspondence for robot manipulation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-27T06:55:53.760595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T06:55:53.760595Z digest=sha256:38e75a3a49cfe2c1bedb75c219d542320a4eb0c6bfbcab59c8698b0153676155

Observation b481fdee-2b71-47ee-8af6-cbd0b9d1ab9e · outbound

This paper cites Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning.

MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-07-03T14:48:32.866010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T06:55:53.760595Z digest=sha256:16edd2ffc2b0fb1befdfce06d71ef2a8228fa30c216af0e0c314e2c8bb23adc4

Observation e30b5a95-79ea-4450-8be4-ba986e49bcf3 · outbound

This paper cites MolmoAct: Action Reasoning Models that can Reason in Space.

MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models MolmoAct: Action Reasoning Models that can Reason in Space

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-07-03T14:48:32.861030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T06:55:53.760595Z digest=sha256:4144747e35f971de411255653c25e59a9ab02de0c81e08a18776d800b19097e8

Observation efcee695-882a-44ea-9e78-c97f1102fa99 · outbound

This paper cites Causal World Modeling for Robot Control.

MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models Causal World Modeling for Robot Control

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-07-03T14:48:32.842258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T06:55:53.760595Z digest=sha256:7a8fc34903b7c4a8d797363388c03302c6e11f1bae63f9d79093c69e6bb9e0cb

Observation 27e3a5be-f2fb-4e2c-aae8-907fc7f80c20 · outbound

This paper cites ControlVLA: Few-shot Object-centric Adaptation for Pre-trained Vision-Language-Action Models.

MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models ControlVLA: Few-shot Object-centric Adaptation for Pre-trained Vision-Language-Action Models

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:48:32.825517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T06:55:53.760595Z digest=sha256:d0f81dcc3b66eaa19ba3820d743e77960d793d70ac5a0c63d4743765d05e8ce0

Observation 893f6c00-759c-462a-b33f-bcb2aa11e45c · outbound

This paper cites Video Generators are Robot Policies.

MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models Video Generators are Robot Policies

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-07-03T14:48:32.858750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T06:55:53.760595Z digest=sha256:96ea31d8568f937082fe690a3908dc5dcd75cd799be826cbb51f55b58ec333d6

Observation dac770c4-9af0-4010-9b08-355be8a0b9eb · outbound

This paper cites Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models.

MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-07-03T14:48:32.850863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T06:55:53.760595Z digest=sha256:4938f89b260b5f59f20b8fe7f7b53c68eee3519a06015ff43cbe0f365af792ed

Observation d15a86f7-447f-43aa-b434-c77e1eeeddb7 · outbound

This paper cites Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation.

MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-07-03T14:48:32.873069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T06:55:53.760595Z digest=sha256:37848335ff6c5be27faa7c10c30325e7bb7807bf23aa16ed5a1ce9d986cbcaea

Observation 14adc3dd-c331-4945-844f-7b8981d089e0 · outbound

This paper cites PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies.

MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-07-03T14:48:32.870836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T06:55:53.760595Z digest=sha256:22636879513b40abefd799563d1567a3b5f0ecc7129b0001f4591e2f9f6cee22

Observation 562d7fb7-7ad7-4889-9eb6-7b97b261f992 · outbound

This paper cites Libero: Benchmarking knowledge transfer for lifelong robot learning.Advances in Neural Information Processing Systems, 36:44776– 44791, 2023.

MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models Libero: Benchmarking knowledge transfer for lifelong robot learning.Advances in Neural Information Processing Systems, 36:44776– 44791, 2023

Reference 32

Resolution
unresolved
no resolver link, observed 2026-06-27T06:55:53.760595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T06:55:53.760595Z digest=sha256:8775c5bc5928c4b561be26e8f7f35b37d8402139254e41408c47098b484e43b7

Observation f567c644-6c7a-4cd8-aa53-7b55fd10ee8a · outbound

This paper cites Moka: Open-world robotic manipulation through mark-based visual prompting.

MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models Moka: Open-world robotic manipulation through mark-based visual prompting

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-27T06:55:53.760595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T06:55:53.760595Z digest=sha256:3a87899f6de15473817b154c429d4161f40d57f2be3885335c1b236e6e4f1a1f

Observation dfea2e90-9106-4adb-bd0f-5df94bc1766b · outbound

This paper cites MOKA: Open-World Robotic Manipulation through Mark-Based Visual Prompting.

MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models MOKA: Open-World Robotic Manipulation through Mark-Based Visual Prompting

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:48:32.868635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T06:55:53.760595Z digest=sha256:8711b14f4f25af45c263e7183fadcdada954cbf3554556f7b9943a652f77800b

Observation 43635b95-30b0-428b-bf07-11f06e418a87 · outbound

This paper cites Rdt-1b: a diffusion foundation model for bimanual manipulation.

MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models Rdt-1b: a diffusion foundation model for bimanual manipulation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-06-27T06:55:53.760595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T06:55:53.760595Z digest=sha256:90004f89ba0093367c36d3987c6abd2ad62865fb86bf7e73b6c0eedd40ea5690

Observation c6c961d5-c2e2-43e6-a244-74c026ed50ad · outbound

This paper cites Mask World Model: Predicting What Matters for Robust Robot Policy Learning.

MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models Mask World Model: Predicting What Matters for Robust Robot Policy Learning

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-07-03T14:48:32.819957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T06:55:53.760595Z digest=sha256:15def95bc7831ff2b9ff16d413e2b65a84c42bf92aa6f8d174ea8f2bf86d4fa4

Observation 250a5740-2610-4e78-b91c-3ed41c135a8d · outbound

This paper cites Dit4dit: Jointly modeling video dynamics and actions for generalizable robot control.

MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models Dit4dit: Jointly modeling video dynamics and actions for generalizable robot control

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:48:32.827083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T06:55:53.760595Z digest=sha256:51f17250772d748fa8d9512efbcd2d6cdd975029475eacf239386c03b144c7fe

Observation 5deda1b6-ab54-4a5d-8ba7-d94412783cea · outbound

This paper cites Grounded human-object interaction hotspots from video.

MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models Grounded human-object interaction hotspots from video

Reference 38

Resolution
unresolved
no resolver link, observed 2026-06-27T06:55:53.760595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T06:55:53.760595Z digest=sha256:c32ace8f55f98cdc5123a30f529632aa51607687ac585f60e3fcdfa45932cc75

Observation 75a4e62e-30e9-4904-8b53-631050a96a97 · outbound

This paper cites Rt-affordance: Affordances are versatile intermediate representations for robot manipulation.

MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models Rt-affordance: Affordances are versatile intermediate representations for robot manipulation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-06-27T06:55:53.760595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T06:55:53.760595Z digest=sha256:bdd8a70a7a18078c9ef9c6ae3ff3d6e9a92f7baf496f9bc4e5a6ee840788229a

Observation daca30db-0550-4d47-9c88-00fc623151ae · outbound

This paper cites mimic-video: Video-Action Models for Generalizable Robot Control Beyond VLAs.

MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models mimic-video: Video-Action Models for Generalizable Robot Control Beyond VLAs

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-07-03T14:48:32.827879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T06:55:53.760595Z digest=sha256:1fbdb14dd951840088d8faf72b12bed879c860ae5db2113a2a529ce472f68059

Observation 022ce25f-4564-4a88-8249-82821eeb49f3 · outbound

This paper cites Qwen3-vl: A frontier multimodal large language model.https://github.com/QwenLM/Qwen3-VL,.

MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models Qwen3-vl: A frontier multimodal large language model.https://github.com/QwenLM/Qwen3-VL,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-06-27T06:55:53.760595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T06:55:53.760595Z digest=sha256:78098f4b42fd49d3086e6fe9485e4ba887af141aa83c23d524da2d9a182c754f

Observation 445993a4-7872-44a2-8c84-05a49789cbc0 · outbound

This paper cites an unresolved cited work.

MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models Unresolved cited work

Reference 42

Resolution
parse uncertain
no resolver link, observed 2026-06-27T06:55:53.760595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T06:55:53.760595Z digest=sha256:6e5c6ac4273c2ec4b6a238d570d3c22f228825ffebb313066776f6984acefd81

Observation 8e44d8b1-5e5d-4223-a3c0-9bc99e23ca1f · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.Journal of machine learning research, 21(140):1–67, 2020.

MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models Exploring the limits of transfer learning with a unified text-to-text transformer.Journal of machine learning research, 21(140):1–67, 2020

Reference 43

Resolution
unresolved
no resolver link, observed 2026-06-27T06:55:53.760595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T06:55:53.760595Z digest=sha256:6b3c360dc0c913d29c966d37b8ec190082c6df0e302c31965380b1337c62ce81

Observation e5d56b44-41ec-4490-9675-159f5d6fcc06 · outbound

This paper cites Open-world object manipulation using pre-trained vision-language models.

MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models Open-world object manipulation using pre-trained vision-language models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-06-27T06:55:53.760595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T06:55:53.760595Z digest=sha256:c85a502df52c0fb817292a03dfb5fe8352317982928a194f72a183684d5b4063

Observation d81bf1e9-63ab-4ad9-bf55-e0e967c36678 · outbound

This paper cites Vla-jepa: Enhancing vision-language-action model with latent world model.

MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models Vla-jepa: Enhancing vision-language-action model with latent world model

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:48:32.788136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T06:55:53.760595Z digest=sha256:515a54a9b46e40d75baf26863e856796c2d573419309dfd1da6920e78669172b

Observation 4b5b82ee-50fc-4ca1-aaf6-b4b515283be0 · outbound

This paper cites Kite: Keypoint-conditioned policies for semantic manipulation, 2023.

MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models Kite: Keypoint-conditioned policies for semantic manipulation, 2023

Reference 46

Resolution
unresolved
no resolver link, observed 2026-06-27T06:55:53.760595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T06:55:53.760595Z digest=sha256:2f02559836853fa1832b7a8cdd72be4460c4f2e664773a2f8a81a391b8d01d22

Observation 285ef188-316e-488b-b048-fbf49b5a0036 · outbound

This paper cites HY-Embodied-0.5: Embodied Foundation Models for Real-World Agents.

MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models HY-Embodied-0.5: Embodied Foundation Models for Real-World Agents

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-07-03T14:48:32.809991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T06:55:53.760595Z digest=sha256:c8e4df10e8adb50061788e67f52f1b67d0acc0bf512096be63fb8853a828d6f0

Observation 0f8d9421-7089-45bf-80c7-7ba47a2ad175 · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-07-03T14:48:32.782125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T06:55:53.760595Z digest=sha256:d47d94c34d02350ec6223c187380350b41d6afd1b3f39aa234f07c9413e3219e

Observation 6eb4a742-426d-4c2f-85c9-60e33cf669e2 · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models Wan: Open and Advanced Large-Scale Video Generative Models

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-07-03T14:48:32.868536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T06:55:53.760595Z digest=sha256:3e29778c23e37c5b616835b681ddba8a8e54bf3ff18d2889fc9a44ef0174cf35

Observation 58fcc746-8cbd-4179-9933-ba321a8a2eba · outbound

This paper cites SKIL: Semantic Keypoint Imitation Learning for Generalizable Data-efficient Manipulation.

MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models SKIL: Semantic Keypoint Imitation Learning for Generalizable Data-efficient Manipulation

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:48:32.863620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T06:55:53.760595Z digest=sha256:0bbb62583d6c3bb2c36add09fc75d6677e21e4b17f0708ba1b13d992b7358838

Observation eec8e05b-a71a-45ad-8576-61ad39230792 · outbound

This paper cites Any-point Trajectory Modeling for Policy Learning.

MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models Any-point Trajectory Modeling for Policy Learning

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-07-03T14:48:32.832756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T06:55:53.760595Z digest=sha256:2d58bbdb11b008df8dc8ab978ef5c5c12a2367abc48888e2094e6914d95bfa18

Observation cc3bc517-242b-425e-b212-a123e059d14d · outbound

This paper cites Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model.

MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-07-03T14:48:32.842966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T06:55:53.760595Z digest=sha256:c5a842e7288fc95d2bb6f479e8950e082cfab58940c41ca1ac5dadca57c1e346

Observation af4f495d-70c4-4e82-b184-ee5e2723eb8e · outbound

This paper cites Unleashing large-scale video generative pre-training for visual robot manipulation, 2023.

MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models Unleashing large-scale video generative pre-training for visual robot manipulation, 2023

Reference 53

Resolution
unresolved
no resolver link, observed 2026-06-27T06:55:53.760595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T06:55:53.760595Z digest=sha256:ba4b9ae7a0e7574f00e668199bda3eba46b9b60e2cb0f2026e2a95c4da44a281

Observation d975eeef-af77-4901-a240-ed919ae4274d · outbound

This paper cites A Pragmatic VLA Foundation Model.

MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models A Pragmatic VLA Foundation Model

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-07-03T14:48:32.837942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T06:55:53.760595Z digest=sha256:852e1b91efebf0a0e23e10c56b68bcf0a9e3fea48ae5f2182e89672c6e9780cc

Observation 908655ab-ed7b-43bc-95e1-c5ef6bca66ab · outbound

This paper cites Flow as the Cross-Domain Manipulation Interface.

MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models Flow as the Cross-Domain Manipulation Interface

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:48:32.848415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T06:55:53.760595Z digest=sha256:591442dfe09f6387704cb88920fe406dc386c41a2401eb2396a17baa382cab55

Observation 7bab5752-2c10-486f-bdb9-3c04350f6bf3 · outbound

This paper cites Flow as the cross-domain manipulation interface.

MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models Flow as the cross-domain manipulation interface

Reference 56

Resolution
unresolved
no resolver link, observed 2026-06-27T06:55:53.760595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T06:55:53.760595Z digest=sha256:6b03bab37c9a1428c6142d2d29636a916e44dd2ad1504c920c92d3379031e075

Observation b3148b1b-39b5-4538-abec-6dd0d846e124 · outbound

This paper cites World Action Models are Zero-shot Policies.

MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models World Action Models are Zero-shot Policies

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-07-03T14:48:32.812375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T06:55:53.760595Z digest=sha256:d1f08b8bee3d9d7f6d8f9dfae9545e342506181129473d3bd2dd6587d5e2e73c

Observation 9a251410-2a8a-4867-8c0c-e1dab205303a · outbound

This paper cites Point what you mean: Visually grounded instruction policy.

MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models Point what you mean: Visually grounded instruction policy

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:48:32.856042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T06:55:53.760595Z digest=sha256:f572c4405511d6afacd225a9582eb5dfb6452044a25bd3cbeabaf125e21c1b1f

Observation 13dd516d-c730-4cb6-8eca-7e7cd782d13f · outbound

This paper cites LM-Gaussian: Boost Sparse-view 3D Gaussian Splatting with Large Model Priors.

MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models LM-Gaussian: Boost Sparse-view 3D Gaussian Splatting with Large Model Priors

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:48:32.807608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T06:55:53.760595Z digest=sha256:6ed0d565b8b582d38b93b6ed76343dbf5c6cd47cc08265fba2d395ff990a1b27

Observation 5420341a-6eff-46af-b91b-29567b7a7d9f · outbound

This paper cites Fast-WAM: Do World Action Models Need Test-time Future Imagination?.

MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models Fast-WAM: Do World Action Models Need Test-time Future Imagination?

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-07-03T14:48:32.845669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T06:55:53.760595Z digest=sha256:fe5e2ac49c17014e1b3f0a0ba48666812b58afc621b0b9d3a83bfe857dac0493

Observation 58227ffd-37e0-4898-a3bd-91655f024069 · outbound

This paper cites RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics.

MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:48:32.779436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T06:55:53.760595Z digest=sha256:3cac97caad8c5ddb62010fc23d9876f6895e9b531ae9895b118b5c1f9df8b60f

Observation 3c4dad22-9513-4d15-b15e-1dbd82a735cf · outbound

This paper cites DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge.

MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-06-27T07:00:39.873785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T06:55:53.760595Z digest=sha256:918966864f0591cb2df79c2481496bd33720e76ab0b62036b803532f065c3580

Observation fd810ebb-154a-4ed6-9bb7-12397e23dd76 · outbound

This paper cites Disentangled Robot Learning via Separate Forward and Inverse Dynamics Pretraining.

MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models Disentangled Robot Learning via Separate Forward and Inverse Dynamics Pretraining

Reference 63

Resolution
verified exact
local_arxiv, observed 2026-07-03T14:48:32.863936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T06:55:53.760595Z digest=sha256:33efd2219986e5a4f49d74b567db52d3d962d4b77ed502b7b3f93d6dddd57740

Observation 482f100e-32d9-4520-a0be-2534e0f551fa · outbound

This paper cites Cot-vla: Visual chain-of-thought reasoning for vision-language-action models, 2025.

MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models Cot-vla: Visual chain-of-thought reasoning for vision-language-action models, 2025

Reference 64

Resolution
unresolved
no resolver link, observed 2026-06-27T06:55:53.760595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T06:55:53.760595Z digest=sha256:6de01f39dff884a68ee5a82f239c837b15ed97c1cf2c78f2f5d2f6e668d05011

Observation 65e7a61e-b97c-4515-88b2-d6f86217d9dc · outbound

This paper cites FLARE: Robot Learning with Implicit World Modeling.

MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models FLARE: Robot Learning with Implicit World Modeling

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-07-03T14:48:32.866264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T06:55:53.760595Z digest=sha256:fdb7ebf34683ce2f76eb8de20326b615386330f501258594765992ebdf32e872

Observation 76047247-7d81-4709-b7ff-90e38df45862 · outbound

This paper cites Act2goal: From world model to general goal-conditioned policy, 2025.

MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models Act2goal: From world model to general goal-conditioned policy, 2025

Reference 66

Resolution
unresolved
no resolver link, observed 2026-06-27T06:55:53.760595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T06:55:53.760595Z digest=sha256:e60c43230354a1dbd360abea0eee6d8d63147bc020cdfbfb80b678a8902c0f8a

Observation afa71aa8-f3d5-44d5-a6f6-4a260acaed36 · outbound

This paper cites RoboDreamer: Learning Compositional World Models for Robot Imagination.

MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models RoboDreamer: Learning Compositional World Models for Robot Imagination

Reference 67

Resolution
verified exact
local_arxiv, observed 2026-07-03T14:48:32.830425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T06:55:53.760595Z digest=sha256:208a19847e047b11cadf674ff7baca80fc01fa67cd4eebc4408cf535dcaa734c

Observation 7af03346-7e7d-4ecd-aad3-3221750951f9 · outbound

This paper cites Unified World Models: Coupling Video and Action Diffusion for Pretraining on Large Robotic Datasets.

MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models Unified World Models: Coupling Video and Action Diffusion for Pretraining on Large Robotic Datasets

Reference 68

Resolution
malformed identifier
local_arxiv, observed 2026-07-03T14:48:32.829469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T06:55:53.760595Z digest=sha256:12344e9ab81b52e6c7b22e1598acfcaf69c6c970e6da1e79f4254a6aedddc2a7

Pith citing papers

Observation fc6a206d-370a-47f8-8eaf-08b9c22d2e2d · inbound

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? cites this paper.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:39:17.382706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:e97756e642f3e60d365867cf8529f888750e5189843784b465223d16277d2435

Observation 52c4186d-e49b-443f-aed4-58f60cdf057e · inbound

WAM-TTT: Steering World-Action Models by Watching Human Play at Test Time cites this paper.

WAM-TTT: Steering World-Action Models by Watching Human Play at Test Time MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-07-09T22:16:36.391030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T22:16:31.529359Z digest=sha256:94927a687eaec1f535bbb233f561f2c1a91898cb6d0973cf55bcb598c7197cfe

Observation 353c7a26-527d-41ce-9909-9322f610238a · inbound

WAM-TTT: Steering World-Action Models by Watching Human Play at Test Time cites this paper.

WAM-TTT: Steering World-Action Models by Watching Human Play at Test Time MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-13T06:48:14.554799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T06:48:14.554799Z digest=sha256:dcc23f34df1fdf96c8cf7d9218ae72ae11fb2367d473a45630d2ee464167b99c

Observation 7e1e5aeb-e2b5-440a-a5cf-0671109db37b · inbound

DC-WAM: Dynamic-Centric Visual Supervision and Reasoning for World-Action Models cites this paper.

DC-WAM: Dynamic-Centric Visual Supervision and Reasoning for World-Action Models MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T01:11:13.569491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T01:11:13.569491Z digest=sha256:aa3daa51f209d67852c2cf2045d315a9c5a8ed154b8a15160639d5d0d3d50af8

Observation 59ddcd4f-cb92-44c3-bc60-21170645685b · inbound

Enfold: Folding World Model Imagination into Predictive Representations for Ultra-Efficient Embodied Control cites this paper.

Enfold: Folding World Model Imagination into Predictive Representations for Ultra-Efficient Embodied Control MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T11:27:05.799304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:27:05.799304Z digest=sha256:e942b96edf58a1a352d89424a35d6e4ff543e32dfafc42772a88cec422b5a029

Observation 00e891fe-26a8-4d63-9025-5e753132fdd1 · inbound

Enfold: Folding World Model Imagination into Predictive Representations for Ultra-Efficient Embodied Control cites this paper.

Enfold: Folding World Model Imagination into Predictive Representations for Ultra-Efficient Embodied Control MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T03:23:47.997711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:23:47.997711Z digest=sha256:aa723063040bee6eaec9f066d8bcf2b7a4a8902ed9fc46e8a5d856ed0fad3fac

Observation 7c83ccbb-0a5b-4667-92bf-b7b875d07856 · inbound

ST-WAM: Semantic-Temporal World Action Model for Robust Manipulation under Visual Distribution Shifts cites this paper.

ST-WAM: Semantic-Temporal World Action Model for Robust Manipulation under Visual Distribution Shifts MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T15:49:34.013520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T15:49:34.013520Z digest=sha256:78af87e60542ddf5cabbe204bbf57deadf86735f0d935a7921c6a3ac90fad3e0