Pith. sign in

Paper Citation Record · LEDGER

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing?

As of 5 August 2026, this Paper Citation Record lists 99 of 99 outbound references and 8 inbound Pith citation observations for arXiv:2606.19531.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.19531 v1

Coverage vector

measured 99 of 99 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-26T21:02:24.792139Z

measured 107 of 107 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T00:44:09.563673Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-09T22:16:36.343123Z

Reference resolution

99 of 99 outbound references displayed

  • verified exact57
  • verified fuzzy0
  • unresolved39
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bb88c99b-f26e-4f06-ac68-7f8144631f14 · outbound

This paper cites Video prediction policy: A generalist robot policy with predictive visual representations.arXiv preprint, 2024.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Video prediction policy: A generalist robot policy with predictive visual representations.arXiv preprint, 2024

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-26T21:02:24.792139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:498edd138f78e479f2b4c61face15098f973d396d8cc75687cbd4c8de3b56fb3

Observation 55cf8d2c-e378-446b-aafe-a18898ff0784 · outbound

This paper cites World Action Models are Zero-shot Policies.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? World Action Models are Zero-shot Policies

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:39:17.467808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:6b94cae7257bda1e5d2516aa2a0bd7235118341b4ad857228208268246caac33

Observation 3e5fb449-95e6-44bb-a884-981bea3df0cc · outbound

This paper cites Causal World Modeling for Robot Control.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Causal World Modeling for Robot Control

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:39:17.491633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:bda70f25ef134e6b28f82adc561268d41f9839a4dbf0bb7b9b35bcc8d2f99b87

Observation 1d603c82-7b96-4ff3-9bc8-944b3ab84c57 · outbound

This paper cites Dit4dit: Jointly modeling video dynamics and actions for generalizable robot control.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Dit4dit: Jointly modeling video dynamics and actions for generalizable robot control

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:39:17.484291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:12b807449963b132d764c7026cbc6c79764200c795827ec3d469beac9ec5da1d

Observation da15809f-ac50-4c19-88ab-a4941754f2d2 · outbound

This paper cites Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:39:17.436744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:517ec67a2336d937e4edf377296e565296543264ca4ab688658cb66bbdb20379

Observation 8806fd25-8b79-4fb0-8f2d-fbc5721e0fef · outbound

This paper cites Bagelvla: Enhancing long-horizon manipulation via interleaved vision-language-action generation.CoRR, abs/2602.09849.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Bagelvla: Enhancing long-horizon manipulation via interleaved vision-language-action generation.CoRR, abs/2602.09849

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:39:17.459384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:60d14bf79bf1d1725dbd2e6e077bed6104f7bc30038c72f8d2daf22e25ee43f1

Observation 2fc7ab0e-a2cc-4f98-8656-9c603164ae31 · outbound

This paper cites UAM: A Dual-Stream Perspective on Forgetting in VLA Training.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? UAM: A Dual-Stream Perspective on Forgetting in VLA Training

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:39:17.457436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:2c96c0602926ced42ca1d04de8d922e391f5304b0c12930283af75f67fda14eb

Observation 8745c12b-469d-43bf-befc-56524f10daa2 · outbound

This paper cites AIM: Intent-Aware Unified world action Modeling with Spatial Value Maps.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? AIM: Intent-Aware Unified world action Modeling with Spatial Value Maps

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:39:17.489624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:2aa8024a3ffebbef28bff39b12a00b77e3d82bab26db144cb391a698a64d5b75

Observation 3cbd8d2b-fb12-46a5-a3b1-5336238b6068 · outbound

This paper cites Unified world models: Coupling video and action diffusion for pretraining on large robotic datasets.arXiv preprint, 2025.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Unified world models: Coupling video and action diffusion for pretraining on large robotic datasets.arXiv preprint, 2025

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-26T21:02:24.792139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:06966fd880e55d18521f2b27f3a55b972a964b7c6940981dfb1ed0154a87294b

Observation 66749db6-7b82-4394-8ced-461f35ebfd52 · outbound

This paper cites LDA-1B: Scaling Latent Dynamics Action Model via Universal Embodied Data Ingestion.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? LDA-1B: Scaling Latent Dynamics Action Model via Universal Embodied Data Ingestion

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T00:39:17.385932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:7391e7816dba30eeaa5d53cb4e469cc1499b4062712dee631a1d8b025a76dde2

Observation c2350027-27ce-45d1-a687-15dbb9000079 · outbound

This paper cites Disentangled Robot Learning via Separate Forward and Inverse Dynamics Pretraining.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Disentangled Robot Learning via Separate Forward and Inverse Dynamics Pretraining

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:39:17.374363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:a59f0ab5bdc5db9ebcf26d1cf193e75e3768fbe4a764ae1472b1bab805111c50

Observation 57a6f44a-2bea-45f8-bdf3-5c200963474c · outbound

This paper cites Motus: A Unified Latent Action World Model.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Motus: A Unified Latent Action World Model

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:39:17.374184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:1831846b8b719aa00aad5c159772447d8581fc4f23e5f10d363fa762bf7e45d2

Observation 25347798-464b-4aa2-8c5a-9e67c39771b9 · outbound

This paper cites Fast-WAM: Do World Action Models Need Test-time Future Imagination?.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Fast-WAM: Do World Action Models Need Test-time Future Imagination?

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:39:17.345770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:3a9158d9cf1920d5fbf2ebe8b6a67856798322fc13e599dbb0e8ed49065ac6df

Observation 97fae678-9072-4b65-9a4f-26a728759995 · outbound

This paper cites arXiv preprint arXiv:2603.17240 , year=.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? arXiv preprint arXiv:2603.17240 , year=

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:39:17.410662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:66a6d6e1a6c8cfb1e5323acc6612b7caad417fa95f0bec83c3a72251a9c69469

Observation fc6a206d-370a-47f8-8eaf-08b9c22d2e2d · outbound

This paper cites MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:39:17.382706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:e97756e642f3e60d365867cf8529f888750e5189843784b465223d16277d2435

Observation 5c779dd0-6daf-4f43-895b-811706055d88 · outbound

This paper cites Reworld: Multi-dimensional reward modeling for embodied world models.arXiv preprint arXiv:2601.12428.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Reworld: Multi-dimensional reward modeling for embodied world models.arXiv preprint arXiv:2601.12428

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:39:17.379964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:68d41e2ed37121eda7dcb18f43c0bbbfb03fc4ef4f7f97cd5f831c2de5853bbc

Observation 4468830d-c137-4bd6-a2b7-526e744b2ea5 · outbound

This paper cites Orv: 4d occupancy-centric robot video generation.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Orv: 4d occupancy-centric robot video generation

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:39:17.318561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:50e21f29380eccea7c356e2b01a24a305b98edc40c2c4e161087015212b482ed

Observation 8a7e792b-03f0-4fe0-ae1e-fc9c18abf86b · outbound

This paper cites TesserAct: Learning 4D Embodied World Models.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? TesserAct: Learning 4D Embodied World Models

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T00:39:17.407562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:147dad3c35c336990c4f2b6874a6d0d81435cb8fa20abb4965222b8ff2bd2e0c

Observation 09a1f3a2-efba-441a-bf4d-3066a5bc7ecb · outbound

This paper cites Scene graph disentanglement and composition for generalizable complex image generation.Advances in Neural Information Processing Systems, 37:98478–98504, 2024.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Scene graph disentanglement and composition for generalizable complex image generation.Advances in Neural Information Processing Systems, 37:98478–98504, 2024

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-26T21:02:24.792139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:087072d50f8dc94f4ccd61d417091a1e565bd841585fd86c4f9bf2dd1e9e4bd6

Observation 68c14e92-5f93-4d71-89ab-50c2bed12ba4 · outbound

This paper cites Nano banana pro.https://deepmind.google/technologies/gemini/, 2025.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Nano banana pro.https://deepmind.google/technologies/gemini/, 2025

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-26T21:02:24.792139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:d02f626d5f59c30ee4fb1b9bfe3440876d4b19c55c612cfe1986a67f504f0fe1

Observation 4f02af53-4df3-4ecf-a7f9-c2c9f8353a79 · outbound

This paper cites GPT-Image-1.5.https://openai.com/index/new-chatgpt-images-is-here/, 2026.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? GPT-Image-1.5.https://openai.com/index/new-chatgpt-images-is-here/, 2026

Reference 21

Resolution
unresolved
no resolver link, observed 2026-06-26T21:02:24.792139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:0878c558cb418985f9ec4620055aaa0ddc221fae47f3ca96ee7a3e61fd82ce5f

Observation 4dfacc50-3044-436e-b1a1-c696d17de981 · outbound

This paper cites ImgEdit: A Unified Image Editing Dataset and Benchmark.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? ImgEdit: A Unified Image Editing Dataset and Benchmark

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:39:17.358514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:b96363322003a6c2b304f2a800b7c3201b8eadfa3abda0f2bf02f2790b730c0d

Observation d3fd09fc-a1e0-4db5-9927-11331ab8be47 · outbound

This paper cites Qwen-Image Technical Report.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Qwen-Image Technical Report

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:39:17.443553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:5dd86376b068f31172b029c9ae9b16955856030c674bd017ad9a11354346bcb8

Observation e5ce8f5c-1ada-4a42-af99-6ebfc034226a · outbound

This paper cites Glm-image.https://huggingface.co/zai-org/GLM-Image, 2026.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Glm-image.https://huggingface.co/zai-org/GLM-Image, 2026

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-26T21:02:24.792139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:ccfe84a6c73e8df689b7d151b8928674c3f9faeb4fa202beb128d65a3adaf22c

Observation 91eb9690-007e-4ce1-8383-78c7d784bd25 · outbound

This paper cites NextStep-1: Toward Autoregressive Image Generation with Continuous Tokens at Scale.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? NextStep-1: Toward Autoregressive Image Generation with Continuous Tokens at Scale

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:39:17.419935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:76d562c78fbc0a068c290b7337e236faccf2c69ca1e1bae852ae61cb7020d948

Observation 7b80f4f7-d35e-47e2-84bf-ab3887ecc7b9 · outbound

This paper cites Longcat-next: Lexicalizing modalities as discrete tokens.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Longcat-next: Lexicalizing modalities as discrete tokens

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:39:17.481511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:e80a527f62ed6d2e3345cf412d84ab419b75ef916aede274ae952eadd7d69e17

Observation c5aaa63f-f241-4dcf-89ff-08905dbda029 · outbound

This paper cites Uni-Edit: Intelligent Editing Is A General Task For Unified Model Tuning.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Uni-Edit: Intelligent Editing Is A General Task For Unified Model Tuning

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:39:17.454286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:6bfc3cf49ff637bcc2c3bef042ae82868e3e8de9ca4f2ccb406973fdb0f188a2

Observation d7782d6b-c758-40a0-8a40-0e99d0d7f834 · outbound

This paper cites Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:39:17.489219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:9196db861230ecdba9eacdbac1eaba56add8a598a9c0f7dd307659c30ed5e6c3

Observation 3b2a4355-4f7a-4393-b59d-445e17d75e08 · outbound

This paper cites Magicbrush: A manually annotated dataset for instruction-guided image editing.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Magicbrush: A manually annotated dataset for instruction-guided image editing

Reference 29

Resolution
unresolved
no resolver link, observed 2026-06-26T21:02:24.792139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:ed6b1bbe25319d109383409ee2bf55d4ab61b9f45aaf0c7aa929954ebfd4913d

Observation 1a9bca96-0817-4ad8-8670-8541b39e6314 · outbound

This paper cites Guiding instruction-based image editing via multimodal large language models.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Guiding instruction-based image editing via multimodal large language models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-06-26T21:02:24.792139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:749323cf7a4a63635e0f5848d1e8bda1ac5e34931969c3a0540eaac0f2c585b9

Observation 85834bbe-24b8-4385-886f-b331ea5d66a2 · outbound

This paper cites Emu edit: Precise image editing via recognition and generation tasks.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Emu edit: Precise image editing via recognition and generation tasks

Reference 31

Resolution
unresolved
no resolver link, observed 2026-06-26T21:02:24.792139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:6878dcf0f422a642119d3d06ec68e5149987fdcd5e4b83128cd5bee634424a40

Observation 60268c47-cab4-4bd4-832b-ed8b01614dc6 · outbound

This paper cites Anyedit: Mastering unified high-quality image editing for any idea.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Anyedit: Mastering unified high-quality image editing for any idea

Reference 32

Resolution
unresolved
no resolver link, observed 2026-06-26T21:02:24.792139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:b99d86daa88c00de71eb552015ccea22aa44bc3248d98b54beae074dd39afdc9

Observation afc015af-0f20-46b7-b41d-27add10fb77a · outbound

This paper cites Image Generators are Generalist Vision Learners.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Image Generators are Generalist Vision Learners

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:39:17.438144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:6e27d905884c752f9388414dccbc33593cb6a5ef15466484822128daacafbf99

Observation d4147ffa-7d69-4dc8-95e0-276839c1df09 · outbound

This paper cites Diffusion Model as a Generalist Segmentation Learner.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Diffusion Model as a Generalist Segmentation Learner

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:39:17.422990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:2b3c140aa6f774ba503a41bd1bf37d4b44af8a6f5a529189f1c2b9b7289e91bd

Observation 526d87a1-7c32-4881-a54b-ddfdd7862d4b · outbound

This paper cites Leveraging Image Generators to Address Data Scarcity: The Gen4Regen Dataset for Forest Regeneration Mapping.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Leveraging Image Generators to Address Data Scarcity: The Gen4Regen Dataset for Forest Regeneration Mapping

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:39:17.416927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:f7c3df7dc354bb33f3fdddb8deb194f2adac057b1ec4520eba7830f928c56627

Observation 0509c2cc-1de9-4542-875d-ce00f9e75795 · outbound

This paper cites pi0: A vision-language-action flow model for general robot control.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? pi0: A vision-language-action flow model for general robot control

Reference 36

Resolution
unresolved
no resolver link, observed 2026-06-26T21:02:24.792139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:3532094ec93ed7c60893e55f43c8639d38168e382d5f4dafb456ebe3954664ce

Observation 375cd041-792b-41f4-8889-e89e8caae296 · outbound

This paper cites pi0.5: a vision-language-action model with open-world generalization.arXiv preprint, 2025.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? pi0.5: a vision-language-action model with open-world generalization.arXiv preprint, 2025

Reference 37

Resolution
unresolved
no resolver link, observed 2026-06-26T21:02:24.792139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:469f1bcdebbacfee775964756468668cf590d0689e789f8482fe53986e85e82a

Observation 105870cd-06ae-4c3c-9d6b-c3ca971de2ec · outbound

This paper cites Gr00t n1: An open foundation model for generalist humanoid robots.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Gr00t n1: An open foundation model for generalist humanoid robots

Reference 38

Resolution
unresolved
no resolver link, observed 2026-06-26T21:02:24.792139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:06086bfabd2a1ca38cd1155c3ac7fc421797c929e8836f6722c95bdcb6df454d

Observation fc0dc4f3-9207-42c1-bbb4-4f0c6824d81f · outbound

This paper cites Dreamvla: A vision-language-action model dreamed with comprehensive world knowledge.arXiv preprint, 2025.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Dreamvla: A vision-language-action model dreamed with comprehensive world knowledge.arXiv preprint, 2025

Reference 39

Resolution
unresolved
no resolver link, observed 2026-06-26T21:02:24.792139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:a3581539edfcbaa7f6e9133ac171a1ea503284a2f9ef29971f8578c56279027e

Observation fc2f5178-863e-4215-b1f6-ba371fc13ec7 · outbound

This paper cites Reconvla: Reconstructive vision-language-action model as effective robot perceiver.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Reconvla: Reconstructive vision-language-action model as effective robot perceiver

Reference 40

Resolution
unresolved
no resolver link, observed 2026-06-26T21:02:24.792139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:93a3d7d031910f604c102566dbd2e5175cd5943d04f2df2c4d7be8848143ee33

Observation c953e55b-3a7a-43a8-aae7-6005ab8d20f8 · outbound

This paper cites HY-Embodied-0.5: Embodied Foundation Models for Real-World Agents.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? HY-Embodied-0.5: Embodied Foundation Models for Real-World Agents

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:39:17.448929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:425fc2e25c99bff7b93c2e9c2826f6241a819f7a05e74608a997b5de39bef4f0

Observation db88cc04-c30a-4c28-864c-8cfed2513b62 · outbound

This paper cites PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:39:17.425754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:ad72251d0cbc6a1790182a9f322f1157da34694d0388821ad99e3d3da5823d14

Observation 4928c709-4293-4dd2-a8c9-ae42f4f9e16b · outbound

This paper cites Depthvla: Enhancing vision-language-action models with depth-aware spatial reasoning.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Depthvla: Enhancing vision-language-action models with depth-aware spatial reasoning

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:39:17.351632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:b3aeccfdf47dcf187b4f163bcfbf9f68b20159149f8caf98fb9595cad8148bd7

Observation 9e20d433-8483-424e-8815-44a13802bb41 · outbound

This paper cites Spatialvla: Exploring spatial representations for visual-language-action model.arXiv preprint, 2025.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Spatialvla: Exploring spatial representations for visual-language-action model.arXiv preprint, 2025

Reference 44

Resolution
unresolved
no resolver link, observed 2026-06-26T21:02:24.792139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:ae614510bc1fb33349edc54d135321067643bd7eb782dc9a79e9115453e2a012

Observation e7488449-f3a2-412a-8bc5-3ffce10961a3 · outbound

This paper cites Predictive inverse dynamics models are scalable learners for robotic manipulation.ICLR, 2024.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Predictive inverse dynamics models are scalable learners for robotic manipulation.ICLR, 2024

Reference 45

Resolution
unresolved
no resolver link, observed 2026-06-26T21:02:24.792139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:52ec7dd03756e09cef3e795e54294d3b94f2f5299b4b94c54c176b7b42d9ba72

Observation 0c0567e8-f909-4605-aefe-da49d1e3874a · outbound

This paper cites Cot-vla: Visual chain-of-thought reasoning for vision-language-action models.arXiv preprint, 2025.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Cot-vla: Visual chain-of-thought reasoning for vision-language-action models.arXiv preprint, 2025

Reference 46

Resolution
unresolved
no resolver link, observed 2026-06-26T21:02:24.792139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:cd922b68ef537a6cfe5539d67baf1a301ba87a00c03a09cd5fe58790ed2f96aa

Observation 8d3d711b-191d-4090-a626-c0bb205772f8 · outbound

This paper cites Transport Discrepancy as a Reliability Signal for Vision-Language-Action Models.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Transport Discrepancy as a Reliability Signal for Vision-Language-Action Models

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:39:17.496728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:efdcbf51b5c2f57e98568dd05b9d910b39886b8d35e8d3a3bbe4ba5063617d3a

Observation 3f43bd6c-0aa9-4096-9776-af5ce84dcc5b · outbound

This paper cites Being-h0: Vision-language-action pretraining from large-scale human videos.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Being-h0: Vision-language-action pretraining from large-scale human videos

Reference 48

Resolution
unresolved
no resolver link, observed 2026-06-26T21:02:24.792139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:60c0e0b2775b94f4d39d5317f851862795f1f71251f8320bac2efa904742ea22

Observation 61d533a4-7712-4317-b84c-043e57ea2e73 · outbound

This paper cites Unified diffusion vla: Vision-language-action model via joint discrete denoising diffusion process.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Unified diffusion vla: Vision-language-action model via joint discrete denoising diffusion process

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:39:17.352756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:18622e6ea3609fa100df076eb4ed2a2b763053368f136bdfe55217ef23785857

Observation ab6841ed-a7aa-46b8-889c-cd84731cb13b · outbound

This paper cites Spatial forcing: Implicit spatial representation alignment for vision- language-action model.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Spatial forcing: Implicit spatial representation alignment for vision- language-action model

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:39:17.361427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:911b1ae19ffbaaba60e3e211089ae887c9c8aa3cff4b8d36533e78414a458ade

Observation 7304a538-edaa-4395-b132-608c61fc49b8 · outbound

This paper cites Vla-jepa: Enhancing vision-language-action model with latent world model.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Vla-jepa: Enhancing vision-language-action model with latent world model

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:39:17.494341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:531882373baac56dae928b1eb8f5427e7a6eec60b420205554aee1eaf7345baf

Observation d5cdf597-1f76-4d68-b72a-c0f1332bb108 · outbound

This paper cites Vla-adapter: An effective paradigm for tiny-scale vision-language-action model.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Vla-adapter: An effective paradigm for tiny-scale vision-language-action model

Reference 52

Resolution
unresolved
no resolver link, observed 2026-06-26T21:02:24.792139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:b4f005b71d34fa39313bd8bb8c686a6248246618fb2156c8cfa2a1e7366fdc5f

Observation a2bb38f6-3ddc-47fa-8f33-a49eaf5c2754 · outbound

This paper cites A Pragmatic VLA Foundation Model.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? A Pragmatic VLA Foundation Model

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:39:17.379412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:afc2d809591e07b5bc3d36414547fab85bbdeb974468923316c8395af420957c

Observation 69cf02e9-7140-473b-b6ba-a15b7c624c4a · outbound

This paper cites MolmoAct: Action Reasoning Models that can Reason in Space.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? MolmoAct: Action Reasoning Models that can Reason in Space

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:39:17.411368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:7d5b04ccd4f070cd865a3cf5dd605455051d504ff52bcdffcca3a670f9cbe066

Observation 29857894-4f8e-42cd-9a14-732efa9d282d · outbound

This paper cites F1: A Vision-Language-Action Model Bridging Understanding and Generation to Actions.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? F1: A Vision-Language-Action Model Bridging Understanding and Generation to Actions

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:39:17.372041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:e3098b316749db176ba0c56cf6772292f65e8bc7077c37ffa9887bbb612eaa2e

Observation f1f3d92f-c05d-4754-b3e8-4e56ea20eac5 · outbound

This paper cites Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:39:17.315926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:9193d0caaa2d3d7b974498bbb5f2c5fae12e23559b17cb1a43a229807b2a20a9

Observation 11fa9e64-fa22-4e49-a0f3-d2cc74c18463 · outbound

This paper cites Seeing to act, prompting to specify: A bayesian factorization of vision language action policy.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Seeing to act, prompting to specify: A bayesian factorization of vision language action policy

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:39:17.377234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:0123727507f7d5a45e582727a0166fb9c92bec8b3d960add12000bea17013d00

Observation 6056000c-c1ab-41ad-b785-3109ca2cd912 · outbound

This paper cites Learning universal policies via text-guided video generation.NeurIPS, 2024.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Learning universal policies via text-guided video generation.NeurIPS, 2024

Reference 58

Resolution
unresolved
no resolver link, observed 2026-06-26T21:02:24.792139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:847d1faf76a40525175ee0a13e47d99beb645be8e827c580625b30a59cfd35da

Observation 784b60f0-9eeb-43f2-8e59-4863e28530af · outbound

This paper cites Zero-shot robotic manipulation with pretrained image-editing diffusion models.arXiv preprint, 2023.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Zero-shot robotic manipulation with pretrained image-editing diffusion models.arXiv preprint, 2023

Reference 59

Resolution
unresolved
no resolver link, observed 2026-06-26T21:02:24.792139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:ecd8035fb79361a9e20a01cfc90e64d8800a927019353fad1722699858a07a70

Observation 78d94391-2ea9-404c-a55d-9f6700deb43d · outbound

This paper cites Generalist bimanual manipulation via foundation video diffusion models.arXiv preprint, 2025.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Generalist bimanual manipulation via foundation video diffusion models.arXiv preprint, 2025

Reference 60

Resolution
unresolved
no resolver link, observed 2026-06-26T21:02:24.792139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:72ca6b45612774ac7378400c0ca778b76f8602b811bc09b72ce0857ca47199ec

Observation 7b5e2f14-2e13-463d-9e72-a8bf5c2c14b8 · outbound

This paper cites Vidman: Exploiting implicit dynamics from video diffusion model for effective robot manipulation.NeurIPS, 2024.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Vidman: Exploiting implicit dynamics from video diffusion model for effective robot manipulation.NeurIPS, 2024

Reference 61

Resolution
unresolved
no resolver link, observed 2026-06-26T21:02:24.792139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:52e39834cac0e6291d420885c343c01fa78603dce490c36c5c4032744634401c

Observation 74fcc968-2799-4902-913c-45a3397f77dd · outbound

This paper cites Murphy, Chelsea Finn, and Yilun Du.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Murphy, Chelsea Finn, and Yilun Du

Reference 62

Resolution
unresolved
no resolver link, observed 2026-06-26T21:02:24.792139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:8c14c078f167837be3d77a8a6f74c25817f682bee75984a1e39e9ee5a1cb1292

Observation a78fae03-fe60-4ad9-a60c-293f8a4c3896 · outbound

This paper cites Large Video Planner Enables Generalizable Robot Control.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Large Video Planner Enables Generalizable Robot Control

Reference 63

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T00:39:17.321453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:4235c5cf69de0800601f4141911a7439c9b2de5c543cca5a88669316287a0ed1

Observation 00920460-fc74-4407-8332-ed14b0520ee3 · outbound

This paper cites Anypos: Automated task-agnostic actions for bimanual manipulation.arXiv preprint, 2025.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Anypos: Automated task-agnostic actions for bimanual manipulation.arXiv preprint, 2025

Reference 64

Resolution
unresolved
no resolver link, observed 2026-06-26T21:02:24.792139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:1c9d8481744fb47a0b90c69019f83b3d9ec28c308c087326409c48bb010f48e3

Observation de19a9f3-fd6e-47e8-9c07-45022c798637 · outbound

This paper cites Tc-idm: Grounding video generation for executable zero-shot robot motion.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Tc-idm: Grounding video generation for executable zero-shot robot motion

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:39:17.428620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:791642a27524275440be86c7da300d8c7d554b506ee6a359a39ec5a1257275d0

Observation e58da869-edcd-492c-95ea-0f9fb0fce6f7 · outbound

This paper cites Veo-act: How far can frontier video models advance generalizable robot manipulation? 2026.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Veo-act: How far can frontier video models advance generalizable robot manipulation? 2026

Reference 66

Resolution
unresolved
no resolver link, observed 2026-06-26T21:02:24.792139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:4c3f6f3b30060e0489497de14c8173a82f111a6d6a653e958b0b13b09bd75806

Observation d98d9d2e-6f0c-4044-900b-1bf18c42b8cd · outbound

This paper cites V AMPO: Policy optimization for improving visual dynamics in video action models.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? V AMPO: Policy optimization for improving visual dynamics in video action models

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:39:17.394923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:0fb68a6ef82714aaba1d74cd2802e6efb2e48bd8227531b760a1726b70cf6fac

Observation fd943eb4-7439-4c96-bd18-58f8cf3a2319 · outbound

This paper cites Do World Action Models Generalize Better than VLAs? A Robustness Study.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Do World Action Models Generalize Better than VLAs? A Robustness Study

Reference 68

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:39:17.427944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:204b86069dbdd78c848fd8ec6128aa9434940455eb39e802a7007d510e5afe09

Observation 3448e806-290e-4388-83ae-79aea7606673 · outbound

This paper cites WorldEval: World Model as Real-World Robot Policies Evaluator.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? WorldEval: World Model as Real-World Robot Policies Evaluator

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:39:17.498204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:dbc4617d70ddee262a793c04e6701168c3518bde844b920d5304ec43335b7719

Observation 0aa67ed4-4f61-4427-9ba1-0dd92c6fa43a · outbound

This paper cites Kinema4d: Kinematic 4d world modeling for spatiotemporal embodied simulation, 2026.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Kinema4d: Kinematic 4d world modeling for spatiotemporal embodied simulation, 2026

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:39:17.486632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:0a14682e52669343f3b0809ad48e307387008ed17fe606881dc21c98092f79ec

Observation 935e3104-67d5-4bb1-ae65-53764a4d2599 · outbound

This paper cites WoVR: World Models as Reliable Simulators for Post-Training VLA Policies with RL.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? WoVR: World Models as Reliable Simulators for Post-Training VLA Policies with RL

Reference 71

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:39:17.433965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:0189b1b019a89c2212f77f7b392ed0f1f17e6c9fb1b3a653fa04ff52af4627ee

Observation dfbbef40-6a09-4aa3-89ed-07011a2a21aa · outbound

This paper cites RoboStereo: Dual-Tower 4D Embodied World Models for Unified Policy Optimization.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? RoboStereo: Dual-Tower 4D Embodied World Models for Unified Policy Optimization

Reference 72

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:39:17.445163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:509d6f9e2ed1ad5bc09ae2a3d2570dd962e9adfbc524deb0cdc9d22123724367

Observation 4ff1ec73-74e3-4a7d-b10c-126d3e147026 · outbound

This paper cites UniT: Toward a Unified Physical Language for Human-to-Humanoid Policy Learning and World Modeling.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? UniT: Toward a Unified Physical Language for Human-to-Humanoid Policy Learning and World Modeling

Reference 73

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:39:17.465109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:075b322f0868498181b4d7e8abb7ff717a8e887dc9221939685c333f21b789aa

Observation 905869cc-0d2d-42ac-b5a1-f95b68dda1dc · outbound

This paper cites Persistent robot world models: Stabilizing multi- step rollouts via reinforcement learning.arXiv preprint arXiv:2603.25685, 2026.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Persistent robot world models: Stabilizing multi- step rollouts via reinforcement learning.arXiv preprint arXiv:2603.25685, 2026

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:39:17.414015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:0f78b6eef6c8d63dfd76958d6905ddd0b0dd29cdd8084b1f331566522fc54ccc

Observation de531357-d7d4-4ba2-8563-86cfc7f0f692 · outbound

This paper cites Fate: Closed-loop feasibility-aware task generation with active repair for physically grounded robotic curricula.arXiv preprint arXiv:2603.01505, 2026.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Fate: Closed-loop feasibility-aware task generation with active repair for physically grounded robotic curricula.arXiv preprint arXiv:2603.01505, 2026

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:39:17.451736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:33e41310a012abc049af7748264592b173857c33b3ef6f9b58fa08eb3aeba97a

Observation 0f438a2f-b26c-4524-8e23-1d02b2425c4a · outbound

This paper cites VAG: Dual-Stream Video-Action Generation for Embodied Data Synthesis.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? VAG: Dual-Stream Video-Action Generation for Embodied Data Synthesis

Reference 76

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:39:17.456675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:c6f2feff279a28846c5a26fddb3e8d578669f182c4834958ba971b2b22402b21

Observation c7a70452-d578-409e-a32a-a9c6d08002a6 · outbound

This paper cites Interactive world simulator for robot policy training and evaluation.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Interactive world simulator for robot policy training and evaluation

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:39:17.451543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:66cbbbf60e4343d5ffeb65faa59c03f243ca5d94cab21ff1b2a999dcbcee51fb

Observation 57c7b703-25db-4956-9d2c-87d108d8d84e · outbound

This paper cites World Action Verifier: Self-Improving World Models via Forward-Inverse Asymmetry.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? World Action Verifier: Self-Improving World Models via Forward-Inverse Asymmetry

Reference 78

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:39:17.462991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:74abf106fc0752e0e39a5b9b099c3c27463d91ee46a2c3f90aa658a250c18ec7

Observation e6c788ca-136a-4193-a8b3-8ecb38839b1b · outbound

This paper cites World-Value-Action Model: Implicit Planning for Vision-Language-Action Systems.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? World-Value-Action Model: Implicit Planning for Vision-Language-Action Systems

Reference 79

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:39:17.470339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:5b495e7dc1984b0a9be8424517f2e8a6954a2ee45b263e8d9a0dee7a283f03cc

Observation bbbb3e9b-cda3-43ae-a63d-5bd3d44b3f2b · outbound

This paper cites Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation

Reference 80

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:39:17.499177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:1d9bc21c55d6043fe6e408bef17d55a23fffe7b0154beeca98df6493bbcbe893

Observation 2251de90-0eb7-4782-8ab7-a819a181c546 · outbound

This paper cites dworldeval: Scalable robotic policy evaluation via discrete diffusion world model.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? dworldeval: Scalable robotic policy evaluation via discrete diffusion world model

Reference 81

Resolution
unresolved
no resolver link, observed 2026-06-26T21:02:24.792139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:a5f5b370ffd585ac01d557921e7ec60c0ce30ed0544c153ecf89a22cb39ed800

Observation 68440a13-9d8c-410b-9de8-5c82605775cf · outbound

This paper cites Interactive world simulator for robot policy training and evaluation.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Interactive world simulator for robot policy training and evaluation

Reference 82

Resolution
unresolved
no resolver link, observed 2026-06-26T21:02:24.792139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:b2aa863916e449916af3068788fd9bffc029f69b994a614cc65c328a4e2d9a34

Observation e537418e-723c-4118-8f8f-2c213335ddbe · outbound

This paper cites an unresolved cited work.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Unresolved cited work

Reference 83

Resolution
unresolved
no resolver link, observed 2026-06-26T21:02:24.792139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:ad12da3b4873f3d35c787eb91acc3959c81da6339748429038e56a4bbc95359d

Observation 891ebd45-3d91-4e59-b077-7546724672b3 · outbound

This paper cites Cosmos 3: Omnimodal World Models for Physical AI.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Cosmos 3: Omnimodal World Models for Physical AI

Reference 84

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:39:17.503977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:ac2e7a365dfe2e77fa557e4947327ba2363c76afc87dc040e70ee5da32f28020

Observation db9e670c-aa8c-4e7f-9acc-74fd93acba5c · outbound

This paper cites OmniGen2: Towards Instruction-Aligned Multimodal Generation.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 85

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:39:17.425424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:3b854e3d4864e476eed60f279b7f067f080dd88b5ef4772f6a09c41a917ddc90

Observation 86578b31-29d0-469b-b3f0-e0c6be404263 · outbound

This paper cites Ovis-U1 Technical Report.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Ovis-U1 Technical Report

Reference 86

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:39:17.483974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:63ccd6bc75d06b9c722c88e549d04cd5a41532aff95e10cc75fc90b05deb1818

Observation 6daf48e1-c49d-4601-b795-4e10bd1bd6ae · outbound

This paper cites FLUX.2: Frontier Visual Intelligence.https://bfl.ai/blog/flux-2, 2025.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? FLUX.2: Frontier Visual Intelligence.https://bfl.ai/blog/flux-2, 2025

Reference 87

Resolution
unresolved
no resolver link, observed 2026-06-26T21:02:24.792139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:3cc6b9433dd75358361c16235994a84190c64184121f71f6a366838cc5d0bec0

Observation f66e74c9-5128-42bb-a262-ac921091081c · outbound

This paper cites Libero: Benchmarking knowledge transfer for lifelong robot learning.arXiv preprint, 2023.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Libero: Benchmarking knowledge transfer for lifelong robot learning.arXiv preprint, 2023

Reference 88

Resolution
unresolved
no resolver link, observed 2026-06-26T21:02:24.792139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:caea122e02c65a0d6120a8b85cd8a15808ddc7b4e33b15aed88f4b0820696383

Observation fcbcddc8-c2a0-40b7-84c6-e42809dcef7b · outbound

This paper cites LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models

Reference 89

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:39:17.501582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:82a6d44a8aa404e343eba1a218dc26e723ea53e203379815297672edcbedd55c

Observation bb81c9a5-2dbb-4999-bd92-34e297a3d90a · outbound

This paper cites RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation

Reference 90

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:39:17.478920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:a143c5a7e6c31ca78500b692a6a717678fa884a9323ac55e619b9d695a7d4a08

Observation 7bd0c31e-2336-4fc0-a92b-09e08f8e7897 · outbound

This paper cites ABot-M0: VLA Foundation Model for Robotic Manipulation with Action Manifold Learning.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? ABot-M0: VLA Foundation Model for Robotic Manipulation with Action Manifold Learning

Reference 91

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:39:17.476110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:aa3e11eccafb81f3e5c1401ea8704db35e5b28a21e828ea74bb1fa8f2b54258c

Observation 358e5826-1b6d-4992-a9c2-c95f290d10b8 · outbound

This paper cites Openvla: An open-source vision-language-action model.arXiv preprint, 2024.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Openvla: An open-source vision-language-action model.arXiv preprint, 2024

Reference 92

Resolution
unresolved
no resolver link, observed 2026-06-26T21:02:24.792139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:8b2bce8a2f5adb6e630b255f9731b6e3026e20b4d10b2e77252e4192946ed081

Observation 9f6d3ece-d196-458e-bb2b-9f30247be5bd · outbound

This paper cites LIBERO: benchmarking knowledge transfer for lifelong robot learning.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? LIBERO: benchmarking knowledge transfer for lifelong robot learning

Reference 93

Resolution
unresolved
no resolver link, observed 2026-06-26T21:02:24.792139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:71abed9b9e1939d875ffc9440a1c8c339324c47e5a92cec7b9a40a260bc82f5a

Observation 9661e013-9c30-4787-b8ba-c16769e88032 · outbound

This paper cites Univla: Learning to act anywhere with task-centric latent actions.arXiv preprint, 2025.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Univla: Learning to act anywhere with task-centric latent actions.arXiv preprint, 2025

Reference 94

Resolution
unresolved
no resolver link, observed 2026-06-26T21:02:24.792139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:5fac7db53c8b3284a3734d476621a9e4fce8af3e142540e2bbbf392d10cc8ddf

Observation 6df536dd-5ba4-4a2e-860b-8f92479e53a7 · outbound

This paper cites Fine-tuning vision-language-action models: Optimizing speed and success.arXiv preprint, 2025.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Fine-tuning vision-language-action models: Optimizing speed and success.arXiv preprint, 2025

Reference 95

Resolution
unresolved
no resolver link, observed 2026-06-26T21:02:24.792139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:6c5c39c60ea2a064c9c9516ed8f77659862793127eb94377b0eb75434b17bbd5

Observation f6a9cb63-31cb-4abd-b7ca-7bf46aea9909 · outbound

This paper cites Fast: Efficient action tokenization for vision-language-action models.arXiv preprint, 2025.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Fast: Efficient action tokenization for vision-language-action models.arXiv preprint, 2025

Reference 96

Resolution
unresolved
no resolver link, observed 2026-06-26T21:02:24.792139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:dfa32d49e0483084411c2387e8a0d0491d3eba93c1274be20a0304a00211d816

Observation fb9c0662-b8c2-4117-9d68-c1dc0dfaf584 · outbound

This paper cites Worldvla: Towards autoregressive action world model.arXiv preprint, 2025.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Worldvla: Towards autoregressive action world model.arXiv preprint, 2025

Reference 97

Resolution
unresolved
no resolver link, observed 2026-06-26T21:02:24.792139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:398945f11ec4c3fe53bc278b1d4229a1f1e435e71f91882d71c57f536a27e353

Observation c25c564c-8273-4211-9a7e-a372c55ad769 · outbound

This paper cites Unified Vision-Language-Action Model.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Unified Vision-Language-Action Model

Reference 98

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:39:17.478757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:9d6bb340cd38d79c9e7ec9442dbf4b9bd7da9015805e8d63167439f05fd96b51

Observation e1f2d5fa-4b5a-461b-9c8a-c65358a4e9a2 · outbound

This paper cites Representationalignmentforgeneration: Trainingdiffusiontransformersiseasierthanyouthink.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Representationalignmentforgeneration: Trainingdiffusiontransformersiseasierthanyouthink

Reference 99

Resolution
unresolved
no resolver link, observed 2026-06-26T21:02:24.792139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:745ddc66e6ce5ef650e0ff1881763a1ecad07a8f7b9f3a9da64885b2cb619c2d

Pith citing papers

Observation db64921a-7e5f-4e69-8e4a-b4f9d1a4cb15 · inbound

ABot-M0.5: Unified Mobility-and-Manipulation World Action Model cites this paper.

ABot-M0.5: Unified Mobility-and-Manipulation World Action Model ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing?

Reference 77

Resolution
verified exact
local_arxiv, observed 2026-07-02T14:27:03.166204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-02T14:24:23.187164Z digest=sha256:934a26431941ecbacbfeea7ef22d072b63411e58a8423ec8f48ecb3a055291e1

Observation 65c4f8c6-d1ef-47d5-a90f-41e7a1289e64 · inbound

ABot-M0.5: Unified Mobility-and-Manipulation World Action Model cites this paper.

ABot-M0.5: Unified Mobility-and-Manipulation World Action Model ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing?

Reference 77

Resolution
unresolved
no resolver link, observed 2026-07-12T09:22:08.000379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:22:08.000379Z digest=sha256:78d068bd33d3bdc4a24f79e43313542d801de917375bc955d72c2ca21793d306

Observation 71d2ccd4-84eb-47f8-b566-dfded78279e5 · inbound

WAM-TTT: Steering World-Action Models by Watching Human Play at Test Time cites this paper.

WAM-TTT: Steering World-Action Models by Watching Human Play at Test Time ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing?

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-07-09T22:16:36.344563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T22:16:31.529359Z digest=sha256:9f5a16429da51839dc475721ae4c6569b003eac4c5cede6ce9700f047846bf61

Observation ec7efade-5d5f-4b4f-ae57-6a1d700fdc21 · inbound

WAM-TTT: Steering World-Action Models by Watching Human Play at Test Time cites this paper.

WAM-TTT: Steering World-Action Models by Watching Human Play at Test Time ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing?

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-13T06:48:14.554799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T06:48:14.554799Z digest=sha256:9c54c4e72421fe5e9610d09cc41e34c1455b990ee50e92c78285590cd5b5a1ed

Observation 95e3db30-fa90-42e0-8294-3544737834fb · inbound

LeapBot-WA: World-Anchor Action Models via Predictive Latent Alignments cites this paper.

LeapBot-WA: World-Anchor Action Models via Predictive Latent Alignments ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing?

Reference 97

Resolution
unresolved
no resolver link, observed 2026-07-31T23:29:15.579570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T23:29:15.579570Z digest=sha256:72c2e98961998f39dae6dcd93cd1ad71a1594d567a36caf84e195af40ddc09c9

Observation 671ec938-b253-4260-961e-f997e3e5d774 · inbound

Enfold: Folding World Model Imagination into Predictive Representations for Ultra-Efficient Embodied Control cites this paper.

Enfold: Folding World Model Imagination into Predictive Representations for Ultra-Efficient Embodied Control ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing?

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T11:27:05.681731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:27:05.681731Z digest=sha256:8f3c6e99e53f329dd92b47a084130387b692844a13651b5301d35f62643ba88d

Observation 95b61b4f-5aed-4deb-8e15-66783e7c5618 · inbound

Enfold: Folding World Model Imagination into Predictive Representations for Ultra-Efficient Embodied Control cites this paper.

Enfold: Folding World Model Imagination into Predictive Representations for Ultra-Efficient Embodied Control ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing?

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T03:23:44.956053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:23:44.956053Z digest=sha256:3944914861070e1ae392c8b6d695a92ebff0ee387ee24ea5ae8a39fd56c26e95

Observation c35ee440-b94a-4d52-af34-7428fc87f338 · inbound

Disentangling Visuo-Tactile Foresight: Oracle-Guided Interface Discovery for World Action Models cites this paper.

Disentangling Visuo-Tactile Foresight: Oracle-Guided Interface Discovery for World Action Models ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing?

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T00:44:09.563673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T00:44:09.563673Z digest=sha256:e556fd3b22db51c34b9d3b9c1134a5bdef1a87547fe69c2e5f57ecbe311d669f