Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-26T21:02:24.792139Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 99 of 99 outbound references and 8 inbound Pith citation observations for arXiv:2606.19531.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-26T21:02:24.792139Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-05T00:44:09.563673Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-09T22:16:36.343123Z
99 of 99 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation bb88c99b-f26e-4f06-ac68-7f8144631f14 · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Video prediction policy: A generalist robot policy with predictive visual representations.arXiv preprint, 2024
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55cf8d2c-e378-446b-aafe-a18898ff0784 · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? World Action Models are Zero-shot Policies
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3e5fb449-95e6-44bb-a884-981bea3df0cc · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Causal World Modeling for Robot Control
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1d603c82-7b96-4ff3-9bc8-944b3ab84c57 · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Dit4dit: Jointly modeling video dynamics and actions for generalizable robot control
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation da15809f-ac50-4c19-88ab-a4941754f2d2 · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8806fd25-8b79-4fb0-8f2d-fbc5721e0fef · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Bagelvla: Enhancing long-horizon manipulation via interleaved vision-language-action generation.CoRR, abs/2602.09849
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2fc7ab0e-a2cc-4f98-8656-9c603164ae31 · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? UAM: A Dual-Stream Perspective on Forgetting in VLA Training
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8745c12b-469d-43bf-befc-56524f10daa2 · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? AIM: Intent-Aware Unified world action Modeling with Spatial Value Maps
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3cbd8d2b-fb12-46a5-a3b1-5336238b6068 · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Unified world models: Coupling video and action diffusion for pretraining on large robotic datasets.arXiv preprint, 2025
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66749db6-7b82-4394-8ced-461f35ebfd52 · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? LDA-1B: Scaling Latent Dynamics Action Model via Universal Embodied Data Ingestion
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c2350027-27ce-45d1-a687-15dbb9000079 · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Disentangled Robot Learning via Separate Forward and Inverse Dynamics Pretraining
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 57a6f44a-2bea-45f8-bdf3-5c200963474c · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Motus: A Unified Latent Action World Model
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 25347798-464b-4aa2-8c5a-9e67c39771b9 · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Fast-WAM: Do World Action Models Need Test-time Future Imagination?
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 97fae678-9072-4b65-9a4f-26a728759995 · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? arXiv preprint arXiv:2603.17240 , year=
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation fc6a206d-370a-47f8-8eaf-08b9c22d2e2d · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5c779dd0-6daf-4f43-895b-811706055d88 · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Reworld: Multi-dimensional reward modeling for embodied world models.arXiv preprint arXiv:2601.12428
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4468830d-c137-4bd6-a2b7-526e744b2ea5 · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Orv: 4d occupancy-centric robot video generation
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8a7e792b-03f0-4fe0-ae1e-fc9c18abf86b · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? TesserAct: Learning 4D Embodied World Models
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 09a1f3a2-efba-441a-bf4d-3066a5bc7ecb · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Scene graph disentanglement and composition for generalizable complex image generation.Advances in Neural Information Processing Systems, 37:98478–98504, 2024
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68c14e92-5f93-4d71-89ab-50c2bed12ba4 · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Nano banana pro.https://deepmind.google/technologies/gemini/, 2025
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f02af53-4df3-4ecf-a7f9-c2c9f8353a79 · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? GPT-Image-1.5.https://openai.com/index/new-chatgpt-images-is-here/, 2026
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4dfacc50-3044-436e-b1a1-c696d17de981 · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? ImgEdit: A Unified Image Editing Dataset and Benchmark
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d3fd09fc-a1e0-4db5-9927-11331ab8be47 · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Qwen-Image Technical Report
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e5ce8f5c-1ada-4a42-af99-6ebfc034226a · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Glm-image.https://huggingface.co/zai-org/GLM-Image, 2026
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91eb9690-007e-4ce1-8383-78c7d784bd25 · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? NextStep-1: Toward Autoregressive Image Generation with Continuous Tokens at Scale
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7b80f4f7-d35e-47e2-84bf-ab3887ecc7b9 · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Longcat-next: Lexicalizing modalities as discrete tokens
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c5aaa63f-f241-4dcf-89ff-08905dbda029 · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Uni-Edit: Intelligent Editing Is A General Task For Unified Model Tuning
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d7782d6b-c758-40a0-8a40-0e99d0d7f834 · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3b2a4355-4f7a-4393-b59d-445e17d75e08 · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Magicbrush: A manually annotated dataset for instruction-guided image editing
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a9bca96-0817-4ad8-8670-8541b39e6314 · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Guiding instruction-based image editing via multimodal large language models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85834bbe-24b8-4385-886f-b331ea5d66a2 · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Emu edit: Precise image editing via recognition and generation tasks
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60268c47-cab4-4bd4-832b-ed8b01614dc6 · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Anyedit: Mastering unified high-quality image editing for any idea
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation afc015af-0f20-46b7-b41d-27add10fb77a · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Image Generators are Generalist Vision Learners
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d4147ffa-7d69-4dc8-95e0-276839c1df09 · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Diffusion Model as a Generalist Segmentation Learner
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 526d87a1-7c32-4881-a54b-ddfdd7862d4b · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Leveraging Image Generators to Address Data Scarcity: The Gen4Regen Dataset for Forest Regeneration Mapping
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0509c2cc-1de9-4542-875d-ce00f9e75795 · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? pi0: A vision-language-action flow model for general robot control
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 375cd041-792b-41f4-8889-e89e8caae296 · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? pi0.5: a vision-language-action model with open-world generalization.arXiv preprint, 2025
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 105870cd-06ae-4c3c-9d6b-c3ca971de2ec · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Gr00t n1: An open foundation model for generalist humanoid robots
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc0dc4f3-9207-42c1-bbb4-4f0c6824d81f · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Dreamvla: A vision-language-action model dreamed with comprehensive world knowledge.arXiv preprint, 2025
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc2f5178-863e-4215-b1f6-ba371fc13ec7 · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Reconvla: Reconstructive vision-language-action model as effective robot perceiver
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c953e55b-3a7a-43a8-aae7-6005ab8d20f8 · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? HY-Embodied-0.5: Embodied Foundation Models for Real-World Agents
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation db88cc04-c30a-4c28-864c-8cfed2513b62 · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4928c709-4293-4dd2-a8c9-ae42f4f9e16b · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Depthvla: Enhancing vision-language-action models with depth-aware spatial reasoning
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9e20d433-8483-424e-8815-44a13802bb41 · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Spatialvla: Exploring spatial representations for visual-language-action model.arXiv preprint, 2025
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7488449-f3a2-412a-8bc5-3ffce10961a3 · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Predictive inverse dynamics models are scalable learners for robotic manipulation.ICLR, 2024
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c0567e8-f909-4605-aefe-da49d1e3874a · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Cot-vla: Visual chain-of-thought reasoning for vision-language-action models.arXiv preprint, 2025
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d3d711b-191d-4090-a626-c0bb205772f8 · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Transport Discrepancy as a Reliability Signal for Vision-Language-Action Models
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3f43bd6c-0aa9-4096-9776-af5ce84dcc5b · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Being-h0: Vision-language-action pretraining from large-scale human videos
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61d533a4-7712-4317-b84c-043e57ea2e73 · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Unified diffusion vla: Vision-language-action model via joint discrete denoising diffusion process
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ab6841ed-a7aa-46b8-889c-cd84731cb13b · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Spatial forcing: Implicit spatial representation alignment for vision- language-action model
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7304a538-edaa-4395-b132-608c61fc49b8 · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Vla-jepa: Enhancing vision-language-action model with latent world model
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d5cdf597-1f76-4d68-b72a-c0f1332bb108 · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Vla-adapter: An effective paradigm for tiny-scale vision-language-action model
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2bb38f6-3ddc-47fa-8f33-a49eaf5c2754 · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? A Pragmatic VLA Foundation Model
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 69cf02e9-7140-473b-b6ba-a15b7c624c4a · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? MolmoAct: Action Reasoning Models that can Reason in Space
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 29857894-4f8e-42cd-9a14-732efa9d282d · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? F1: A Vision-Language-Action Model Bridging Understanding and Generation to Actions
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f1f3d92f-c05d-4754-b3e8-4e56ea20eac5 · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 11fa9e64-fa22-4e49-a0f3-d2cc74c18463 · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Seeing to act, prompting to specify: A bayesian factorization of vision language action policy
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6056000c-c1ab-41ad-b785-3109ca2cd912 · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Learning universal policies via text-guided video generation.NeurIPS, 2024
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 784b60f0-9eeb-43f2-8e59-4863e28530af · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Zero-shot robotic manipulation with pretrained image-editing diffusion models.arXiv preprint, 2023
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78d94391-2ea9-404c-a55d-9f6700deb43d · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Generalist bimanual manipulation via foundation video diffusion models.arXiv preprint, 2025
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b5e2f14-2e13-463d-9e72-a8bf5c2c14b8 · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Vidman: Exploiting implicit dynamics from video diffusion model for effective robot manipulation.NeurIPS, 2024
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74fcc968-2799-4902-913c-45a3397f77dd · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Murphy, Chelsea Finn, and Yilun Du
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a78fae03-fe60-4ad9-a60c-293f8a4c3896 · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Large Video Planner Enables Generalizable Robot Control
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 00920460-fc74-4407-8332-ed14b0520ee3 · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Anypos: Automated task-agnostic actions for bimanual manipulation.arXiv preprint, 2025
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de19a9f3-fd6e-47e8-9c07-45022c798637 · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Tc-idm: Grounding video generation for executable zero-shot robot motion
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e58da869-edcd-492c-95ea-0f9fb0fce6f7 · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Veo-act: How far can frontier video models advance generalizable robot manipulation? 2026
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d98d9d2e-6f0c-4044-900b-1bf18c42b8cd · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? V AMPO: Policy optimization for improving visual dynamics in video action models
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation fd943eb4-7439-4c96-bd18-58f8cf3a2319 · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Do World Action Models Generalize Better than VLAs? A Robustness Study
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3448e806-290e-4388-83ae-79aea7606673 · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? WorldEval: World Model as Real-World Robot Policies Evaluator
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0aa67ed4-4f61-4427-9ba1-0dd92c6fa43a · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Kinema4d: Kinematic 4d world modeling for spatiotemporal embodied simulation, 2026
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 935e3104-67d5-4bb1-ae65-53764a4d2599 · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? WoVR: World Models as Reliable Simulators for Post-Training VLA Policies with RL
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation dfbbef40-6a09-4aa3-89ed-07011a2a21aa · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? RoboStereo: Dual-Tower 4D Embodied World Models for Unified Policy Optimization
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4ff1ec73-74e3-4a7d-b10c-126d3e147026 · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? UniT: Toward a Unified Physical Language for Human-to-Humanoid Policy Learning and World Modeling
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 905869cc-0d2d-42ac-b5a1-f95b68dda1dc · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Persistent robot world models: Stabilizing multi- step rollouts via reinforcement learning.arXiv preprint arXiv:2603.25685, 2026
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation de531357-d7d4-4ba2-8563-86cfc7f0f692 · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Fate: Closed-loop feasibility-aware task generation with active repair for physically grounded robotic curricula.arXiv preprint arXiv:2603.01505, 2026
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0f438a2f-b26c-4524-8e23-1d02b2425c4a · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? VAG: Dual-Stream Video-Action Generation for Embodied Data Synthesis
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c7a70452-d578-409e-a32a-a9c6d08002a6 · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Interactive world simulator for robot policy training and evaluation
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 57c7b703-25db-4956-9d2c-87d108d8d84e · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? World Action Verifier: Self-Improving World Models via Forward-Inverse Asymmetry
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e6c788ca-136a-4193-a8b3-8ecb38839b1b · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? World-Value-Action Model: Implicit Planning for Vision-Language-Action Systems
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation bbbb3e9b-cda3-43ae-a63d-5bd3d44b3f2b · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2251de90-0eb7-4782-8ab7-a819a181c546 · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? dworldeval: Scalable robotic policy evaluation via discrete diffusion world model
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68440a13-9d8c-410b-9de8-5c82605775cf · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Interactive world simulator for robot policy training and evaluation
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e537418e-723c-4118-8f8f-2c213335ddbe · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Unresolved cited work
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 891ebd45-3d91-4e59-b077-7546724672b3 · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Cosmos 3: Omnimodal World Models for Physical AI
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation db9e670c-aa8c-4e7f-9acc-74fd93acba5c · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 86578b31-29d0-469b-b3f0-e0c6be404263 · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Ovis-U1 Technical Report
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6daf48e1-c49d-4601-b795-4e10bd1bd6ae · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? FLUX.2: Frontier Visual Intelligence.https://bfl.ai/blog/flux-2, 2025
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f66e74c9-5128-42bb-a262-ac921091081c · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Libero: Benchmarking knowledge transfer for lifelong robot learning.arXiv preprint, 2023
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fcbcddc8-c2a0-40b7-84c6-e42809dcef7b · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation bb81c9a5-2dbb-4999-bd92-34e297a3d90a · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7bd0c31e-2336-4fc0-a92b-09e08f8e7897 · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? ABot-M0: VLA Foundation Model for Robotic Manipulation with Action Manifold Learning
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 358e5826-1b6d-4992-a9c2-c95f290d10b8 · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Openvla: An open-source vision-language-action model.arXiv preprint, 2024
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f6d3ece-d196-458e-bb2b-9f30247be5bd · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? LIBERO: benchmarking knowledge transfer for lifelong robot learning
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9661e013-9c30-4787-b8ba-c16769e88032 · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Univla: Learning to act anywhere with task-centric latent actions.arXiv preprint, 2025
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6df536dd-5ba4-4a2e-860b-8f92479e53a7 · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Fine-tuning vision-language-action models: Optimizing speed and success.arXiv preprint, 2025
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6a9cb63-31cb-4abd-b7ca-7bf46aea9909 · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Fast: Efficient action tokenization for vision-language-action models.arXiv preprint, 2025
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb9c0662-b8c2-4117-9d68-c1dc0dfaf584 · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Worldvla: Towards autoregressive action world model.arXiv preprint, 2025
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c25c564c-8273-4211-9a7e-a372c55ad769 · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Unified Vision-Language-Action Model
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e1f2d5fa-4b5a-461b-9c8a-c65358a4e9a2 · outbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Representationalignmentforgeneration: Trainingdiffusiontransformersiseasierthanyouthink
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db64921a-7e5f-4e69-8e4a-b4f9d1a4cb15 · inbound
ABot-M0.5: Unified Mobility-and-Manipulation World Action Model ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing?
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 65c4f8c6-d1ef-47d5-a90f-41e7a1289e64 · inbound
ABot-M0.5: Unified Mobility-and-Manipulation World Action Model ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing?
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71d2ccd4-84eb-47f8-b566-dfded78279e5 · inbound
WAM-TTT: Steering World-Action Models by Watching Human Play at Test Time ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing?
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ec7efade-5d5f-4b4f-ae57-6a1d700fdc21 · inbound
WAM-TTT: Steering World-Action Models by Watching Human Play at Test Time ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing?
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95e3db30-fa90-42e0-8294-3544737834fb · inbound
LeapBot-WA: World-Anchor Action Models via Predictive Latent Alignments ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing?
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 671ec938-b253-4260-961e-f997e3e5d774 · inbound
Enfold: Folding World Model Imagination into Predictive Representations for Ultra-Efficient Embodied Control ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing?
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95b61b4f-5aed-4deb-8e15-66783e7c5618 · inbound
Enfold: Folding World Model Imagination into Predictive Representations for Ultra-Efficient Embodied Control ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing?
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c35ee440-b94a-4d52-af34-7428fc87f338 · inbound
Disentangling Visuo-Tactile Foresight: Oracle-Guided Interface Discovery for World Action Models ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing?
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.