Pith. sign in

Paper Citation Record · LEDGER

World Simulation with Video Foundation Models for Physical AI

As of 5 August 2026, this Paper Citation Record lists 100 of 103 outbound references and 100 inbound Pith citation observations for arXiv:2511.00062.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2511.00062 v2

Coverage vector

measured 100 of 103 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-12T23:01:13.546110Z

measured 200 of 200 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 100 of 134 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T18:05:46.912489Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

100 of 103 outbound references displayed

  • verified exact62
  • verified fuzzy34
  • unresolved1
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch2

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation ab3cb743-68a0-4ace-b948-d3a1d9d89180 · outbound

This paper cites V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning.

World Simulation with Video Foundation Models for Physical AI V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-12T23:01:13.764797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:c6f4f55e9763d982ee21b6a2c4204bbd448131db71783d4db06ae11a8de6f3c1

Observation 412f4703-524d-4c24-8067-82cce219e50b · outbound

This paper cites Edify Image: High-Quality Image Generation with Pixel Space Laplacian Diffusion Models.

World Simulation with Video Foundation Models for Physical AI Edify Image: High-Quality Image Generation with Pixel Space Laplacian Diffusion Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:13.610147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:8143cf5ab17332e71093ccaf079f66aa9247fbedcabd7e525dfeb9a9edc229e6

Observation 8d4e97d9-0307-45a8-9a93-615803083ce8 · outbound

This paper cites Recammaster: Camera-controlled generative rendering from a single video.

World Simulation with Video Foundation Models for Physical AI Recammaster: Camera-controlled generative rendering from a single video

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:01:14.114818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:0208206a4df5871b055c7113b6b98a2036a3713ea58ccb4ae8c2b8064e3baf3e

Observation 05cfa97a-37d4-47f8-a963-f576c59a03bb · outbound

This paper cites Syncammaster: Synchronizing multi-camera video generation from diverse viewpoints.

World Simulation with Video Foundation Models for Physical AI Syncammaster: Synchronizing multi-camera video generation from diverse viewpoints

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:01:14.080265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:1e616cdcd52dae6499768aa23c1cb6f72b45446c7bfba4889404e3989502106c

Observation 8ad112cd-35b9-4b7d-9eab-055948c21a6f · outbound

This paper cites Qwen2.5-VL Technical Report.

World Simulation with Video Foundation Models for Physical AI Qwen2.5-VL Technical Report

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-12T23:01:13.900179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:c24dce0b1b8c5da035035cae9351be36b944af28e0acd20596618cae599907ad

Observation 8c71e354-e252-43f9-8c9c-d59d67da6b9d · outbound

This paper cites Genie 3: A new frontier for world models.

World Simulation with Video Foundation Models for Physical AI Genie 3: A new frontier for world models

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:01:14.090308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:cce027b992578540b9e7f42c081667eab721daecce988796cef616d06fe529ad

Observation c17a0f9c-3cfe-4b8b-a98d-f1dec6f899ab · outbound

This paper cites VideoPhy: Evaluating Physical Commonsense for Video Generation.

World Simulation with Video Foundation Models for Physical AI VideoPhy: Evaluating Physical Commonsense for Video Generation

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:34:37.994076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:afd2fc0b3f57869d22a11bc0260c23531e012f22f05f3310fb1a7fbc5cb885db

Observation feea1a15-9db2-4ab4-8253-1d0e571742b9 · outbound

This paper cites VideoPhy-2: A Challenging Action-Centric Physical Commonsense Evaluation in Video Generation.

World Simulation with Video Foundation Models for Physical AI VideoPhy-2: A Challenging Action-Centric Physical Commonsense Evaluation in Video Generation

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:13.918325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:8a4e103c8e13a13ab686cd3f5af3bac1f4349cbb1dc62febe132bb1878b41aa5

Observation 8b76e304-375e-4841-a6bc-bf9387a2be39 · outbound

This paper cites GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.

World Simulation with Video Foundation Models for Physical AI GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-12T23:01:13.924336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:d75aa93ad26a4b34ef490cdd2cb739a202a68ccd5cf59bd556d7992dfc1df88f

Observation 87aa9c59-6556-4152-a084-af0d4e6de054 · outbound

This paper cites an unresolved cited work.

World Simulation with Video Foundation Models for Physical AI Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-05-12T23:01:14.108994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:ac336b8de23d38cdc1c408627feba1cc391e0534534aea53a373a99a78e7a061

Observation d5e3258f-1975-43a1-a456-a2d2404812eb · outbound

This paper cites IntPhys 2: Benchmarking Intuitive Physics Understanding In Complex Synthetic Environments.

World Simulation with Video Foundation Models for Physical AI IntPhys 2: Benchmarking Intuitive Physics Understanding In Complex Synthetic Environments

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:13.930429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:7191521537351a8d5e50212dac3f46693980e5c5484d5ce18d1f1bffa8ebd5bc

Observation be046de4-45a5-4780-8d99-c997383fe7c2 · outbound

This paper cites Agibot world colosseo: A large-scale manipulation platform for scalable and intelligent embodied systems.

World Simulation with Video Foundation Models for Physical AI Agibot world colosseo: A large-scale manipulation platform for scalable and intelligent embodied systems

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:01:14.120244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:761d38cc19be871ec4d596d6ceabbe7e9b044e3c9f4369ea8f9041655ff74249

Observation 78979042-1691-4b31-97a4-7a2df6df84a2 · outbound

This paper cites Planning with Reasoning using Vision Language World Model.

World Simulation with Video Foundation Models for Physical AI Planning with Reasoning using Vision Language World Model

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:13.939379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:4d085f7f73f737a94e6437da333a0d916b1f0fcfa38780b6bf65d68db4ba3ff5

Observation b4249491-da51-4d0b-938f-825a0180ee23 · outbound

This paper cites Video depth anything: Consistent depth estimation for super-long videos.

World Simulation with Video Foundation Models for Physical AI Video depth anything: Consistent depth estimation for super-long videos

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:01:14.130867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:61f96ddb9191802b795c69d74a34c77d6ba8a0339a4202ba81be90cbbbbc46fb

Observation b2e3c0e0-06c0-4587-9278-38e3d2fc6054 · outbound

This paper cites On the Importance of Noise Scheduling for Diffusion Models.

World Simulation with Video Foundation Models for Physical AI On the Importance of Noise Scheduling for Diffusion Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:13.947275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:4915d9307589447d8de45419870b313bf4fd23902cb8b2c9f4eae8c3e0e3efc9

Observation 55b562ef-1d13-4b22-996a-c77360e10b06 · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action diffusion.

World Simulation with Video Foundation Models for Physical AI Diffusion policy: Visuomotor policy learning via action diffusion

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:01:14.152163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:35a7595da1e0f32f7c8c6c460c0c4b26cafe5e51586737b649fb0e4f07bec982

Observation 9458ccae-8aac-41b5-afea-ceabf2f6645d · outbound

This paper cites Delta lake: Open-source storage framework that enables building lakehouses.https: //delta.io/.

World Simulation with Video Foundation Models for Physical AI Delta lake: Open-source storage framework that enables building lakehouses.https: //delta.io/

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:01:14.158787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:79e260cf3dd6177d0a2d7153eb9bfb6c2b31becba679c4d6feaaa0a081b90462

Observation 49a52bbf-cba0-4db8-bdb9-078863bdba7e · outbound

This paper cites Veo 3, 5 2025.

World Simulation with Video Foundation Models for Physical AI Veo 3, 5 2025

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:01:14.163715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:1950bda43ddc83bdb543bfd4dd72af653e255e5ee349498a60ec193d9d7742bd

Observation c29ec697-2aeb-436e-ade5-3d8d1c0c930d · outbound

This paper cites Worldscore: A unified evaluation benchmark for world generation.

World Simulation with Video Foundation Models for Physical AI Worldscore: A unified evaluation benchmark for world generation

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:13.954077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:3c72fdd6a0c5ca329096b1ffada49ce3d4240288213e3745a07998a2985b9a0f

Observation 29ff6d43-a026-4531-8901-be2947ef9b2f · outbound

This paper cites Scaling rectified flow transformers for high-resolution image synthesis.

World Simulation with Video Foundation Models for Physical AI Scaling rectified flow transformers for high-resolution image synthesis

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:01:14.176994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:8b55d4bf6d82327a734c0c83d3b8c9169d1b8239bff1ac370a8a1fbc9a9837f0

Observation 3464cb8c-58e6-4fa1-906f-ac882462c28b · outbound

This paper cites LLM-based Realistic Safety-Critical Driving Video Generation.

World Simulation with Video Foundation Models for Physical AI LLM-based Realistic Safety-Critical Driving Video Generation

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-12T23:01:13.962088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:c0da1c61608f786c4f6152ff313bb33a9c9006698f3f37e7b715b382c4ed8de0

Observation 1f8223a3-7728-4dc5-845a-deb5eb77378a · outbound

This paper cites Diffusion models and gaussian flow matching: Two sides of the same coin.

World Simulation with Video Foundation Models for Physical AI Diffusion models and gaussian flow matching: Two sides of the same coin

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:01:14.195347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:7ee8b68249d97d364064a647652ae4f05e3f8f414d57dab9c3f528d8a2d840f4

Observation edea4d0f-8dc8-424d-9fcb-f19dc243d1b4 · outbound

This paper cites Seedance 1.0: Exploring the Boundaries of Video Generation Models.

World Simulation with Video Foundation Models for Physical AI Seedance 1.0: Exploring the Boundaries of Video Generation Models

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-12T23:01:13.971409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:65ea5a11180bf3b7f6bcb76a0279f96127fd6aea012c0513759d20bff7cca9d5

Observation eaa9cbc9-f6d9-4048-9310-6e68115ac020 · outbound

This paper cites YOLOX: Exceeding YOLO Series in 2021.

World Simulation with Video Foundation Models for Physical AI YOLOX: Exceeding YOLO Series in 2021

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-13T10:31:31.716715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:7ba531d9e596f20e27e53d8d9f6d14d7d8c867517649491fe7a66f059ec543d2

Observation 3dadca2a-d312-45cb-9b6f-62d0f3855602 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

World Simulation with Video Foundation Models for Physical AI DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-12T23:01:13.986635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:269771e53d440fb469137e0adf7a6382da423fe0e794500410843ea6cb82c58f

Observation 8dc1b711-ac24-4cf6-b365-8735ce972af6 · outbound

This paper cites T2VPhysBench: A First-Principles Benchmark for Physical Consistency in Text-to-Video Generation.

World Simulation with Video Foundation Models for Physical AI T2VPhysBench: A First-Principles Benchmark for Physical Consistency in Text-to-Video Generation

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:13.994170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:edcb482fbfbb486a3c3f5ccdb4f1c671862f605c6382801fa09509fd9f191d60

Observation f99882ba-cb42-46ea-9153-109dca9c5367 · outbound

This paper cites World Models.

World Simulation with Video Foundation Models for Physical AI World Models

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-12T23:01:14.002924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:45a51dbc90f9cf8a7a1e353407106bb7daf4e3698a22fa24f934a7239e5c4cf9

Observation 906f848f-38b0-44f0-99bf-c1808cfbddd3 · outbound

This paper cites LTX-Video: Realtime Video Latent Diffusion.

World Simulation with Video Foundation Models for Physical AI LTX-Video: Realtime Video Latent Diffusion

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-12T23:01:14.011752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:a23761d30130804d342630358a5374ca705879f21866eb638da22537b8f92e5d

Observation 4a1f96cd-9daa-47a3-b344-64bcabdb5e9f · outbound

This paper cites Dream to Control: Learning Behaviors by Latent Imagination.

World Simulation with Video Foundation Models for Physical AI Dream to Control: Learning Behaviors by Latent Imagination

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-12T23:01:14.018849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:cb41c9ae0847e8d4c38ef9f1f731162dec05b45babbd1882a790cb496edd0a06

Observation aa7ef977-8a45-4454-abc7-f06974e97305 · outbound

This paper cites Generalized Neighborhood Attention: Multi-dimensional Sparse Attention at the Speed of Light.

World Simulation with Video Foundation Models for Physical AI Generalized Neighborhood Attention: Multi-dimensional Sparse Attention at the Speed of Light

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:14.026420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:339f42355d597aac7e51740da5e74947410190c4684595f3667757fd5c1cc475

Observation 1aac1657-8923-4a1e-9f87-cc03d0c41c31 · outbound

This paper cites UniRelight: Learning Joint Decomposition and Synthesis for Video Relighting.

World Simulation with Video Foundation Models for Physical AI UniRelight: Learning Joint Decomposition and Synthesis for Video Relighting

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:14.035929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:5c17967eb922ae152a3881c467eb598bf758dc4c71a2ed062ffbee1a26ee425b

Observation cd329c03-ce76-45d0-9f9d-4782b071fd88 · outbound

This paper cites simple diffusion: End-to-end diffusion for high resolution images.

World Simulation with Video Foundation Models for Physical AI simple diffusion: End-to-end diffusion for high resolution images

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:01:14.264196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:b11ff684f3c97d58c8cefb4b5ea2bcaf988ba6ec32ec813d4d4a88ad550965b4

Observation f9dcd7e3-e2e4-4e43-8037-dba17d16ab2d · outbound

This paper cites ViPE: Video Pose Engine for 3D Geometric Perception.

World Simulation with Video Foundation Models for Physical AI ViPE: Video Pose Engine for 3D Geometric Perception

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-16T16:41:08.910540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:c083a27177be0fb7aa07deabf808c5e51bfca69aa75ef317b8da9b7f8b29c00f

Observation 3842889c-98aa-432d-93c6-397d01e3b08f · outbound

This paper cites LET-3D-AP: Longitudinal Error Tolerant 3D Average Precision for Camera-Only 3D Detection.

World Simulation with Video Foundation Models for Physical AI LET-3D-AP: Longitudinal Error Tolerant 3D Average Precision for Camera-Only 3D Detection

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:14.050176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:f719a55d1eb041386c6310c298ebf7ece6ebdcfc991f07190c10bd2cf73bc723

Observation d4e9e9b6-b193-4e97-8800-42732f720e3e · outbound

This paper cites GPT-4o System Card.

World Simulation with Video Foundation Models for Physical AI GPT-4o System Card

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-12T23:01:14.056202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:a6d7e66d5c737f6a89358442e1eabb61611290b44866db99e3a49f526e7cb9e7

Observation 284aa987-6557-4fc1-aa91-dcc35acd264d · outbound

This paper cites DreamGen: Unlocking Generalization in Robot Learning through Video World Models.

World Simulation with Video Foundation Models for Physical AI DreamGen: Unlocking Generalization in Robot Learning through Video World Models

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-15T23:50:45.800685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:2a07339edfde14136073d593c57d6d9b567612803ad6432db17fe86a13832e4d

Observation f0765cd3-e892-4bbe-9d26-fa149d80302a · outbound

This paper cites RTMPose: Real-Time Multi-Person Pose Estimation based on MMPose.

World Simulation with Video Foundation Models for Physical AI RTMPose: Real-Time Multi-Person Pose Estimation based on MMPose

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:14.070641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:0de70af36037902f5c7a53b71c28cb9cb0c6138a23e4b63280e2aa7dde60a67d

Observation 6dd5d3a1-74ef-486c-b209-98da15fff0a2 · outbound

This paper cites Elucidating the design space of diffusion-based generative models.NeurIPS.

World Simulation with Video Foundation Models for Physical AI Elucidating the design space of diffusion-based generative models.NeurIPS

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:01:14.304402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:b5d3283252427b5953b97fd72d61b5724906f10a1f7ce97e8da71bcb60081dda

Observation 3537471d-bbb2-4c45-8296-29eb79762abc · outbound

This paper cites DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset.

World Simulation with Video Foundation Models for Physical AI DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-12T23:01:14.075355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:3a2874c1a3a80e7513875ae7e7c0155a70415f3c7f54dfc8dffb8fbcb8cde65e

Observation 665fad89-9f5f-469c-a9e5-a035537641f5 · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

World Simulation with Video Foundation Models for Physical AI HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-12T23:01:13.624574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:b59e24ea91e15023c26786a05118a01e166a396ddd167debe09e36607aa66f74

Observation b29f4ddb-f6b4-498a-9da7-c2e97480c021 · outbound

This paper cites Kling.

World Simulation with Video Foundation Models for Physical AI Kling

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:01:14.321923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:b2c79eb662500572a3ab3c70d6b7ed8c8b9e35bb76cc97a14b327466089b4bf4

Observation f2ee8ce3-f025-44bc-a2cf-8780d356ba4e · outbound

This paper cites BEVFormer: Learning Bird's-Eye-View Representation from Multi-Camera Images via Spatiotemporal Transformers.

World Simulation with Video Foundation Models for Physical AI BEVFormer: Learning Bird's-Eye-View Representation from Multi-Camera Images via Spatiotemporal Transformers

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:13.632850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:2dbfd265429a86a8711bf8471c720609bc445b1b4fc873f138a63f57ed043191

Observation 5b6e1514-92d1-45c1-ada2-7903ddd6408a · outbound

This paper cites Won- derplay: Dynamic 3d scene generation from a single image and actions.arXiv preprint arXiv:2505.18151.

World Simulation with Video Foundation Models for Physical AI Won- derplay: Dynamic 3d scene generation from a single image and actions.arXiv preprint arXiv:2505.18151

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:13.639296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:824c1d977a90e14148d1b5dfa4fc5f8d9baf43b3b518a0dc1b6f7e9d2260b4d1

Observation cbcafa83-5f1e-4391-ab60-06f49959080b · outbound

This paper cites Torchtitan: One-stop pytorch native solution for production ready LLM pretraining.

World Simulation with Video Foundation Models for Physical AI Torchtitan: One-stop pytorch native solution for production ready LLM pretraining

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:01:14.099057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:74ec84cdeac95297fae81fb9d9b4c3e377b6f1917011ccbd496c6ac08ec59f0f

Observation 8dd4a2d0-814a-4cef-b4ad-421c3dd83fff · outbound

This paper cites Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation.

World Simulation with Video Foundation Models for Physical AI Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-15T21:28:42.077022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:f5382bdf61a840810dcf245d1e22e2886f3a5ea9538a7b63c1a988f1c8474147

Observation adbef28e-01d6-4b66-b831-e5918f9aadb2 · outbound

This paper cites Flow Matching for Generative Modeling.

World Simulation with Video Foundation Models for Physical AI Flow Matching for Generative Modeling

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-05-12T23:01:13.653660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:85a6724fbeda4655d416078a2e7b0d33895713e8108915a329b64d6e418767e6

Observation 44d2be4c-486d-448f-9aa2-ea991d47a487 · outbound

This paper cites Improving Video Generation with Human Feedback.

World Simulation with Video Foundation Models for Physical AI Improving Video Generation with Human Feedback

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-13T15:30:03.116566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:c178cb8ee6345c770bcc318a0d3ab088f52ac551cad5316a1cc36d5528f77402

Observation f99bbb63-c9bd-4b33-aa44-b2e411deb33e · outbound

This paper cites Dynamicscaler: Seamless and scalable video generation for panoramic scenes.

World Simulation with Video Foundation Models for Physical AI Dynamicscaler: Seamless and scalable video generation for panoramic scenes

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:01:14.138025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:e9694667b10e7b61c8f7b0e73af18707ea06b6cead5434665ea57c3c6cb4192e

Observation 08efc4e9-13c8-481a-aca0-8d5367ee7e4b · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

World Simulation with Video Foundation Models for Physical AI Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-05-12T23:01:13.667121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:a2c045932456324387baffbfdabd1e2ebfaddedbfa8ae201dd1496c4592de477

Observation fd492237-9055-48ba-a752-7c5e6a233445 · outbound

This paper cites Simplifying, Stabilizing and Scaling Continuous-Time Consistency Models.

World Simulation with Video Foundation Models for Physical AI Simplifying, Stabilizing and Scaling Continuous-Time Consistency Models

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-13T10:26:23.772771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:a320fa5f3ddce2042705a64646ae3f898d0d557f1cbc825eba2bb7653a6fd43b

Observation ef66dd02-13ac-41d7-9beb-570513e9222e · outbound

This paper cites LATR: 3D Lane Detection from Monocular Images with Transformer.

World Simulation with Video Foundation Models for Physical AI LATR: 3D Lane Detection from Monocular Images with Transformer

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:13.682645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:c9222c77ff029c6cba743c2579efa8e27a52c6d0f7746782b52e2134a63e0c83

Observation e9e99ea9-b2e0-4f61-a539-5affb253ebf9 · outbound

This paper cites Hailuo.

World Simulation with Video Foundation Models for Physical AI Hailuo

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:01:14.206489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:b95502309266e65649f553ece02741109ac22482f603d8576897be08f08a2c14

Observation b073b607-1fdc-4a9e-a313-abd9b69fd17e · outbound

This paper cites Do generative video models understand physical principles?.

World Simulation with Video Foundation Models for Physical AI Do generative video models understand physical principles?

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:47:06.031259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:72a721b9568c669ff1e97edb10871b5e6a59b90d9b9f1371473d710a0e3f8f97

Observation edf2c7b8-8352-4d53-ba76-1603399fd485 · outbound

This paper cites RoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist Robots.

World Simulation with Video Foundation Models for Physical AI RoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist Robots

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:46:30.425882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:bba1dd747202e9031bd873c9a7d5d212ffaf2a32acb8c081916e4a1fca030a1a

Observation 32f39998-bff7-4dce-8971-8a381add6192 · outbound

This paper cites Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning.

World Simulation with Video Foundation Models for Physical AI Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:47:10.348602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:bb56e6467a1fac7778cfb5a9f74940ede06527a4401599b9833ee56b89d02ccf

Observation f344ecf1-a33d-49c0-8634-5443f93647f8 · outbound

This paper cites Cosmos-Transfer1: Conditional World Generation with Adaptive Multimodal Control.

World Simulation with Video Foundation Models for Physical AI Cosmos-Transfer1: Conditional World Generation with Adaptive Multimodal Control

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:13.712342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:1b17462a2e7feaec2cb18f0f01b59bf57cd2adc28cb167710c71565a66fc486e

Observation 6f8d0196-3608-45c3-a9f9-cdf0d070daea · outbound

This paper cites Cosmos World Foundation Model Platform for Physical AI.

World Simulation with Video Foundation Models for Physical AI Cosmos World Foundation Model Platform for Physical AI

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-05-12T23:01:13.716640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:fac4760bb607814dbe25ff5e78664707d4f6cd012742bff247ad3ca758bce711

Observation ba398b82-ef48-45ec-8759-07d233eea72d · outbound

This paper cites an unresolved cited work.

World Simulation with Video Foundation Models for Physical AI Unresolved cited work

Reference 58

Resolution
parse uncertain
raw_fallback, observed 2026-05-12T23:01:14.248509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:52e5c156daec96d4e06bd9d5fa3a2a3abfa08a159c522e83482b5a171ecdc4c0

Observation 89c3a877-9d7f-402b-9b79-fefb1ab37802 · outbound

This paper cites Sora.

World Simulation with Video Foundation Models for Physical AI Sora

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:01:14.256392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:1b9077f9f892e8b9813b5f49d36a65b3e52b5fd265f106d208165842ee6d86d9

Observation 2ee28bcd-8222-4143-82a5-fe69c5d9dabd · outbound

This paper cites Training language models to follow instructions with human feedback.NeurIPS.

World Simulation with Video Foundation Models for Physical AI Training language models to follow instructions with human feedback.NeurIPS

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:01:14.269581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:8f86df58b5a29c131a625aaaf4daa1e489f7b1f2f86909b0e78b5902d482e90b

Observation 19018ce4-6981-4d95-9720-68469fb6f6d1 · outbound

This paper cites YaRN: Efficient Context Window Extension of Large Language Models.

World Simulation with Video Foundation Models for Physical AI YaRN: Efficient Context Window Extension of Large Language Models

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-05-12T23:01:13.721168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:998fc31a6dbdf0940e304523137bae90480440bfe510fcc20cd70a72762efd00

Observation 74b9026e-75e5-420d-80d1-c74b354abed5 · outbound

This paper cites Movie Gen: A Cast of Media Foundation Models.

World Simulation with Video Foundation Models for Physical AI Movie Gen: A Cast of Media Foundation Models

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-05-12T23:01:13.725884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:710b7d7cd5c11769a878750af498804b3ad96adaa1299d0d911be41155bcb8b3

Observation a2727a34-99a7-459f-adef-a336243af3b0 · outbound

This paper cites Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters.

World Simulation with Video Foundation Models for Physical AI Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:01:14.287690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:0ee33f45875580d3aadab5a5e56668f8d6d8d38cdd0660bac25ed6d4865a26a2

Observation 66e125d8-e0fc-4ea1-8ee6-072a45efb597 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

World Simulation with Video Foundation Models for Physical AI SAM 2: Segment Anything in Images and Videos

Reference 64

Resolution
verified exact
local_arxiv, observed 2026-05-12T23:01:13.730843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:f63770f78e580130d5f43327ea5e1ee0530d2329cd7e79115d9d4634845f2e11

Observation 95d616d7-bd92-46a3-b35c-e77a2a868564 · outbound

This paper cites Ren, Justin Lidard, Lars Lien Ankile, Anthony Simeonov, Pulkit Agrawal, Anirudha Majumdar, Benjamin Burchfiel, Hongkai Dai, and Max Simchowitz.

World Simulation with Video Foundation Models for Physical AI Ren, Justin Lidard, Lars Lien Ankile, Anthony Simeonov, Pulkit Agrawal, Anirudha Majumdar, Benjamin Burchfiel, Hongkai Dai, and Max Simchowitz

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:01:14.310856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:67acdaea64104e7416fbe0293a7379fe60bf28cb5225efc5db8553668713d119

Observation 283d1d9f-ef01-48eb-b6ed-3b806058b365 · outbound

This paper cites Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks.

World Simulation with Video Foundation Models for Physical AI Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-05-12T23:01:13.736123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:cb5178c1593f19085b3f6f789c019dcd86336cb52cd26fb4a6d203d6eebb5f5d

Observation 32c4917d-c431-423f-97c5-4237fd9f6cbc · outbound

This paper cites Cosmos-Drive-Dreams: Scalable Synthetic Driving Data Generation with World Foundation Models.

World Simulation with Video Foundation Models for Physical AI Cosmos-Drive-Dreams: Scalable Synthetic Driving Data Generation with World Foundation Models

Reference 67

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T23:01:13.741361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:a54b8b59aeb83cefd1043c2c9cb2f36b1f30333d3e43954813e9aa8d3bd19cde

Observation 654844e1-d016-48e3-a839-c110ecb5fcc1 · outbound

This paper cites Gen3c: 3d-informed world-consistent video generation with precise camera control.

World Simulation with Video Foundation Models for Physical AI Gen3c: 3d-informed world-consistent video generation with precise camera control

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:01:14.094409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:c37739b63665a2a3e68a50df4176964fcdcfb61cb4c76dd5a8fbe4b88bac1704

Observation 26ee39e3-336d-4d1e-b6ec-aeba3a8aa358 · outbound

This paper cites Gen 3.

World Simulation with Video Foundation Models for Physical AI Gen 3

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:01:14.104453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:74a06682bc60b96b0c9d055d5642c09d021060de53507459682719532b1ec1b8

Observation 769806dc-f5fe-4f9d-be52-832c99f440da · outbound

This paper cites very scattered.

World Simulation with Video Foundation Models for Physical AI very scattered

Reference 70

Resolution
verified exact
doi, observed 2026-05-12T23:01:13.602107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:136fa1a5b1f18a1abe883b5483acfa0814c25780202187ef585eb5b9ae13b035

Observation 1a43ad0e-d7a0-4a16-ba2f-6fa68d3d1a41 · outbound

This paper cites Proximal Policy Optimization Algorithms.

World Simulation with Video Foundation Models for Physical AI Proximal Policy Optimization Algorithms

Reference 71

Resolution
verified exact
local_arxiv, observed 2026-05-12T23:01:13.746628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:4eda8b585e472683465b0431dfd03a3d8d1bb848ccb98e921f42bbc6582d4e1c

Observation 43e802b8-64de-4ea1-b362-8ab0529edb7b · outbound

This paper cites Text-To-4D Dynamic Scene Generation.

World Simulation with Video Foundation Models for Physical AI Text-To-4D Dynamic Scene Generation

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:13.753011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:d4022cf35a24d2a5a9e96377459d563f9f3a447a34406dd6fc31c3f40e5d2dee

Observation fc0617d6-52bd-4021-8f69-a2ccb5acd549 · outbound

This paper cites Light field networks: Neural scene representations with single-evaluation rendering.NeurIPS.

World Simulation with Video Foundation Models for Physical AI Light field networks: Neural scene representations with single-evaluation rendering.NeurIPS

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:01:14.187667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:bc668c7bf8ea186a4a4345dd2324d90400863361151d2f9d954813fb65b9a36a

Observation 6ea03bf2-a2d8-492f-b240-f06edf88a77c · outbound

This paper cites Stylegan-v: A continuous video generator with the price, image quality and perks of stylegan2.

World Simulation with Video Foundation Models for Physical AI Stylegan-v: A continuous video generator with the price, image quality and perks of stylegan2

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:01:14.201259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:73e14d759b38a405086d9e482fa8d3e0a8d71f8334b60dc6d0af2da6fddab747

Observation 488ccc1b-86f2-49fb-82ce-e6bfcb82f4a9 · outbound

This paper cites cuRobo: Parallelized Collision-Free Minimum-Jerk Robot Motion Generation.

World Simulation with Video Foundation Models for Physical AI cuRobo: Parallelized Collision-Free Minimum-Jerk Robot Motion Generation

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:13.759032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:68c1c0d79512f0f8427b3effd55ae71cbe261b149c06743ff3f6ea3c66a8cf8c

Observation 8833b5d3-c4f8-445d-920f-8debd9cf26b0 · outbound

This paper cites 1x technologies | safe humanoids for the home.

World Simulation with Video Foundation Models for Physical AI 1x technologies | safe humanoids for the home

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:01:14.223399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:b2b16f535f8e84dd08bc1797e2c434c31885c34c2b357a80d3fb463b3ab2003a

Observation 2f7c672a-6f2c-446d-a1bd-d7f359c57fe8 · outbound

This paper cites Open x-embodiment: Robotic learning datasets and rt-x models.

World Simulation with Video Foundation Models for Physical AI Open x-embodiment: Robotic learning datasets and rt-x models

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:01:14.228413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:c82bcbf6b516d367652180ba60fd7afdabbb4cf4059ef948f628ae2fbfb331c6

Observation 944ad03f-8f43-4bcd-b564-08fefcf4f9cb · outbound

This paper cites Bridgedata v2: A dataset for robot learning at scale.

World Simulation with Video Foundation Models for Physical AI Bridgedata v2: A dataset for robot learning at scale

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:01:14.233658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:17a4a599cf28e6f4981fe44a905b65a3582adcdc777db9d0c51378441ed763b3

Observation 49be7c1a-e43e-48fb-b04d-dd4d4cf4ebe1 · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

World Simulation with Video Foundation Models for Physical AI Wan: Open and Advanced Large-Scale Video Generative Models

Reference 79

Resolution
verified exact
local_arxiv, observed 2026-05-12T23:01:13.617255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:1d27e8ca56aa8c8aa8a73a05d5084f77d7629ebb987659d04476bec14ec558ed

Observation 99eae887-fd5f-4c07-9951-138dbeff236e · outbound

This paper cites A comprehensive study of decoder-only llms for text-to-image generation.

World Simulation with Video Foundation Models for Physical AI A comprehensive study of decoder-only llms for text-to-image generation

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:01:14.274833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:654dc57c4ca634b68e90d34fdaf6e319a8cc409da67d1f93465e7b9557e93d46

Observation 44c15761-9b25-49a0-aa9c-1407d55be937 · outbound

This paper cites Frame in-n-out: Unbounded con- trollable image-to-video generation.

World Simulation with Video Foundation Models for Physical AI Frame in-n-out: Unbounded con- trollable image-to-video generation

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:13.770365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:4e489f1d04e73aa0831a1a1126e26ea5b277ccb4bfac6e612546ad6e2f4566ab

Observation fb2dc8e2-b8b4-43a5-a17b-f3ce4ba162fb · outbound

This paper cites Koala-36m: A large-scale video dataset improving consistency between fine-grained conditions and video content.

World Simulation with Video Foundation Models for Physical AI Koala-36m: A large-scale video dataset improving consistency between fine-grained conditions and video content

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:01:14.296587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:7946526fe6eae0aa42ee262169acca08b358aeed8d14265d6a634eb22595aede

Observation 866afd17-3d65-4d32-95ae-5ddf1427481f · outbound

This paper cites Internvideo2: Scaling foundation models for multimodal video understanding.

World Simulation with Video Foundation Models for Physical AI Internvideo2: Scaling foundation models for multimodal video understanding

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:01:14.315657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:cda6a2876c9ddad87de460c3b7b3ec6f2ec4de472633b083b7c096c93bfb45d0

Observation f53d5fe8-576d-4941-9640-12dca55e4056 · outbound

This paper cites Controlling Space and Time with Diffusion Models.

World Simulation with Video Foundation Models for Physical AI Controlling Space and Time with Diffusion Models

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:13.776381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:0ea4f095477edf278ced6395113a95bb809cfad8fa0e84911d74d4debe3be59f

Observation ed7ba837-e571-4f04-b041-a9f4e77c7902 · outbound

This paper cites Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time.

World Simulation with Video Foundation Models for Physical AI Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:01:14.125708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:0ad3869fbffd8a0b7ced53d2737374b3f4c4b970a499bb979ef0f1a85dc6387a

Observation 3bdeaf82-ebd2-456b-8442-0534bb253c71 · outbound

This paper cites Exploring video quality assessment on user generated contents from aesthetic and technical perspectives.

World Simulation with Video Foundation Models for Physical AI Exploring video quality assessment on user generated contents from aesthetic and technical perspectives

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:01:14.169733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:df9e7987629196de76fdfe7e670648eb81cc5dd42afd5e5de1dda64132a14df4

Observation 68418d26-aa82-45e3-a780-4df51bc9b894 · outbound

This paper cites RoboMIND: Benchmark on Multi-embodiment Intelligence Normative Data for Robot Manipulation.

World Simulation with Video Foundation Models for Physical AI RoboMIND: Benchmark on Multi-embodiment Intelligence Normative Data for Robot Manipulation

Reference 87

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:14:18.432154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:fccfb50d16f003f644cbefb42703eb842745c2ab2a9f4b26a897ab9666c09f92

Observation 1c2f1d5a-30c3-4efa-972e-32f679470874 · outbound

This paper cites Ties-merging: Resolving interference when merging models.NeurIPS.

World Simulation with Video Foundation Models for Physical AI Ties-merging: Resolving interference when merging models.NeurIPS

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:01:14.240424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:f5ba96ea987abb8b41961ee76a4203d58c701337d88e5385906224b385cccaaf

Observation 51988b6b-b002-4279-ba30-465ed87158a8 · outbound

This paper cites Model Merging in LLMs, MLLMs, and Beyond: Methods, Theories, Applications and Opportunities.

World Simulation with Video Foundation Models for Physical AI Model Merging in LLMs, MLLMs, and Beyond: Methods, Theories, Applications and Opportunities

Reference 89

Resolution
verified exact
arxiv_id, observed 2026-05-17T22:16:05.189340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:7f9be82ce26331bedc4619a6d9e8710d35600a425a39b07f7091fbca67ebeae8

Observation 0c3c7d1d-2359-447e-acd6-b3946edd022e · outbound

This paper cites EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents.

World Simulation with Video Foundation Models for Physical AI EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

Reference 90

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:24:45.426358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:a56f165ec723cac5c18f4b28bbfd118d08d4afc2ecc07b328422073e9dc02479

Observation 74b42442-84f4-46f9-b307-b658f57c8fe1 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

World Simulation with Video Foundation Models for Physical AI CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 91

Resolution
verified exact
local_arxiv, observed 2026-05-12T23:01:13.799820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:acedbdf535ff5c6dc35ea8fb257642da88d709cb555fcaeb6704f3395eabdec1

Observation f5188f6a-8cf5-40e6-b1b3-820765e7236c · outbound

This paper cites Data-regularized reinforcement learning for diffusion models at scale.

World Simulation with Video Foundation Models for Physical AI Data-regularized reinforcement learning for diffusion models at scale

Reference 92

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:13.811776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:92b4aa943a1ecc2380277647f27024e17ceebe2dc539c56ff9abc473061d5d22

Observation 252b42ad-53b3-4bc5-a695-a07a8dea90ff · outbound

This paper cites Language models are super mario: Absorbing abilities from homologous models as a free lunch.

World Simulation with Video Foundation Models for Physical AI Language models are super mario: Absorbing abilities from homologous models as a free lunch

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:01:14.085751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:71e96dc779d1feda17f694a8d1cdaee3076885834f059dcdb21eadf6e3d62cb3

Observation 8af2157c-8527-42e4-97d2-2543d4aba109 · outbound

This paper cites EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models.

World Simulation with Video Foundation Models for Physical AI EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:13.822345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:a3c229f8d26b430b321f1a049a55aaaa176651461e760f8db8efbbf9d12502d7

Observation 2ec731de-1750-4793-9558-1281b08fe6fe · outbound

This paper cites Waver: Wave Your Way to Lifelike Video Generation.

World Simulation with Video Foundation Models for Physical AI Waver: Wave Your Way to Lifelike Video Generation

Reference 95

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:13.834624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:ee2e89ae33343c72778d39fe2428cac4a111663ec487681a6d858198b404178b

Observation e6b72cb2-cb9b-4d0f-ae62-3342d8ab69a7 · outbound

This paper cites GenXD: Generating Any 3D and 4D Scenes.

World Simulation with Video Foundation Models for Physical AI GenXD: Generating Any 3D and 4D Scenes

Reference 96

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:13.848208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:88b4692e87b43cae02a834b8a96974f3a18ff1c250935b26dfedba477e03a8d2

Observation 49ce5cb8-f139-4fb8-9c88-1d606b8f9a2b · outbound

This paper cites Are Synthetic Videos Useful? A Benchmark for Retrieval-Centric Evaluation of Synthetic Videos.

World Simulation with Video Foundation Models for Physical AI Are Synthetic Videos Useful? A Benchmark for Retrieval-Centric Evaluation of Synthetic Videos

Reference 97

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:13.858638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:bdd0a34634bda548c8e68db6a6773aa044d82f29f2893ab8c2509d01d52c7f4c

Observation 333d6195-f51a-44ff-bbea-23e061f4f542 · outbound

This paper cites Large Scale Diffusion Distillation via Score-Regularized Continuous-Time Consistency.

World Simulation with Video Foundation Models for Physical AI Large Scale Diffusion Distillation via Score-Regularized Continuous-Time Consistency

Reference 98

Resolution
verified exact
local_arxiv, observed 2026-05-12T23:01:13.868632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:22cf5ec28955e851ef3c574b6d790bd5fb6108a5b9e4c74bbd18196139e6d1e6

Observation ba33fbdf-c62f-4392-a184-c2149205d762 · outbound

This paper cites Unipc: A unified predictor-corrector framework for fast sampling of diffusion models.Advances in Neural Information Processing Systems, 36:49842–49869.

World Simulation with Video Foundation Models for Physical AI Unipc: A unified predictor-corrector framework for fast sampling of diffusion models.Advances in Neural Information Processing Systems, 36:49842–49869

Reference 99

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T23:01:13.875615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:c98fdbfb2fa1658ca5b087063d750c0b2c7772d864efb503697afcd63855267f

Observation b9d7bddd-8785-4d5c-bb5e-1031eda8cec5 · outbound

This paper cites VLM4D: Towards Spatiotemporal Awareness in Vision Language Models.

World Simulation with Video Foundation Models for Physical AI VLM4D: Towards Spatiotemporal Awareness in Vision Language Models

Reference 100

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:13.883037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:19ec2fce35ea5df9a67a7bf2846e9117159dd30e9980b4bd8b3551195175c100

Pith citing papers

Observation 3f536bf9-a60d-45b9-8dfa-2c26111dbfac · inbound

Non-invasive Assessment of Pancreatic Duct Hypertension Using Computational Flow Modeling cites this paper.

Non-invasive Assessment of Pancreatic Duct Hypertension Using Computational Flow Modeling World Simulation with Video Foundation Models for Physical AI

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T18:05:46.912489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:05:46.912489Z digest=sha256:03097167981eb1cc3c9dbebe38f977c3aa6eaae87915e716b1cefb02a6035ad4

Observation b3a173d5-3c1b-43f1-b9e8-e75981a48582 · inbound

SWITCH: Benchmarking Modeling and Handling of Tangible Interfaces in Long-horizon Embodied Scenarios cites this paper.

SWITCH: Benchmarking Modeling and Handling of Tangible Interfaces in Long-horizon Embodied Scenarios World Simulation with Video Foundation Models for Physical AI

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T21:14:08.241793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:14:08.241793Z digest=sha256:4c1b1fcb6787370707cbf0020f9ef130f5272d7e9de3e9ca08a7d38bd3e3c7d1

Observation 19866105-002a-4d37-a1c4-952a478a6906 · inbound

Action-guided generation of 3D functionality segmentation data cites this paper.

Action-guided generation of 3D functionality segmentation data World Simulation with Video Foundation Models for Physical AI

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-17T04:51:31.961334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T04:49:06.269256Z digest=sha256:5abae9d46f9df96d530de1c45875cbe1762f71ce6828cacb0084380029b74655

Observation ffb8ebf1-a0b8-4ab6-b101-e4f911ddd5bf · inbound

mimic-video: Video-Action Models for Generalizable Robot Control Beyond VLAs cites this paper.

mimic-video: Video-Action Models for Generalizable Robot Control Beyond VLAs World Simulation with Video Foundation Models for Physical AI

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-05-15T10:41:00.371735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T10:41:00.142543Z digest=sha256:d03b51d45aa7598dbe52003d3ea411142e0579e06e3c0a3ccc22a1568113787c

Observation 0ea168d0-4a6f-44bd-9c15-f0726a94a156 · inbound

LangDriveCTRL: Natural Language Controllable Driving Scene Editing with Multi-modal Agents cites this paper.

LangDriveCTRL: Natural Language Controllable Driving Scene Editing with Multi-modal Agents World Simulation with Video Foundation Models for Physical AI

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-16T20:58:31.824406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T20:56:58.770875Z digest=sha256:60668b686f1d380f11abcece1ca55e8007e7724dc9dc2ba55748ee0ec841ba27

Observation 42fe7b14-e06f-41cd-ac2d-85ccf240be77 · inbound

AstraNav-World: World Model for Foresight Control and Consistency cites this paper.

AstraNav-World: World Model for Foresight Control and Consistency World Simulation with Video Foundation Models for Physical AI

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-16T19:28:20.868354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T19:23:58.769472Z digest=sha256:327fc0b4fc000318df73eabfd580c762eee1703284d5e721af3ce8c10e77d697

Observation 0c5207d5-c917-4f4d-a024-e34e56eab943 · inbound

DriveLaW:Unifying Planning and Video Generation in a Latent Driving World cites this paper.

DriveLaW:Unifying Planning and Video Generation in a Latent Driving World World Simulation with Video Foundation Models for Physical AI

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-16T19:38:21.104574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T19:34:39.518649Z digest=sha256:1ccbb236464d497c96c14abd857322a5a907da601c02bb7c1c609fd4070b9358

Observation 0110820c-4239-499e-a7bd-9fce4a723f27 · inbound

CycleVLA: Proactive Self-Correcting Vision-Language-Action Models via Subtask Backtracking and Minimum Bayes Risk Decoding cites this paper.

CycleVLA: Proactive Self-Correcting Vision-Language-Action Models via Subtask Backtracking and Minimum Bayes Risk Decoding World Simulation with Video Foundation Models for Physical AI

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-03T12:37:28.494262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:37:28.494262Z digest=sha256:e8d11d033ec4a48c912b0b39482cbd3471a26d8133f23d4602f0b6deda243f9a

Observation cd8451bc-67de-427e-94c8-e4efbcaf06f5 · inbound

Advancing Open-source World Models cites this paper.

Advancing Open-source World Models World Simulation with Video Foundation Models for Physical AI

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-05-16T09:07:00.942546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T09:07:00.904794Z digest=sha256:82502086967c3ffc256c8a74fc96adb99dc6e3130302f41548eade5415ae4f15

Observation f9612596-cb44-414f-bc59-5f979c5d309c · inbound

World-VLA-Loop: Closed-Loop Learning of Video World Model and VLA Policy cites this paper.

World-VLA-Loop: Closed-Loop Learning of Video World Model and VLA Policy World Simulation with Video Foundation Models for Physical AI

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T03:56:47.126834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:56:47.126834Z digest=sha256:37d2ac5e7164a4524099de425a65f136d6432482fe394f47f2c14d0dc53e0105

Observation e9d9abb1-f484-4fe4-8a09-de5a2e057d81 · inbound

DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos cites this paper.

DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos World Simulation with Video Foundation Models for Physical AI

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-16T17:02:34.267430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T17:02:33.997887Z digest=sha256:44f5e91e0856ef5036266921eac43677df7d0fc1d0ddd538de753563979f95cd

Observation 7332bf28-128b-4f1d-95c3-82f7f3f881f3 · inbound

Rolling Sink: Bridging Limited-Horizon Training and Open-Ended Testing in Autoregressive Video Diffusion cites this paper.

Rolling Sink: Bridging Limited-Horizon Training and Open-Ended Testing in Autoregressive Video Diffusion World Simulation with Video Foundation Models for Physical AI

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T07:07:30.032697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T07:02:38.876518Z digest=sha256:cf94bd58ec158fa90eda89329b19a722d1551ba44d8de0b9df90dd3ed5fe37a6

Observation d218176b-7e28-444f-b351-30d5d2c1c658 · inbound

RISE: Self-Improving Robot Policy with Compositional World Model cites this paper.

RISE: Self-Improving Robot Policy with Compositional World Model World Simulation with Video Foundation Models for Physical AI

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-16T02:30:32.148569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T02:28:37.997148Z digest=sha256:7df7f25eecf92f74a585d2706031a9f887cfe01ce2fc4cfdeadea7b60938fa30

Observation e8de21ed-7a5e-4bd5-aa18-5e3d79c67d72 · inbound

World Action Models are Zero-shot Policies cites this paper.

World Action Models are Zero-shot Policies World Simulation with Video Foundation Models for Physical AI

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:14.323574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T16:18:15.003371Z digest=sha256:04625068ba17b1ad1befd3b0f8b4b8d1c8cb7efa13bccb4c3d16e82b085c7c79

Observation fd1e209b-bbf7-4fd5-8078-18c3e1614fa1 · inbound

Controllable Egocentric Video Generation via Occlusion-Aware Sparse 3D Hand Joints cites this paper.

Controllable Egocentric Video Generation via Occlusion-Aware Sparse 3D Hand Joints World Simulation with Video Foundation Models for Physical AI

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T12:20:00.593719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T12:18:45.538658Z digest=sha256:a5ff330200cfa7f90a99e02f9ae476ce6d8c7b7815ec161f0424bb968adf0724

Observation 829bb067-274d-43b5-9162-ab1af4f452bf · inbound

Video Generation Models as World Models: Efficient Paradigms, Architectures and Algorithms cites this paper.

Video Generation Models as World Models: Efficient Paradigms, Architectures and Algorithms World Simulation with Video Foundation Models for Physical AI

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-05-14T01:38:36.297567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T01:35:14.878069Z digest=sha256:e89b50e34a64d082c9a42e57fca40d0b8326d6bd238a707c7b9e19b46b94a307

Observation a553edc2-0934-4d58-81eb-de4ec5248f1e · inbound

TRACE: High-Fidelity 3D Scene Editing via Tangible Reconstruction and Geometry-Aligned Contextual Video Masking cites this paper.

TRACE: High-Fidelity 3D Scene Editing via Tangible Reconstruction and Geometry-Aligned Contextual Video Masking World Simulation with Video Foundation Models for Physical AI

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-03T02:31:02.007136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:31:02.007136Z digest=sha256:ce50270ecfaa6a9f498223409e261f1f72bd7ea1b369a3749695e7130077caa4

Observation f428f70b-ebc1-4390-bd81-c670992afeec · inbound

Lifting Unlabeled Internet-level Data for 3D Scene Understanding cites this paper.

Lifting Unlabeled Internet-level Data for 3D Scene Understanding World Simulation with Video Foundation Models for Physical AI

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-13T22:18:21.015364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T22:16:57.890955Z digest=sha256:a00888ce5e6fb013f486badc59a4182d16b81d96858569d060636bb1222e8466

Observation d7239c86-ffbd-4f65-8fc4-5d87d982b5f2 · inbound

OpenWorldLib: A Unified Codebase and Definition of Advanced World Models cites this paper.

OpenWorldLib: A Unified Codebase and Definition of Advanced World Models World Simulation with Video Foundation Models for Physical AI

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:14.323574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:36:42.100191Z digest=sha256:8025ab120a69d68799c898939676ba32a26c038d996f23cfe4d0bad5366bf98e

Observation d6ab8ce6-250c-4191-aa3f-8b7db6070145 · inbound

OpenWorldLib: A Unified Codebase and Definition of Advanced World Models cites this paper.

OpenWorldLib: A Unified Codebase and Definition of Advanced World Models World Simulation with Video Foundation Models for Physical AI

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-13T09:42:23.808691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T09:42:23.808691Z digest=sha256:1782c5d77aae753bd38c435118d13582afd1f86676ddf2c88884abfd87dff95b

Observation abdbe0d3-d15a-446e-b6c7-e80340e0ec62 · inbound

Action Images: End-to-End Policy Learning via Multiview Video Generation cites this paper.

Action Images: End-to-End Policy Learning via Multiview Video Generation World Simulation with Video Foundation Models for Physical AI

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T23:01:14.323574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:51:05.206602Z digest=sha256:b6828006ea5a98371c62db4febf3be3fc6ddec743a98bf82500ea9fce432c4c8

Observation 422b8182-9835-4b95-95b1-4473873e0e50 · inbound

SceneScribe-1M: A Large-Scale Video Dataset with Comprehensive Geometric and Semantic Annotations cites this paper.

SceneScribe-1M: A Large-Scale Video Dataset with Comprehensive Geometric and Semantic Annotations World Simulation with Video Foundation Models for Physical AI

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:14.323574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T17:16:23.355682Z digest=sha256:87df6c982d82ccff8545e4789f8f6a071ba92e928dc12fdaee22ac1ca8f9fce6

Observation 83cdf01c-32a7-4a92-9d8d-2a7220f58fed · inbound

MotionScape: A Large-Scale Real-World Highly Dynamic UAV Video Dataset for World Models cites this paper.

MotionScape: A Large-Scale Real-World Highly Dynamic UAV Video Dataset for World Models World Simulation with Video Foundation Models for Physical AI

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:14.323574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T17:11:14.242522Z digest=sha256:d686098799156ae4bdda7e7f6d3aefe56ebf30ffb999faf0dd1673f18c06b2ab

Observation 50bb55b6-4159-413c-abc3-14c00f504715 · inbound

Representations Before Pixels: Semantics-Guided Hierarchical Video Prediction cites this paper.

Representations Before Pixels: Semantics-Guided Hierarchical Video Prediction World Simulation with Video Foundation Models for Physical AI

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T23:01:14.323574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T16:30:53.578491Z digest=sha256:a9ff3b699ea1cc2fb7de56fba0d0166d9f967eee57b13b2b27b03b1bdfc46846

Observation 780fa249-5e2b-40b1-9161-f5c2d82e6cbf · inbound

SyncFix: Fixing 3D Reconstructions via Multi-View Synchronization cites this paper.

SyncFix: Fixing 3D Reconstructions via Multi-View Synchronization World Simulation with Video Foundation Models for Physical AI

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T23:01:14.323574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:20:04.440192Z digest=sha256:5dff294c77e4c3d7901de719923bf5066a9abb6297cf54f1bfb732faab764d73

Observation 479e6cdb-9b9b-4c34-8193-940ee358f451 · inbound

ShapeGen: Robotic Data Generation for Category-Level Manipulation cites this paper.

ShapeGen: Robotic Data Generation for Category-Level Manipulation World Simulation with Video Foundation Models for Physical AI

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:14.323574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T10:17:05.939312Z digest=sha256:4a61da7f05c7d94ccfba9cbe7776f6eae0e2669c32f65d866075dc2a6874b7a1

Observation 0ed5f6d6-1d1c-4220-bacc-ff5fe532e494 · inbound

From Seeing to Simulating: Generative High-Fidelity Simulation with Digital Cousins for Generalizable Robot Learning and Evaluation cites this paper.

From Seeing to Simulating: Generative High-Fidelity Simulation with Digital Cousins for Generalizable Robot Learning and Evaluation World Simulation with Video Foundation Models for Physical AI

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:14.323574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T08:45:46.944379Z digest=sha256:fceffe13dfa474324dde993a0e1bd82432fe1a39a65c5f45e2ee3b389e218244

Observation 062bf222-1e5a-4795-80f6-c44d909a736a · inbound

Human Cognition in Machines: A Unified Perspective of World Models cites this paper.

Human Cognition in Machines: A Unified Perspective of World Models World Simulation with Video Foundation Models for Physical AI

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T23:01:14.323574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T08:12:15.663761Z digest=sha256:3d8b52f4f66379f266701b6eae7b6baae96077a8fa268c714d8e30e3cf072cf0

Observation 476c1df0-c9b8-43a1-ba13-518b7294a8ab · inbound

UniGeo: Unifying Geometric Guidance for Camera-Controllable Image Editing via Video Models cites this paper.

UniGeo: Unifying Geometric Guidance for Camera-Controllable Image Editing via Video Models World Simulation with Video Foundation Models for Physical AI

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T23:01:14.323574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:03:06.920592Z digest=sha256:12bb458615939ea95f93c713b974f2d4441a9b35d5c7ea0e0917c8563caea1f2

Observation a36559ad-8a7a-46a8-9299-77092c1d3db8 · inbound

UniGeo: Unifying Geometric Guidance for Camera-Controllable Image Editing via Video Models cites this paper.

UniGeo: Unifying Geometric Guidance for Camera-Controllable Image Editing via Video Models World Simulation with Video Foundation Models for Physical AI

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T23:01:14.323574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T02:52:58.232585Z digest=sha256:f95c5fceb28657c632dd71246f0ce44a832c59523af9d669c03851a166f60fd7

Observation 9510abdc-f1c1-4fae-86d0-b88035f03019 · inbound

UniGeo: Unifying Geometric Guidance for Camera-Controllable Image Editing via Video Models cites this paper.

UniGeo: Unifying Geometric Guidance for Camera-Controllable Image Editing via Video Models World Simulation with Video Foundation Models for Physical AI

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-07-05T16:51:14.240617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-05T16:47:32.853010Z digest=sha256:3c359ae8873681fc89f42c29e31ad1648c6a5dca68317624c62098fdaad871aa

Observation a1d3144e-8776-4ac3-b493-34198b3f9357 · inbound

MultiWorld: Scalable Multi-Agent Multi-View Video World Models cites this paper.

MultiWorld: Scalable Multi-Agent Multi-View Video World Models World Simulation with Video Foundation Models for Physical AI

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:14.323574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T05:06:11.514186Z digest=sha256:36886734b795688dc824d279fdaf13fb936d552c8007ef8cece1507024f019b8

Observation 7d185eac-58e0-45a9-8009-1b1b9f3e4da6 · inbound

RoboWM-Bench: A Benchmark for Evaluating World Models in Robotic Manipulation cites this paper.

RoboWM-Bench: A Benchmark for Evaluating World Models in Robotic Manipulation World Simulation with Video Foundation Models for Physical AI

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T23:01:14.323574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T03:02:26.084185Z digest=sha256:8c9b223ec3997af39cc625acb8a765f60064ecc1f1c9f1bcc627a15b1b71e993

Observation 08d65f79-7c4b-4ebf-98d5-bb5732f81ade · inbound

RoboWM-Bench: A Benchmark for Evaluating World Models in Robotic Manipulation cites this paper.

RoboWM-Bench: A Benchmark for Evaluating World Models in Robotic Manipulation World Simulation with Video Foundation Models for Physical AI

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T06:19:50.279253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T06:15:32.881140Z digest=sha256:f4aa2e8c281a0375ab5c8ae223eec38bdbfd9dda2b7d3df8999ef0a92062c1e5

Observation cc575329-78a5-4ad9-8af7-fe0ea9cf7c5d · inbound

Mask World Model: Predicting What Matters for Robust Robot Policy Learning cites this paper.

Mask World Model: Predicting What Matters for Robust Robot Policy Learning World Simulation with Video Foundation Models for Physical AI

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T23:01:14.323574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T02:14:17.676675Z digest=sha256:cd9fac1389afebbac33b1a96ab5cf6e17e652b3b7f301593bf7fd87490c5c0a7

Observation e736b80b-c559-45cf-96f9-f4ffd6c23873 · inbound

Open-H-Embodiment: A Large-Scale Dataset for Enabling Foundation Models in Medical Robotics cites this paper.

Open-H-Embodiment: A Large-Scale Dataset for Enabling Foundation Models in Medical Robotics World Simulation with Video Foundation Models for Physical AI

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:14.323574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-09T23:53:09.614994Z digest=sha256:dd27f0119798d2e5e3645d11a296c60ab107d47c216f3d2842e069e1614e077a

Observation 4e5f7877-17c3-4ec6-a432-db1e23b5cccf · inbound

Hi-WM: Human-in-the-World-Model for Scalable Robot Post-Training cites this paper.

Hi-WM: Human-in-the-World-Model for Scalable Robot Post-Training World Simulation with Video Foundation Models for Physical AI

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:14.323574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-09T21:26:26.540403Z digest=sha256:2b7d3cb56a4c4278c2f1c6ad709565888d0ca1009af962ad720515a745aba7b0

Observation b0566a21-f8cf-4eff-9d9e-a1eba39f51b7 · inbound

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling cites this paper.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling World Simulation with Video Foundation Models for Physical AI

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:14.323574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:1fcff2f56430abb710211f660c10a127cabd4cc04666fb907907c766aae93c08

Observation 34795d18-378b-4dc0-aa2b-d1cebad163ee · inbound

Learning physically grounded traffic accident reconstruction from public accident reports cites this paper.

Learning physically grounded traffic accident reconstruction from public accident reports World Simulation with Video Foundation Models for Physical AI

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:14.323574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-09T19:59:27.929899Z digest=sha256:99144cd2fd1581568a1d14fd15b2d91d0581c906f81d1b9dcf2e77bb22591575

Observation 53bbaa4a-b583-462f-9ea3-f96b48cdd36e · inbound

World Model for Robot Learning: A Comprehensive Survey cites this paper.

World Model for Robot Learning: A Comprehensive Survey World Simulation with Video Foundation Models for Physical AI

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:14.323574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-09T20:38:12.709629Z digest=sha256:1cb93afe48fb6b861f94ea15e13164777ab31b8b13856d93942399fd5dcf0aea

Observation 096cd1c8-87f2-4264-806f-9488893c8fc6 · inbound

Seeing Realism from Simulation: Efficient Video Transfer for Vision-Language-Action Data Augmentation cites this paper.

Seeing Realism from Simulation: Efficient Video Transfer for Vision-Language-Action Data Augmentation World Simulation with Video Foundation Models for Physical AI

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:14.323574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T18:21:00.089755Z digest=sha256:5084c94c3af79bb6a0cf52c3f1a0f9e14b495f2accc9aa643c345ab639e063c7

Observation 530c4c0e-acb4-45b7-a087-145d7c9c3d87 · inbound

RLDX-1 Technical Report cites this paper.

RLDX-1 Technical Report World Simulation with Video Foundation Models for Physical AI

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:14.323574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:06:10.595164Z digest=sha256:c19469b6d4cf2d8e0840838d4ab69d043ef0afa3ce028cedd821535b6b897af9

Observation 3014a16a-ab53-4b04-ae5b-40792c4d8723 · inbound

RLDX-1 Technical Report cites this paper.

RLDX-1 Technical Report World Simulation with Video Foundation Models for Physical AI

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:14.323574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T19:02:19.756092Z digest=sha256:b87350a94d86033b0b718206ff904cdc269ba60d650c06725a2fca4e3d956d5c

Observation 9fb38d36-dac3-47be-8b49-1d3c523199e9 · inbound

Is the Future Compatible? Diagnosing Dynamic Consistency in World Action Models cites this paper.

Is the Future Compatible? Diagnosing Dynamic Consistency in World Action Models World Simulation with Video Foundation Models for Physical AI

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:14.323574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T02:31:30.530595Z digest=sha256:a4091b03602182a0cb62a5dad59abcf25820a6e846fdabba6617a037b448d6d8

Observation c8695a0e-d55a-4eee-ab2c-18803b91979e · inbound

CollabVR: Collaborative Video Reasoning with Vision-Language and Video Generation Models cites this paper.

CollabVR: Collaborative Video Reasoning with Vision-Language and Video Generation Models World Simulation with Video Foundation Models for Physical AI

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:14.323574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T02:17:30.227755Z digest=sha256:95a6ec42509fbd52f3db19dc9c6dae7ec73ca3cce759e7e37989cbbe184a4b9f

Observation a07cccf4-f49a-42f0-af2b-663b6299d52d · inbound

Reinforcing VLAs in Task-Agnostic World Models cites this paper.

Reinforcing VLAs in Task-Agnostic World Models World Simulation with Video Foundation Models for Physical AI

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-13T04:27:14.277708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T04:17:51.349213Z digest=sha256:a079dd47a9dad579b0c7e6884ced9901f35c61b9989ab781c7348c9ea69ac8c7

Observation 986f8d80-af33-4be2-8288-fdd7ba5dba9f · inbound

Reinforcing VLAs in Task-Agnostic World Models cites this paper.

Reinforcing VLAs in Task-Agnostic World Models World Simulation with Video Foundation Models for Physical AI

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-21T08:14:03.037059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T08:13:40.975340Z digest=sha256:61ab1b324bca855896148ec1edd6c922735f8c17f740319306bb8875246863dc

Observation 516c767b-71c5-4c2f-91d5-68956fcce97e · inbound

Di-BiLPS: Denoising induced Bidirectional Latent-PDE-Solver under Sparse Observations cites this paper.

Di-BiLPS: Denoising induced Bidirectional Latent-PDE-Solver under Sparse Observations World Simulation with Video Foundation Models for Physical AI

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-14T19:07:50.987224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T19:06:18.193097Z digest=sha256:d3a0adb0ce9e135543aacd7b5e2193d6df04e8f2b33971ae292594fed153c537

Observation d4c23c5c-4dd4-489f-a9bd-f90ffb6b02f6 · inbound

Coding Agent Is Good As World Simulator cites this paper.

Coding Agent Is Good As World Simulator World Simulation with Video Foundation Models for Physical AI

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-15T02:23:32.034608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T02:19:49.149075Z digest=sha256:0847e8614e8ab15e90afbc9c020ec12ccf6eaedff8d904b38a6df5eb94bc4d06

Observation a078f570-9f76-4132-8f15-06ccece8828e · inbound

Coding Agent Is Good As World Simulator cites this paper.

Coding Agent Is Good As World Simulator World Simulation with Video Foundation Models for Physical AI

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:35:46.852607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T20:55:16.040505Z digest=sha256:5a336d64c3af2accb629993fcf0656277785c69399f882bfaae8da0782031e47

Observation b5ed68ce-66e3-44de-8b02-9e582067c172 · inbound

DriveCtrl: Conditioned Sim-to-Real Driving Video Generation cites this paper.

DriveCtrl: Conditioned Sim-to-Real Driving Video Generation World Simulation with Video Foundation Models for Physical AI

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-06-30T21:15:04.787272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T21:06:06.548538Z digest=sha256:76374985ef51be3a99bd9f341c844a4f64294a5216d86e12340370d0bc103df0

Observation d91c1c62-87f3-4fc8-a354-8c5c442bcfca · inbound

Pelican-Unify 1.0: A Unified Embodied Intelligence Model for Understanding, Reasoning, Imagination and Action cites this paper.

Pelican-Unify 1.0: A Unified Embodied Intelligence Model for Understanding, Reasoning, Imagination and Action World Simulation with Video Foundation Models for Physical AI

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-15T03:19:44.187738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T03:15:35.678011Z digest=sha256:e936ca87df768d2c545a698ab826b26251bbaa213ffbce98e458b5eb9b6fbf29

Observation 173d31fe-008e-4bbe-86ce-73d7b0502d74 · inbound

Pelican-Unify 1.0: A Unified Embodied Intelligence Model for Understanding, Reasoning, Imagination and Action cites this paper.

Pelican-Unify 1.0: A Unified Embodied Intelligence Model for Understanding, Reasoning, Imagination and Action World Simulation with Video Foundation Models for Physical AI

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:46:22.452132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T09:45:13.742534Z digest=sha256:6f9beb61f7ea2e2d8b98fd09815a635a509501fe329e871efc2d488505284d6f

Observation 425ee546-4408-47c4-bd50-17460cbc0c11 · inbound

How to Instruct Your Robot: Dense Language Annotations Power Robot Policy Learning cites this paper.

How to Instruct Your Robot: Dense Language Annotations Power Robot Policy Learning World Simulation with Video Foundation Models for Physical AI

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-20T15:23:25.279166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T15:18:40.457780Z digest=sha256:ebd1f642db80c5dc2737748623c60c930b186ac0fff0f43419daef8476fe54b9

Observation fa290429-9457-46bb-9934-66c82d90dacd · inbound

Self-supervised Hierarchical Visual Reasoning with World Model cites this paper.

Self-supervised Hierarchical Visual Reasoning with World Model World Simulation with Video Foundation Models for Physical AI

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T12:43:16.866382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T12:41:17.559901Z digest=sha256:50adc7055765d69c2b4973f3f8d1e74b0156c219973bdba9c9b6ace9f0710821

Observation 95302fe6-4ace-4927-a355-6ff92b627d65 · inbound

Self-supervised Hierarchical Visual Reasoning with World Model cites this paper.

Self-supervised Hierarchical Visual Reasoning with World Model World Simulation with Video Foundation Models for Physical AI

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T19:05:00.273959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T19:04:33.433907Z digest=sha256:7685578e6b72fc3e70d3cee4ed432f4846eafd7e52bc3307df205b1d88ddcb2f

Observation 88a07589-9b1c-4766-8c1e-09818ccd60fa · inbound

WorldArena 2.0: Extending Embodied World Model Benchmarking on Modality, Functionality and Platform cites this paper.

WorldArena 2.0: Extending Embodied World Model Benchmarking on Modality, Functionality and Platform World Simulation with Video Foundation Models for Physical AI

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-05-20T11:03:13.688410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T10:59:54.879907Z digest=sha256:6a8f1ff4ebb919b907f6f2a461a1b69a0348b2eebccca9ec91d0f1cb6de740be

Observation 5d7afa79-4545-41d9-832b-d6af11a6f445 · inbound

NEWTON: Agentic Planning for Physically Grounded Video Generation cites this paper.

NEWTON: Agentic Planning for Physically Grounded Video Generation World Simulation with Video Foundation Models for Physical AI

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-20T10:58:13.800270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T10:56:22.343333Z digest=sha256:b9a53bfe1081ad56c6899c448209a61c0fe7d49b99811dfca7d796de404e0c51

Observation d1e93c55-63b7-496a-bac3-7e435c202e64 · inbound

PhyWorld: Physics-Faithful World Model for Video Generation cites this paper.

PhyWorld: Physics-Faithful World Model for Video Generation World Simulation with Video Foundation Models for Physical AI

Reference 64

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T07:33:07.562468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T07:28:20.248452Z digest=sha256:c046efececc3241336ce400b3ea0391440e3c1c2cc958365d54c31fcae2cee7d

Observation 2d959405-5891-4b63-899e-9b1337ae3a52 · inbound

World-Ego Modeling for Long-Horizon Evolution in Hybrid Embodied Tasks cites this paper.

World-Ego Modeling for Long-Horizon Evolution in Hybrid Embodied Tasks World Simulation with Video Foundation Models for Physical AI

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-20T06:38:05.636274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T06:36:47.265734Z digest=sha256:95235435670c816b4fc4f73809a8b705f4e2474fdd287b6f0d8779260197b2e4

Observation 9c206a40-a2dd-4020-a16e-0e2c99729dab · inbound

Accelerating Video Inverse Problem Solvers with Autoregressive Diffusion Models cites this paper.

Accelerating Video Inverse Problem Solvers with Autoregressive Diffusion Models World Simulation with Video Foundation Models for Physical AI

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-21T06:13:59.578801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T06:12:31.783615Z digest=sha256:76d0587204f00e6843fedb28924ebad867fc8086893599bd0297ab7202f929db

Observation 1a8f9b0a-1dbb-49ec-a680-f27e3d4dae36 · inbound

CRONOS: Benchmarking Counterfactual Physical Consistency in Video Models cites this paper.

CRONOS: Benchmarking Counterfactual Physical Consistency in Video Models World Simulation with Video Foundation Models for Physical AI

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-25T04:40:23.268359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-25T04:39:22.400458Z digest=sha256:70ffd86601180e7a2ef956e76d1f2332fb1c10b51ab88040431e6db38d56d32a

Observation 7cfd86f1-b242-4fd6-92ef-9e9ec2370a0f · inbound

What-If World: A Causal Benchmark for General World Models in Embodied Scenarios cites this paper.

What-If World: A Causal Benchmark for General World Models in Embodied Scenarios World Simulation with Video Foundation Models for Physical AI

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-06-29T18:23:50.356445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T18:23:22.987086Z digest=sha256:c691677832b62d221e4a0e9da6af853eb1fe7cb9da36a629e105287e2fb4db57

Observation 68f05e78-715d-4bcd-a2b8-b9b3a4a8522a · inbound

Proprio: Latent Self-Scoring and Inference-Time Refinement for Physically Plausible Video Generation cites this paper.

Proprio: Latent Self-Scoring and Inference-Time Refinement for Physically Plausible Video Generation World Simulation with Video Foundation Models for Physical AI

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:03:26.638052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T12:55:24.689338Z digest=sha256:10b60269f35018a7212c9c535576bcf4ecb8ab6b36a078c88a2932e5a17fe82f

Observation 2e702759-30b6-42cf-aadf-a3d2d85f50ad · inbound

Gamma-World: Generative Multi-Agent World Modeling Beyond Two Players cites this paper.

Gamma-World: Generative Multi-Agent World Modeling Beyond Two Players World Simulation with Video Foundation Models for Physical AI

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:53:29.115357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T13:34:03.020051Z digest=sha256:59c0d5c9798eb82a7135f16ef6d203bc37872094e2554c9a94b1037d793d8e45

Observation 226e4e9c-60ca-4416-9714-e137062acefc · inbound

World Models: A Comprehensive Survey of Architectures, Methodologies, Reasoning Paradigms, and Applications cites this paper.

World Models: A Comprehensive Survey of Architectures, Methodologies, Reasoning Paradigms, and Applications World Simulation with Video Foundation Models for Physical AI

Reference 108

Resolution
verified exact
local_arxiv, observed 2026-06-29T08:43:15.654464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T08:36:23.776293Z digest=sha256:c03bd40c966900d89b4cab1bdcc26ce9c45a55916e0646c79ef3196c63b0f6b3

Observation 0713b3dc-b9e4-4d64-a9ae-05704cec4938 · inbound

OptiWorld: Optimal Control for Video World Generation under Physical Constraints cites this paper.

OptiWorld: Optimal Control for Video World Generation under Physical Constraints World Simulation with Video Foundation Models for Physical AI

Reference 91

Resolution
verified exact
local_arxiv, observed 2026-06-28T19:32:35.509546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T19:02:51.848742Z digest=sha256:409a96438a4984cd1f0cfea97cfa1dd143ad7c384becd4f5c773c636c91ae6d0

Observation 49d26b9e-d761-4967-be0f-f5593d6b32d7 · inbound

$\tau_0$-WM: A Unified Video-Action World Model for Robotic Manipulation cites this paper.

$\tau_0$-WM: A Unified Video-Action World Model for Robotic Manipulation World Simulation with Video Foundation Models for Physical AI

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-06-28T17:22:24.504637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T17:21:06.079898Z digest=sha256:fb0bffb021766359f1076383fd9fef5d42c8bc8ef0ae42700b30e828a4fa3783

Observation ee9ffe6e-34bc-44e0-994b-e200ff4a3f33 · inbound

Beyond Task Success: Behavioral and Representational Diagnostics for WAM and VLA cites this paper.

Beyond Task Success: Behavioral and Representational Diagnostics for WAM and VLA World Simulation with Video Foundation Models for Physical AI

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-06-28T17:22:24.413659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T17:21:15.500408Z digest=sha256:876c36ca651f815d99fdfbd9b37c60edc1ae580055693fcf7e88f7fc9977e94e

Observation 0866c75f-eb61-44fb-a5c6-f73e10f09d6a · inbound

Dexterity-BEV: Aligning 3D World and Actions for Generalizable Robot Policies Learning cites this paper.

Dexterity-BEV: Aligning 3D World and Actions for Generalizable Robot Policies Learning World Simulation with Video Foundation Models for Physical AI

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-07-01T23:06:20.428785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T14:41:50.084254Z digest=sha256:30cebabd6778e2fb56a35f616b9b4edda55806e2e60b13da26962091d65afb7c

Observation 47e0ba62-cfe3-4f77-8281-3863146feb2d · inbound

RoboDream: Compositional World Models for Scalable Robot Data Synthesis cites this paper.

RoboDream: Compositional World Models for Scalable Robot Data Synthesis World Simulation with Video Foundation Models for Physical AI

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-07-01T23:36:23.584650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T14:07:31.809341Z digest=sha256:3bdc425c7d153e6faac97442f712fb2ae9751e593ffa9b465f475657adf550d2

Observation b294cd98-cc39-4856-941a-61a9052f498e · inbound

NVIDIA OmniDreams: Real-Time Generative World Model for Closed-Loop Autonomous Vehicle Simulation cites this paper.

NVIDIA OmniDreams: Real-Time Generative World Model for Closed-Loop Autonomous Vehicle Simulation World Simulation with Video Foundation Models for Physical AI

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:06:27.451431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T11:11:46.428727Z digest=sha256:435fc0da82be00ea50229e3c4f82dac6ed92c6302278a770a6ea4c63610edb86

Observation ba28d442-6e81-405f-8219-aa41da504410 · inbound

NVIDIA OmniDreams: Real-Time Generative World Model for Closed-Loop Autonomous Vehicle Simulation cites this paper.

NVIDIA OmniDreams: Real-Time Generative World Model for Closed-Loop Autonomous Vehicle Simulation World Simulation with Video Foundation Models for Physical AI

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T12:34:17.389528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:34:17.389528Z digest=sha256:885476c34df78da65fb5795819b0afaed0d82ad449fbaf40da38f41bc8281e87

Observation 1177adce-4513-4cb5-997f-ef3e4beaeebd · inbound

PointAction: 3D Points as Universal Action Representations for Robot Control cites this paper.

PointAction: 3D Points as Universal Action Representations for Robot Control World Simulation with Video Foundation Models for Physical AI

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-02T03:16:34.866751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T10:09:48.280446Z digest=sha256:d844e3be20f06c50153c80d92c12b03948723d61056b02a9faebc96cad217dd3

Observation b7c62be3-a5c4-4a6c-a8f0-2cfdb861045b · inbound

OSCAR: Omni-Embodiment Action-Conditioned World Model for Robotics cites this paper.

OSCAR: Omni-Embodiment Action-Conditioned World Model for Robotics World Simulation with Video Foundation Models for Physical AI

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-02T07:56:47.049610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T06:35:26.480518Z digest=sha256:8665c6b23311921a14bc7cc6050b12fb19a109568a03639cae1a26d95124344e

Observation 5a214f1b-b761-493a-8343-d66e691ec6f6 · inbound

Dream-Tac: A Unified Tactile World Action Model for Contact-Rich Robot Manipulation cites this paper.

Dream-Tac: A Unified Tactile World Action Model for Contact-Rich Robot Manipulation World Simulation with Video Foundation Models for Physical AI

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T23:27:26.854388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T18:17:06.288698Z digest=sha256:b7879ebcf547061c31611d1a3821bdc7320c9a4a60e6f64e1ace336be5f0ab8c

Observation 26aa952f-0ba2-47d9-ab50-2cacda776694 · inbound

MotionWAM: Towards Foundation World Action Models for Real-Time Humanoid Loco-Manipulation cites this paper.

MotionWAM: Towards Foundation World Action Models for Real-Time Humanoid Loco-Manipulation World Simulation with Video Foundation Models for Physical AI

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-07-03T01:37:30.721968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:26:25.136391Z digest=sha256:baa7c785bb72dedff08296c4d43f575d62a271b806e1e5517587f9ddad28ee1c

Observation b27474e3-af7a-4011-9ae8-ecbb0a44072b · inbound

Targeting World Models to Compromise Robot Learning Pipelines cites this paper.

Targeting World Models to Compromise Robot Learning Pipelines World Simulation with Video Foundation Models for Physical AI

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-06-27T16:41:03.690648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:05:01.264700Z digest=sha256:e543865705fcf8d927cee83ff46ad5259725a5bc28c7ae2586c8acfa8d39a834

Observation 8485be8d-3234-48c7-9d2e-6114cd9a9e1a · inbound

Prisma-World: Camera-Controllable Multi-Agent Video World Model cites this paper.

Prisma-World: Camera-Controllable Multi-Agent Video World Model World Simulation with Video Foundation Models for Physical AI

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-07-03T00:47:30.157870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T17:03:29.676482Z digest=sha256:3f43c195cb351ccc0a0d2a872a3e310ce0d7d9f0eb608d4aa964224ebc164544

Observation e6590b4a-8ccb-4bc7-bf74-c81556e78738 · inbound

Echo-Memory: A Controlled Study of Memory in Action World Models cites this paper.

Echo-Memory: A Controlled Study of Memory in Action World Models World Simulation with Video Foundation Models for Physical AI

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-03T00:57:29.762501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:58:37.552036Z digest=sha256:a2f5f90314c3beec5092d153e0b4c0cf7237e798b985604458f371b6fdce1def

Observation decda045-f23d-4377-8905-fc6f7ecfd57a · inbound

MemoryVLA++: Temporal Modeling via Memory and Imagination in Vision-Language-Action Models cites this paper.

MemoryVLA++: Temporal Modeling via Memory and Imagination in Vision-Language-Action Models World Simulation with Video Foundation Models for Physical AI

Reference 64

Resolution
verified exact
local_arxiv, observed 2026-07-03T01:57:32.223652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:13:19.238535Z digest=sha256:203b613de34d96d90828522de8f9651a1e84e7c085ff4ddcb8e84d320dbb2aa2

Observation 8e9fe27f-cdae-47c3-a8db-356965475274 · inbound

Hierarchical Policies from Verbal and Egocentric Human Signals for Natural Human-Robot Interaction cites this paper.

Hierarchical Policies from Verbal and Egocentric Human Signals for Natural Human-Robot Interaction World Simulation with Video Foundation Models for Physical AI

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-03T04:57:38.603718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T13:31:02.640975Z digest=sha256:4afef153a0894bff1595bdd46e0d5fc8e96f66e0f51ff9d977fc9604a699af33

Observation bb4167c1-d3f8-43e7-8d01-f4cdd868012f · inbound

Envision4D: Envisioning Visual Futures via Feed-forward 4D Gaussian Splatting for Autonomous Driving cites this paper.

Envision4D: Envisioning Visual Futures via Feed-forward 4D Gaussian Splatting for Autonomous Driving World Simulation with Video Foundation Models for Physical AI

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-07-03T04:37:36.898884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T13:48:15.379724Z digest=sha256:01e19358e0e22d16a293f69e42027acf0b2365de62eb503056633cbf584b996b

Observation 7e5f1ff1-3baf-4992-b0b3-3dc0fbbafba8 · inbound

WorldOlympiad: Can Your World Model Survive a Triathlon? cites this paper.

WorldOlympiad: Can Your World Model Survive a Triathlon? World Simulation with Video Foundation Models for Physical AI

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-03T05:47:41.258522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T13:05:26.397711Z digest=sha256:722f85aa14e165d291986e2e4b425f57883abd489c1daa8baef41080c04faf2c

Observation db0b85f4-8434-45ba-a162-ce39c96f6224 · inbound

World Pilot: Steering Vision-Language-Action Models with World-Action Priors cites this paper.

World Pilot: Steering Vision-Language-Action Models with World-Action Priors World Simulation with Video Foundation Models for Physical AI

Reference 71

Resolution
verified exact
local_arxiv, observed 2026-07-03T11:18:03.367479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T09:40:02.137152Z digest=sha256:e0c839876696f442aa539819868f10e198440ff20ba29016cb9a32411e3f0995

Observation f712b579-d06a-42cb-a67e-fab5be5b2317 · inbound

RepWAM: World Action Modeling with Representation Visual-Action Tokenizers cites this paper.

RepWAM: World Action Modeling with Representation Visual-Action Tokenizers World Simulation with Video Foundation Models for Physical AI

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-03T14:58:33.402657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T06:47:05.236028Z digest=sha256:2b1737c70eac4661db2e5e45664987e032fe6265be920680bb84aafb507169e5

Observation 07740f9e-264d-4bb3-9d19-d02040054ca3 · inbound

JoyAI-Sim: A Simulation-Enabled Interconversion Toolchain for the Embodied Data Pyramid cites this paper.

JoyAI-Sim: A Simulation-Enabled Interconversion Toolchain for the Embodied Data Pyramid World Simulation with Video Foundation Models for Physical AI

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-06-30T10:34:36.285000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T10:31:54.292897Z digest=sha256:f1181fc3a04f83a978ea08a78e4e2303509a8819d4123741d5a7d85f2f626259

Observation ea456dda-ca92-4b09-b322-d91e83309386 · inbound

Unified Motion-Action Modeling for Heterogeneous Robot Learning cites this paper.

Unified Motion-Action Modeling for Heterogeneous Robot Learning World Simulation with Video Foundation Models for Physical AI

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-07-03T17:28:44.450366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T04:09:06.885372Z digest=sha256:47b76a946caf699fb6e4b888baeaa258488a0f0a53cfe7e5ff1c6af753a6ee4f

Observation a0374d14-7131-4951-9926-ab50da5b3f6f · inbound

PAIWorld: A 3D-Consistent World Foundation Model for Robotic Manipulation cites this paper.

PAIWorld: A 3D-Consistent World Foundation Model for Robotic Manipulation World Simulation with Video Foundation Models for Physical AI

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-07-03T21:38:58.746884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T00:19:33.170645Z digest=sha256:b2ed9b1205f5ddbee5dc3c47aeaded36203aed7b5ccf52158a123fa68bdcf873

Observation 53cf68d8-2acc-4d4f-96d2-9a873d364131 · inbound

Data-Forcing Distillation: Restoring Diversity and Fidelity in Few-Step Video Generation cites this paper.

Data-Forcing Distillation: Restoring Diversity and Fidelity in Few-Step Video Generation World Simulation with Video Foundation Models for Physical AI

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-03T21:08:58.095744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T00:55:20.691480Z digest=sha256:d2ac4611596dda9f0db740f45ea49ab01ce0a553f2f4f770c5ad639ce753736e

Observation f2a039eb-731e-4d73-81c0-460d909bde03 · inbound

SC3-Eval: Evaluating Robot Foundation Models via Self-Consistent Video Generation cites this paper.

SC3-Eval: Evaluating Robot Foundation Models via Self-Consistent Video Generation World Simulation with Video Foundation Models for Physical AI

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:19:13.396326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T21:18:07.494794Z digest=sha256:4f96c835d1b70062225a4111ba01501fcb1fb63b6b7594d3f21db7bd71326dda

Observation 43ee7b49-2995-4cb3-8038-0041d7dca188 · inbound

SC3-Eval: Evaluating Robot Foundation Models via Self-Consistent Video Generation cites this paper.

SC3-Eval: Evaluating Robot Foundation Models via Self-Consistent Video Generation World Simulation with Video Foundation Models for Physical AI

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-06-29T18:23:51.418723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T05:06:30.594399Z digest=sha256:e31ac521211c8c1437c6b2e167a24d58de4f8cd8c040f8cb259775c52c30efc8

Observation fd63755f-e389-441f-af67-1969f010d5ae · inbound

Mem-World: Memory-Augmented Action-Conditioned World Models for Persistent Robot Manipulation cites this paper.

Mem-World: Memory-Augmented Action-Conditioned World Models for Persistent Robot Manipulation World Simulation with Video Foundation Models for Physical AI

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-07-03T23:39:04.851790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T21:50:26.223702Z digest=sha256:1c2c95571a78af8a0db265ddb5406061232cf81d27c3f860ed684201e07aad9f

Observation 0186237b-8460-4fd2-ac03-af8946a85dba · inbound

World Engine: Towards the Era of Post-Training for Autonomous Driving cites this paper.

World Engine: Towards the Era of Post-Training for Autonomous Driving World Simulation with Video Foundation Models for Physical AI

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-07-04T03:59:33.870765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T17:21:15.456982Z digest=sha256:68567a50f9b5fdfa7ef0929b95e8b9909670d79a7fa4635ad679348925f3c575

Observation 143010e0-b851-4b85-8b31-7d2f1e0053cc · inbound

FOCA: Future-Oriented Conditioning for Data-Efficient Vision-Language-Action Adaptation cites this paper.

FOCA: Future-Oriented Conditioning for Data-Efficient Vision-Language-Action Adaptation World Simulation with Video Foundation Models for Physical AI

Reference 66

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T03:29:30.391838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-26T18:01:08.616612Z digest=sha256:571d21875b7e8f5a3313fa4d177928e702035e895dc99c810bea7c21e2f03179

Observation f6cfe78c-ec3b-4456-92f7-a7984f1a9b44 · inbound

Foresight: Failure Detection for Long-Horizon Robotic Manipulation with Action-Conditioned World Model Latents cites this paper.

Foresight: Failure Detection for Long-Horizon Robotic Manipulation with Action-Conditioned World Model Latents World Simulation with Video Foundation Models for Physical AI

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-07-04T10:39:44.765169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T08:43:06.983870Z digest=sha256:adbe6c11111d55292f8889246ad499b02beeedefbe9a300866f7fd94bee2ca28

Observation c884a9f7-45c6-4d56-a02b-b543c1a27859 · inbound

Qwen-AgentWorld: Language World Models for General Agents cites this paper.

Qwen-AgentWorld: Language World Models for General Agents World Simulation with Video Foundation Models for Physical AI

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-04T17:09:59.232510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-25T23:52:31.403419Z digest=sha256:b2709367744e0c1c387c13c8134d160f8f7805e88c86e3a7f9da14a5cbb9da9a

Observation 345b4521-a0b1-4df9-9d05-5040bfd2a2a2 · inbound

Physics Question Scene Graph: Fine-grained Evaluation of Physical Plausibility in Text-to-Video Generation cites this paper.

Physics Question Scene Graph: Fine-grained Evaluation of Physical Plausibility in Text-to-Video Generation World Simulation with Video Foundation Models for Physical AI

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T19:20:05.904375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-25T21:33:38.643889Z digest=sha256:54dc8889bfa1ce96ad434b86072bd4d6151db0c0b17e6647c2e3be40a5be5479

Observation 2e214b17-4e5c-4ee0-8ac1-1fa7967f7900 · inbound

Causal-rCM: A Unified Teacher-Forcing and Self-Forcing Open Recipe for Autoregressive Diffusion Distillation in Streaming Video Generation and Interactive World Models cites this paper.

Causal-rCM: A Unified Teacher-Forcing and Self-Forcing Open Recipe for Autoregressive Diffusion Distillation in Streaming Video Generation and Interactive World Models World Simulation with Video Foundation Models for Physical AI

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-04T19:50:11.367469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-25T20:57:30.765802Z digest=sha256:11cd59f3fb0516fb94d4ef10bebca7ef84ab49d96086f341786d3b23e0b13810

Observation b0d95287-a3c6-4785-a94c-83f36ee7c640 · inbound

Not All Actions Are Equal: Rethinking Conditioning for Dexterous World Model cites this paper.

Not All Actions Are Equal: Rethinking Conditioning for Dexterous World Model World Simulation with Video Foundation Models for Physical AI

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-07-04T13:19:50.401100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T05:23:07.046527Z digest=sha256:59d8788c40e752baaa3fec7b8ebf4c66c0ba723e1a20b09281abfb96089d21e1