Pith. sign in

Paper Citation Record · LEDGER

Video World Models with Long-term Spatial Memory

As of 18 August 2026, this Paper Citation Record lists 91 of 91 outbound references and 44 inbound Pith citation observations for arXiv:2506.05284.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.05284 v1

Coverage vector

measured 91 of 91 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:28:43.713731Z

measured 135 of 135 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 44 of 44 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T05:06:52.002085Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T17:58:47.769652Z

Reference resolution

91 of 91 outbound references displayed

  • verified exact0
  • verified fuzzy25
  • unresolved66
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 91a5c251-1319-4aca-b8bf-4882500c1441 · outbound

This paper cites Diffusion for world modeling: Visual details matter in atari.

Video World Models with Long-term Spatial Memory Diffusion for world modeling: Visual details matter in atari

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:35.832506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:35.832506Z digest=sha256:434afc0ae5ed5a538d32b3ec88789347516ea96843764416b37d4628b9fab8c1

Observation 6bf78689-d1cd-4efc-a1d0-84e1097b0921 · outbound

This paper cites Genesis: A universal and generative physics engine for robotics and beyond.

Video World Models with Long-term Spatial Memory Genesis: A universal and generative physics engine for robotics and beyond

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:35.902884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:35.902884Z digest=sha256:7c4f68f6d465559c7772927c9b121b05e7749fb30c8b1e6ca0c506a57d733605

Observation f27438cb-c0fd-4dac-ab76-fb27d11c4915 · outbound

This paper cites Essentials of human memory.

Video World Models with Long-term Spatial Memory Essentials of human memory

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:35.974939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:35.974939Z digest=sha256:425ccc98fc87dbb7c25f549f4c6d3353d40fc0222b33434f60e0bec32cb382ce

Observation 6392592d-07b3-41ca-a3f2-5c8c1645d581 · outbound

This paper cites Lindell, and Sergey Tulyakov.

Video World Models with Long-term Spatial Memory Lindell, and Sergey Tulyakov

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:36.071006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:36.071006Z digest=sha256:781d3423c38dc45092a61bcade585d9982637992a579c28b6edaefec250f3d67

Observation 80fbbd4e-104c-4f2b-ba09-235053dc10ee · outbound

This paper cites Lindell, and Sergey Tulyakov.

Video World Models with Long-term Spatial Memory Lindell, and Sergey Tulyakov

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:36.163159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:36.163159Z digest=sha256:856b86aedfaa8ae45c4ea2283525ead9df65799bb4a057f8fd3aaa2b68332ccc

Observation c885a79c-d2f5-4946-be26-8558814b2f3d · outbound

This paper cites ReCamMaster: Camera-Controlled Generative Rendering from A Single Video.

Video World Models with Long-term Spatial Memory ReCamMaster: Camera-Controlled Generative Rendering from A Single Video

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:36.252319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:36.252319Z digest=sha256:a507b2156121bf6f62e4b1ef1051f6d5021d002662a9c06ee6ec64a19607dd61

Observation bfbd776c-2d70-482a-b724-59789be04f9f · outbound

This paper cites Syncammaster: Synchronizing multi-camera video generation from diverse viewpoints.

Video World Models with Long-term Spatial Memory Syncammaster: Synchronizing multi-camera video generation from diverse viewpoints

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:36.339653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:36.339653Z digest=sha256:473473a1295483d8441c92205d4b172bb5ced09887b684cfd6b7f58e3053c2ef

Observation 4862582a-f8fc-41a9-a43b-5be1a04433a4 · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

Video World Models with Long-term Spatial Memory Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:36.546649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:36.546649Z digest=sha256:e020704e22b18084428c6ffe2fa73ee581e29a1a13415b93fa842cf44d967326

Observation 923a043e-f882-477b-b327-baef14bfdb3d · outbound

This paper cites Video generation models as world simulators.

Video World Models with Long-term Spatial Memory Video generation models as world simulators

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:36.680441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:36.680441Z digest=sha256:08aabd0b3452826489036f07825b41306e68b6d3db17cebdc8a5a824bf9e60f9

Observation 1c1d54ed-5491-4d28-aed4-993cdebfc635 · outbound

This paper cites Video generation models as world simulators.

Video World Models with Long-term Spatial Memory Video generation models as world simulators

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:36.763659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:36.763659Z digest=sha256:a5f33f8e0d275d9df5ab2e722fec63f45c7f53bd7c0f4b64f3970d5eaa79aa35

Observation 5f241a0b-d1f9-4468-8b2e-ef6c19675a03 · outbound

This paper cites GameGen-X: Interactive Open-world Game Video Generation.

Video World Models with Long-term Spatial Memory GameGen-X: Interactive Open-world Game Video Generation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:36.827996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:36.827996Z digest=sha256:330fdaed000b058c518f0d93d2da5b3836cac382267211b5d6a9e8a0fcb89919

Observation 2dc9de1d-22af-4dcd-b894-9e2cee932aad · outbound

This paper cites Diffusion forcing: Next-token prediction meets full-sequence diffusion.

Video World Models with Long-term Spatial Memory Diffusion forcing: Next-token prediction meets full-sequence diffusion

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:36.898038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:36.898038Z digest=sha256:f712720fdbc50b7587108ab3753927d2ff11514db3815ac60ea7b434cff08fe9

Observation 93eaeb15-0c62-4f99-87ab-e5a61f4c15aa · outbound

This paper cites Skyreels-v2: Infinite-length film generative model.

Video World Models with Long-term Spatial Memory Skyreels-v2: Infinite-length film generative model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:36.959518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:36.959518Z digest=sha256:0fe4ea5f40849d3246667e718bae76171eb6400dcf6a57402de8c90349314795

Observation d1cfa4da-a4d7-4852-a9e8-f397c0aa5e7f · outbound

This paper cites FlexWorld: Progressively Expanding 3D Scenes for Flexiable-View Synthesis.

Video World Models with Long-term Spatial Memory FlexWorld: Progressively Expanding 3D Scenes for Flexiable-View Synthesis

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:37.035329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:37.035329Z digest=sha256:1523352a98d5b18b0ddb180559332e729f75bd73973a9e4b18b5c588b0005e67

Observation faf5cd2c-3ffc-40b3-a2ec-26c23d9b1102 · outbound

This paper cites Seine: Short-to-long video diffusion model for generative transition and prediction.

Video World Models with Long-term Spatial Memory Seine: Short-to-long video diffusion model for generative transition and prediction

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:37.101723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:37.101723Z digest=sha256:0bf8c70c27d787f7d4daae0fec363823c11d3b5017b6783f41c551b629b7deee

Observation 567fd608-ebdf-4eeb-b0f1-419488d02439 · outbound

This paper cites Oasis: A universe in a transformer.

Video World Models with Long-term Spatial Memory Oasis: A universe in a transformer

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:50.011566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:28:37.168324Z digest=sha256:587a4851989c1ab7d888daeb7e1c4568cfafb24878900ee234c4073a580e3f71

Observation f177e406-38fd-411d-a3ae-97ee406d55a2 · outbound

This paper cites Oasis: A universe in a transformer.

Video World Models with Long-term Spatial Memory Oasis: A universe in a transformer

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:49.729793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:28:37.259088Z digest=sha256:15fc8627b49e0e0177f2c876d5c63f1c6f97967d7e6405198344a49dce28c67d

Observation 1d6484bd-3529-4b7d-b7e0-813a1e1e18c5 · outbound

This paper cites Diffusion Models Beat GANs on Image Synthesis.

Video World Models with Long-term Spatial Memory Diffusion Models Beat GANs on Image Synthesis

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:37.342105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:37.342105Z digest=sha256:54f542e8de829adda497fc8af586d80c977180b01b7cc900748d2fc99b612f80

Observation 9dda24dc-473b-4a7c-8fe5-ae754093ebe2 · outbound

This paper cites Skyreels-a2: Compose anything in video diffusion transformers.

Video World Models with Long-term Spatial Memory Skyreels-a2: Compose anything in video diffusion transformers

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:49.537809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:28:37.412842Z digest=sha256:cbd5c0be036ee5f5bee7e3df009b6d2d729dcadbf05d11d62912af164e41a96c

Observation 457f001d-7f25-4616-9ce8-e5060f72732d · outbound

This paper cites The Matrix: Infinite-Horizon World Generation with Real-Time Moving Control.

Video World Models with Long-term Spatial Memory The Matrix: Infinite-Horizon World Generation with Real-Time Moving Control

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:37.515228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:37.515228Z digest=sha256:3fcfa73b4b59089e89438531b8767b0207cf0c027336cd92bcb436deaa2808f8

Observation 9fe474e8-794c-45fa-b6ea-aea7aff9f106 · outbound

This paper cites ViD-GPT: Introducing GPT-style Autoregressive Generation in Video Diffusion Models.

Video World Models with Long-term Spatial Memory ViD-GPT: Introducing GPT-style Autoregressive Generation in Video Diffusion Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:37.579879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:37.579879Z digest=sha256:1884ab4666db333e3210df4d1c33383a4616fb17c8625f871516db28006ddedf

Observation 80aac502-dcbf-4189-9864-0b149aa2a270 · outbound

This paper cites Ca2-VDM: Efficient Autoregressive Video Diffusion Model with Causal Generation and Cache Sharing.

Video World Models with Long-term Spatial Memory Ca2-VDM: Efficient Autoregressive Video Diffusion Model with Causal Generation and Cache Sharing

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:37.670335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:37.670335Z digest=sha256:7f6de29a4fb33448ea6643c7cf3e07298716b4b39d5db16a614769c4d724cb43

Observation 8da434c8-387a-4965-a4a8-655836b3b94c · outbound

This paper cites Long-Context Autoregressive Video Modeling with Next-Frame Prediction.

Video World Models with Long-term Spatial Memory Long-Context Autoregressive Video Modeling with Next-Frame Prediction

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:37.767483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:37.767483Z digest=sha256:1c487f2923a75399073e6319acd18e481acb395a23700234fcb731600715fdd2

Observation b2602980-6b33-4dd4-87d9-7a67ca80bfc8 · outbound

This paper cites Diffusion as Shader: 3D-aware Video Diffusion for Versatile Video Generation Control.

Video World Models with Long-term Spatial Memory Diffusion as Shader: 3D-aware Video Diffusion for Versatile Video Generation Control

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:37.873352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:37.873352Z digest=sha256:79a36f00c0d46f88bc92f9a2aef9ca9df6255975670f530c438bc202ce04221d

Observation 2ea90706-e38e-4690-8667-0c9fcd51b660 · outbound

This paper cites Photorealistic video generation with diffusion models.

Video World Models with Long-term Spatial Memory Photorealistic video generation with diffusion models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:49.316341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:28:37.961424Z digest=sha256:417e5b7596b9ecedd08f276a2b1218e84c4ca0603ca96c0de97aa3b6c67551e3

Observation 0c757e0e-edd2-4879-8014-0db92081aa5a · outbound

This paper cites Recurrent world models facilitate policy evolution.

Video World Models with Long-term Spatial Memory Recurrent world models facilitate policy evolution

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:49.027141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:28:38.041172Z digest=sha256:b827996fbbe4f1622ed17f12efd95bfd098b7027ae87c7ca5ef3f56c97c3fbeb

Observation 35f74367-9eb2-4c6f-9c02-dab29c2e1aab · outbound

This paper cites LTX-Video: Realtime Video Latent Diffusion.

Video World Models with Long-term Spatial Memory LTX-Video: Realtime Video Latent Diffusion

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:38.104382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:38.104382Z digest=sha256:b42ed7b8bd9accfe0a9866a935609a7f6c18b54888060f7a509a2841771f0d58

Observation 948e9737-8f24-49ba-b688-82b102fc7d28 · outbound

This paper cites Dream to Control: Learning Behaviors by Latent Imagination.

Video World Models with Long-term Spatial Memory Dream to Control: Learning Behaviors by Latent Imagination

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:38.173494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:38.173494Z digest=sha256:ddaa89785f9616b487391b8ab23e7a7ece13536684dcb8d68ab3e1c8b3083c31

Observation bab5eb94-9035-47ed-aa7c-fed1d6fe2ee9 · outbound

This paper cites Cameractrl: Enabling camera control for video diffusion models.

Video World Models with Long-term Spatial Memory Cameractrl: Enabling camera control for video diffusion models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:48.771677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:28:38.253795Z digest=sha256:b7e064fcf8275434db50fe17d45346f5c7f86d3a4476f012db7080b8be944b6f

Observation 0bc675b5-4a32-4197-ae02-6ab13e16b98f · outbound

This paper cites CameraCtrl II: Dynamic Scene Exploration via Camera-controlled Video Diffusion Models.

Video World Models with Long-term Spatial Memory CameraCtrl II: Dynamic Scene Exploration via Camera-controlled Video Diffusion Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:38.330512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:38.330512Z digest=sha256:9f7009b8ae6ae1802930a4cd5e9adfad3c7545c9c8ab9cc2a8239d04a1703984

Observation 8bcf8818-a277-4f88-839b-f0f9a3320742 · outbound

This paper cites Latent Video Diffusion Models for High-Fidelity Long Video Generation.

Video World Models with Long-term Spatial Memory Latent Video Diffusion Models for High-Fidelity Long Video Generation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:38.402967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:38.402967Z digest=sha256:c727adab971aedfa0891773ac7cb9f3f5681d109afd13e079b4e82a70642f005

Observation 91ae7894-17fd-4800-9cbe-1d0b00dd160e · outbound

This paper cites StreamingT2V: Consistent, Dynamic, and Extendable Long Video Generation from Text.

Video World Models with Long-term Spatial Memory StreamingT2V: Consistent, Dynamic, and Extendable Long Video Generation from Text

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:38.472290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:38.472290Z digest=sha256:4301c90a1c8bc59de234cf91b8981450b0375c4f02229496d4513b1dd70b0a58

Observation b55d5107-a595-49ea-91d6-a02d183c69c3 · outbound

This paper cites Denoising diffusion probabilistic models.

Video World Models with Long-term Spatial Memory Denoising diffusion probabilistic models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:38.562690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:38.562690Z digest=sha256:73acbfa16c40c28f3e5e8c1681738e7c6e58980205d44a98c0722b9761bcbe4a

Observation 6c6b97e6-8af1-47e5-a362-1b05b407af06 · outbound

This paper cites Video diffusion models.

Video World Models with Long-term Spatial Memory Video diffusion models

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:48.558564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:28:38.638729Z digest=sha256:a668b0a1619d40c74f4ec4abc5479f588a0d61f795df983c7bec41ca99c42459

Observation b1da9053-045e-4dda-9e71-347a4250fc7a · outbound

This paper cites Vbench: Comprehensive benchmark suite for video generative models.

Video World Models with Long-term Spatial Memory Vbench: Comprehensive benchmark suite for video generative models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:38.674751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:38.674751Z digest=sha256:758132487af3f2eb2c97c9a7f8fa578daa09c94effe9d1737b09e5ba23c5d208

Observation bed39db2-a54b-42fa-ab93-bd6cc55a0835 · outbound

This paper cites Pyramidal flow matching for efficient video generative modeling.

Video World Models with Long-term Spatial Memory Pyramidal flow matching for efficient video generative modeling

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:38.764229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:38.764229Z digest=sha256:c3fa0800728cca6bbaaa3a0316d1fba76004e85865c4137bde9683f4d8ce1ab1

Observation 953bf0cc-7db2-4310-afde-33fa65007a9c · outbound

This paper cites Pyramidal flow matching for efficient video generative modeling.

Video World Models with Long-term Spatial Memory Pyramidal flow matching for efficient video generative modeling

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:38.855269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:38.855269Z digest=sha256:ebe9743ac773b2ec4b818363d094fb69dc1cafcb17bc7c7005f675623a4d2003

Observation 8bebd46a-7809-49b1-82df-feac7956dab0 · outbound

This paper cites MiraData: A Large-Scale Video Dataset with Long Durations and Structured Captions.

Video World Models with Long-term Spatial Memory MiraData: A Large-Scale Video Dataset with Long Durations and Structured Captions

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:38.940419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:38.940419Z digest=sha256:d2f19aae7d3f66d3a5bef9ca23e1c15c376950c152cb074f10a1a9d2a5499479

Observation 920a28e6-176f-449b-bcb9-03face9391f9 · outbound

This paper cites Videopoet: A large language model for zero-shot video generation.

Video World Models with Long-term Spatial Memory Videopoet: A large language model for zero-shot video generation

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:48.339249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:28:39.015650Z digest=sha256:bb8439a3f0225e7037017f637e3780de03807b40dc9ad0c0a7430bc22726eeb8

Observation 898988a5-0fba-4b7e-8b41-1dc16d0b15fb · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

Video World Models with Long-term Spatial Memory HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:39.075441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:39.075441Z digest=sha256:e0d5ec208a7fa07cdd92aad766854ba6c1e0692841b25ba08c30bb8b77c49a15

Observation 6f0f827e-766a-425b-91d6-8d6d9ed10ccb · outbound

This paper cites Stochastic Adversarial Video Prediction.

Video World Models with Long-term Spatial Memory Stochastic Adversarial Video Prediction

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:39.157223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:39.157223Z digest=sha256:12baf608ce4a70672d332ff0bb59dc9d1ec9772a73ff9b217d0d27d57f88fffd

Observation f9770c06-45ca-4b84-8ad9-e86cede4e095 · outbound

This paper cites Efficient spatially sparse inference for conditional gans and diffusion models.

Video World Models with Long-term Spatial Memory Efficient spatially sparse inference for conditional gans and diffusion models

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:48.096228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:28:39.239472Z digest=sha256:7e46d19edb4acf1a4fb98422fbd419abf07f17fbf63b41cdf5fda74b581409cc

Observation 26982cc4-6593-4f19-aaad-666465b4e2d3 · outbound

This paper cites Megasam: Accurate, fast and robust structure and motion from casual dynamic videos.

Video World Models with Long-term Spatial Memory Megasam: Accurate, fast and robust structure and motion from casual dynamic videos

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:39.287888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:39.287888Z digest=sha256:4f5b63404ac6d9cbfe7a074c8286f257d3954ab1ca384af4ceeef65c81e8b3d2

Observation fc57380e-c90a-400f-bce8-bfb6d55c932e · outbound

This paper cites Open-Sora Plan: Open-Source Large Video Generation Model.

Video World Models with Long-term Spatial Memory Open-Sora Plan: Open-Source Large Video Generation Model

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:39.369402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:39.369402Z digest=sha256:f8fb4d9c9e7494017651d9091ad1ade4ae53dd3ef0dcf2450cfe1a52effc90ed

Observation 923bbfeb-34d1-416d-8c1d-d0c3d911c1c0 · outbound

This paper cites Flow Matching for Generative Modeling.

Video World Models with Long-term Spatial Memory Flow Matching for Generative Modeling

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:39.425630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:39.425630Z digest=sha256:717942151edd7963ad06e2a4aedad07e604d83c7eb84056d84127eb2473ef794

Observation d98c8544-43ef-4d71-ba2b-244e82ba604d · outbound

This paper cites LinFusion: 1 GPU, 1 Minute, 16K Image.

Video World Models with Long-term Spatial Memory LinFusion: 1 GPU, 1 Minute, 16K Image

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:39.512049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:39.512049Z digest=sha256:dc137c4b3d7a75f42de1019f8d19f0ae5df1bb7312882f521626eca169e647d3

Observation 150d21ee-eda9-4bc3-9bcb-7f202f1ac7d9 · outbound

This paper cites Deep multi-scale video prediction beyond mean square error.

Video World Models with Long-term Spatial Memory Deep multi-scale video prediction beyond mean square error

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:39.609601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:39.609601Z digest=sha256:7d9c85d5bb931edfa1502a8b302ed1f666e1d067bbd85835103b870aeee15c47

Observation 8a0f1c6c-608a-491d-abfe-ccea54cb7c7a · outbound

This paper cites SSM Meets Video Diffusion Models: Efficient Long-Term Video Generation with Structured State Spaces.

Video World Models with Long-term Spatial Memory SSM Meets Video Diffusion Models: Efficient Long-Term Video Generation with Structured State Spaces

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:39.676384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:39.676384Z digest=sha256:5da0922fee2a1c3e16ccffd42372f44a5f31ed218ad28698a585c5f754d64b7d

Observation 650fe828-f1c7-44ca-a01c-9b2a9817dab6 · outbound

This paper cites Genie 2: A large-scale foundation world model.

Video World Models with Long-term Spatial Memory Genie 2: A large-scale foundation world model

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:39.746094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:39.746094Z digest=sha256:26dbeebb52385cf319a4d6e02f9d58c442270152dcdb5f1e69d604ec76ab2680

Observation 42c0bced-6059-4722-85cd-b8f9fff41589 · outbound

This paper cites Genie 2: A large-scale foundation world model.

Video World Models with Long-term Spatial Memory Genie 2: A large-scale foundation world model

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:47.912744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:28:39.818205Z digest=sha256:ae6c2996ebc72a340731df30adc0fb09b1d0491854d5603640febdd80bdd7277

Observation d2b0527d-676e-40e1-bf04-ab271b50f2d4 · outbound

This paper cites State of the art on diffusion models for visual computing.

Video World Models with Long-term Spatial Memory State of the art on diffusion models for visual computing

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:47.725079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:28:39.894471Z digest=sha256:13ec09d9c1ecbfc18e65cde0665ef473a37ee710034aa798a710e682edf7f8d2

Observation bac918dd-2c4f-4780-a159-9a6088c77634 · outbound

This paper cites SkyReels-A1: Expressive Portrait Animation in Video Diffusion Transformers.

Video World Models with Long-term Spatial Memory SkyReels-A1: Expressive Portrait Animation in Video Diffusion Transformers

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:39.944935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:39.944935Z digest=sha256:96ae01270fa2e987c669b3c652f84c451b6c4c65922b6b7a4c986d7739ae8c6d

Observation 1bb67816-0a8d-4b2f-8a41-361d86c94510 · outbound

This paper cites Gen3c: 3d-informed world- consistent video generation with precise camera control.

Video World Models with Long-term Spatial Memory Gen3c: 3d-informed world- consistent video generation with precise camera control

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:47.479521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:28:40.031042Z digest=sha256:b13e243fb2e01f216990208da2c0dcb70a3c99bc68384ab1e1564dc6376140e9

Observation dd58db0f-731a-4d55-8836-1d40c29b5c18 · outbound

This paper cites Structure-from-motion revisited.

Video World Models with Long-term Spatial Memory Structure-from-motion revisited

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:47.270600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:28:40.128107Z digest=sha256:5ccfd37dd03a6560b55db79a35359360aa2de47f6507645d12b2c03904c3ba16

Observation 5640641d-4758-41f7-90b7-c9831286f6e4 · outbound

This paper cites Make-a-video: Text-to-video generation without text-video data.

Video World Models with Long-term Spatial Memory Make-a-video: Text-to-video generation without text-video data

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:40.210493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:40.210493Z digest=sha256:98e1efba6d2a61dfd424634943a6718a565e3afe345f7bed1f6c3950c0da140b

Observation c6b775c6-4cc2-4aae-a5a7-088e433cff21 · outbound

This paper cites Deep unsuper- vised learning using nonequilibrium thermodynamics.

Video World Models with Long-term Spatial Memory Deep unsuper- vised learning using nonequilibrium thermodynamics

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:47.116687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:28:40.280357Z digest=sha256:1813610e0a5917fc4f7e0f6a59d6859ff642ebd06331e2c0afe85a6b92d4b5d0

Observation 6198f0ed-8a9f-4334-8a3e-e5bdff7b889f · outbound

This paper cites History-Guided Video Diffusion.

Video World Models with Long-term Spatial Memory History-Guided Video Diffusion

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:40.360174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:40.360174Z digest=sha256:75502459494ac6e4b1678e9672f720bbcc49bbec9881299646cbd6050cbdaf97

Observation 78ffb01a-386c-4b56-89f4-ecfe978b6a74 · outbound

This paper cites Score-based generative modeling through stochastic differential equations.

Video World Models with Long-term Spatial Memory Score-based generative modeling through stochastic differential equations

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:40.445905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:40.445905Z digest=sha256:114460fcae355937c82eccf24cd67e9f75b62f4020c4142d130f0f046d147edf

Observation fb396d7f-bd13-4e2b-8828-13e7b8d3babe · outbound

This paper cites Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation.

Video World Models with Long-term Spatial Memory Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:40.535351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:40.535351Z digest=sha256:346ec3101491a5d1441cdc2af8aa971e4ffd6d111a755070aba24d9fe8ca6ddb

Observation 33048bd2-577b-4e5a-97b2-1f0dc8a1447e · outbound

This paper cites VidTok: A Versatile and Open-Source Video Tokenizer.

Video World Models with Long-term Spatial Memory VidTok: A Versatile and Open-Source Video Tokenizer

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:40.611770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:40.611770Z digest=sha256:9ef864ed084c41adc9f1244a02a733c06a4e2677a12a395ef8a643f5dd23b38a

Observation b1705ef8-544a-44c0-acaf-b86f0eeccbce · outbound

This paper cites Tulyakov, Ming-Yu Liu, Xiaodong Yang, and Jan Kautz.

Video World Models with Long-term Spatial Memory Tulyakov, Ming-Yu Liu, Xiaodong Yang, and Jan Kautz

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:46.994720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:28:40.680078Z digest=sha256:3cdd349e070c5daa6fdfe795cc4b15e040ffdd0b90fb0dd6d3e14e7b02071e23

Observation 4ba40cbd-d454-4f05-88da-2bc89c7e760e · outbound

This paper cites Diffusion Models Are Real-Time Game Engines.

Video World Models with Long-term Spatial Memory Diffusion Models Are Real-Time Game Engines

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:40.743811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:40.743811Z digest=sha256:333a2b4a147affe84b22443e528269ac45c0c3c4d7a35c06d240b31ecfd832ec

Observation b608fb8a-9e4b-46b7-93ca-4c742eefa483 · outbound

This paper cites Phenaki: Variable length video generation from open domain textual descriptions.

Video World Models with Long-term Spatial Memory Phenaki: Variable length video generation from open domain textual descriptions

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:46.824742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:28:40.806367Z digest=sha256:5eb55dcf446dc51bb9d2ba70f5de7cf54f996bd0abdd12d85ddb3631736999ca

Observation efe980eb-a7dc-4451-8b91-b9a084c35640 · outbound

This paper cites Generating the future with adversarial transformers.2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , pages 2992–3000,.

Video World Models with Long-term Spatial Memory Generating the future with adversarial transformers.2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , pages 2992–3000,

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:46.600190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:28:40.943015Z digest=sha256:abb5e85dc36c7e45b3799e2b0ad05353ffb7a3cf32f7e08b63452bc7945e3936

Observation 9756e0d0-e1ce-43e5-885d-66139b2525c0 · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

Video World Models with Long-term Spatial Memory Wan: Open and Advanced Large-Scale Video Generative Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:41.155178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:41.155178Z digest=sha256:a9c3b82e8e1c33f34c8904f5a7b2fa2abc1dc39103ecdc6a3a0681669cfb0cbb

Observation 0c076b86-9971-4220-859b-86de36be2480 · outbound

This paper cites LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity.

Video World Models with Long-term Spatial Memory LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:41.279115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:41.279115Z digest=sha256:8d9c39b397c99e7a7c318bd5f9ffa01f4cc601b44ff12629e99d480fb40cfa05

Observation 35dfcb79-a2ca-4277-a81e-bee7481ebd00 · outbound

This paper cites Vggt: Visual geometry grounded transformer.

Video World Models with Long-term Spatial Memory Vggt: Visual geometry grounded transformer

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:41.344821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:41.344821Z digest=sha256:b8162b682c395001864e9656047c643254fd3ea1cf67d53d9bfa3dcc65803a11

Observation c79fc6c5-3e2a-453a-b376-4d5fc398147b · outbound

This paper cites Efros, and Angjoo Kanazawa.

Video World Models with Long-term Spatial Memory Efros, and Angjoo Kanazawa

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:46.242640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:28:41.445162Z digest=sha256:c5fa37fe586a699389e1d67d9e45fd41d8631922d9848a102fee1bce937b907c

Observation 82dba669-720b-4ccc-ad3f-5a561fbd876a · outbound

This paper cites Dust3r: Geometric 3d vision made easy.

Video World Models with Long-term Spatial Memory Dust3r: Geometric 3d vision made easy

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:45.998420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:28:41.541169Z digest=sha256:ab7c6d9be992108c5f63c7b74f20a2e80c74b88e4ace5d50d3e163db7a3d1715

Observation ecf31cbd-af7b-4861-b57c-eee64139f826 · outbound

This paper cites Loong: Generating Minute-level Long Videos with Autoregressive Language Models.

Video World Models with Long-term Spatial Memory Loong: Generating Minute-level Long Videos with Autoregressive Language Models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:41.605178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:41.605178Z digest=sha256:541481d40e2cd982812ebd9c8eab91e1737159341fbe0040bf6a2f76c8d663f9

Observation 346b2173-a322-4bcf-ba62-ea411bf4285d · outbound

This paper cites Motionctrl: A unified and flexible motion controller for video generation.

Video World Models with Long-term Spatial Memory Motionctrl: A unified and flexible motion controller for video generation

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:45.760855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:28:41.683046Z digest=sha256:513745733d82faa5f590e968a653438dd883771c0d4201cebc538edf919b097f

Observation a9c3a7dd-7d1b-4b93-b67a-a1951d8c7b08 · outbound

This paper cites Art-v: Auto-regressive text-to-video generation with diffusion models.

Video World Models with Long-term Spatial Memory Art-v: Auto-regressive text-to-video generation with diffusion models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:41.761985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:41.761985Z digest=sha256:3def47b111ea3136f828f175434ccb83399e4c07576872d2ae4e65ddb39352f7

Observation e9ca8eb2-5088-473b-be6c-5bb893b5dbe2 · outbound

This paper cites Day- dreamer: World models for physical robot learning.

Video World Models with Long-term Spatial Memory Day- dreamer: World models for physical robot learning

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:41.851167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:41.851167Z digest=sha256:c60e143bd2fc487052e139f893b67bc472271c1620d8685b781f8caf2ddcf5b3

Observation 77f13cd1-39e1-457b-9fe4-b9b73e9c5642 · outbound

This paper cites SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformers.

Video World Models with Long-term Spatial Memory SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformers

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:41.924378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:41.924378Z digest=sha256:fe4da8f500bd5d4cb8266c5f96fe370712432ad98733bc914956c9b104cedbbd

Observation 578cb817-fdf2-44c5-97b2-5af84d59e2e8 · outbound

This paper cites VideoGPT: Video Generation using VQ-VAE and Transformers.

Video World Models with Long-term Spatial Memory VideoGPT: Video Generation using VQ-VAE and Transformers

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:42.033727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:42.033727Z digest=sha256:1dde048d73f2facf92046cf9fbbb34eba518863bc78d61496bae53f9e3c84c23

Observation a576982a-8a90-4d48-8c81-f064dfeaffd5 · outbound

This paper cites Qwen2.5 Technical Report.

Video World Models with Long-term Spatial Memory Qwen2.5 Technical Report

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:42.170961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:42.170961Z digest=sha256:4c252ce1b5bf84b2c4130393bd778f12c2be8209b0f5a2210b0fbb7cd561091e

Observation fa6f67a1-71d8-419d-bf7c-402827964e8f · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

Video World Models with Long-term Spatial Memory CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:42.341717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:42.341717Z digest=sha256:c0020fd9345a533266c5e92255d968bc8cae3ad3976c8ef2fd1e3c9bf03076b2

Observation ec7b3033-cde7-4ca8-a057-45b06dcf9570 · outbound

This paper cites From slow bidirectional to fast autoregressive video diffusion models.

Video World Models with Long-term Spatial Memory From slow bidirectional to fast autoregressive video diffusion models

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:42.440172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:42.440172Z digest=sha256:d44d3f1660834c431466876d1a5ccb9e03cfe0ab2de12dcb26b8795304bb0848

Observation 541332ae-a3ed-4498-8c20-6ed833bca460 · outbound

This paper cites Gamefactory: Creating new games with generative interactive videos.

Video World Models with Long-term Spatial Memory Gamefactory: Creating new games with generative interactive videos

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:42.536739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:42.536739Z digest=sha256:5c4962ad7ba309197b09db9976d25e44d0397c2a1303f41911c3a380ab9804f4

Observation c231cf10-62dc-4d3c-b1c2-8289b9868f84 · outbound

This paper cites TrajectoryCrafter: Redirecting Camera Trajectory for Monocular Videos via Diffusion Models.

Video World Models with Long-term Spatial Memory TrajectoryCrafter: Redirecting Camera Trajectory for Monocular Videos via Diffusion Models

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:42.630970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:42.630970Z digest=sha256:19c4d03b447f0649689c7643bdde467729a25dda1f15f5cbbf28817e51fc191f

Observation 8f70431b-57b1-48e1-93e2-85f0803f82a5 · outbound

This paper cites ViewCrafter: Taming Video Diffusion Models for High-fidelity Novel View Synthesis.

Video World Models with Long-term Spatial Memory ViewCrafter: Taming Video Diffusion Models for High-fidelity Novel View Synthesis

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:42.743161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:42.743161Z digest=sha256:9445f31ee52dee01e89c1195690b7fabfec722a92d5a934c275b33392513edf3

Observation 2412f064-393b-43f9-a086-20f1648b7850 · outbound

This paper cites 3dmatch: Learning local geometric descriptors from rgb-d reconstructions.

Video World Models with Long-term Spatial Memory 3dmatch: Learning local geometric descriptors from rgb-d reconstructions

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:45.494065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:28:42.848662Z digest=sha256:91d8f380f93e10864019d6b19d5eb72abb4cd010c59d4f821d0c70289207c989

Observation 882a93ac-51f8-4de7-8ee8-80fc75ec777e · outbound

This paper cites ReCapture: Generative Video Camera Controls for User-Provided Videos using Masked Video Fine-Tuning.

Video World Models with Long-term Spatial Memory ReCapture: Generative Video Camera Controls for User-Provided Videos using Masked Video Fine-Tuning

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:42.973994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:42.973994Z digest=sha256:8d64f460c4e43f93031bda65d133bcd10482e928881e27e03a398aa81d50bb4a

Observation 475e9458-5a6d-4f50-ae53-2bf000095282 · outbound

This paper cites MonST3r: A simple approach for estimating geometry in the presence of motion.

Video World Models with Long-term Spatial Memory MonST3r: A simple approach for estimating geometry in the presence of motion

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:45.301963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:28:43.082267Z digest=sha256:d260ddaa6ce8529cf2ec5bf1aecf3ff664055b9463600b43efe0d01ea52cf79b

Observation 7bb73879-df32-4fe5-9870-8cd0aeacf65c · outbound

This paper cites Packing input frame context in next-frame prediction models for video generation.

Video World Models with Long-term Spatial Memory Packing input frame context in next-frame prediction models for video generation

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:43.164649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:43.164649Z digest=sha256:8d98a8c5bcddb5b06920cd10d8e26dc705aac4cc8618fb68a75bd06cf7de5b10

Observation e9b8db33-8f4a-402d-a8ed-cd749b9b1cbc · outbound

This paper cites Adding conditional control to text-to-image diffusion models.

Video World Models with Long-term Spatial Memory Adding conditional control to text-to-image diffusion models

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:45.117681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:28:43.283556Z digest=sha256:94edad1badad4432386b9f8d93ae1610054f8b1ce6aabf62163a266e7e066e3a

Observation a3442e55-5134-470b-b5d3-8a41f3b9ab5a · outbound

This paper cites Flare: Feed-forward geometry, appearance and camera estimation from uncalibrated sparse views.

Video World Models with Long-term Spatial Memory Flare: Feed-forward geometry, appearance and camera estimation from uncalibrated sparse views

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:44.920227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:28:43.364916Z digest=sha256:c14ee91b33d0116599d33cd05c14b8a834a2c4cd1f1652246bbf2ad7f2fcd08a

Observation 4c8090fb-4167-469e-8880-e49ae45c6563 · outbound

This paper cites Extdm: Distribu- tion extrapolation diffusion model for video prediction.

Video World Models with Long-term Spatial Memory Extdm: Distribu- tion extrapolation diffusion model for video prediction

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:44.736677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:28:43.465659Z digest=sha256:8bfe4b5d7d53ce662fe0be3f065856bf2e2761ba8568d746ce4847a659e7dc72

Observation bfa0114f-2ef5-44e8-b93d-d320438d1c59 · outbound

This paper cites Open-Sora: Democratizing Efficient Video Production for All.

Video World Models with Long-term Spatial Memory Open-Sora: Democratizing Efficient Video Production for All

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:43.616440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:43.616440Z digest=sha256:d8d4b7443c47fc00b8e0899f79627389c00ffa0c18c7d6e7c40e48bcb7e2e528

Observation b0e7afd5-b60e-4514-875c-0105f89f858f · outbound

This paper cites Is sora a world simulator? a comprehensive survey on general world models and beyond.

Video World Models with Long-term Spatial Memory Is sora a world simulator? a comprehensive survey on general world models and beyond

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:43.713731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:43.713731Z digest=sha256:0bde45bb960d76a7daaae7c71a3c3d884987236bc084350b6d9212a7c84ca350

Observation cc429729-fe76-41ad-ae3b-8da671383f04 · outbound

This paper cites an unresolved cited work.

Video World Models with Long-term Spatial Memory Unresolved cited work

Reference 2017

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:28:46.414049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:28:41.060766Z digest=sha256:0a233a1d918037f1117196fa17faefec3826e927f195f107990c332a15b9f012

Pith citing papers

Observation ee0a90ed-2aea-40b7-8c77-4888a76cc81e · inbound

Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World Modeling cites this paper.

Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World Modeling Video World Models with Long-term Spatial Memory

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-05-19T05:17:06.641553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-19T05:13:28.767788Z digest=sha256:2f14a9ac1f693d00b9df5e5109bbfc4384efdd1fe9b983fff23099beb8ce2b7d

Observation 636d5055-d7e8-4139-831c-824b9ece7313 · inbound

HunyuanWorld 1.0: Generating Immersive, Explorable, and Interactive 3D Worlds from Words or Pixels cites this paper.

HunyuanWorld 1.0: Generating Immersive, Explorable, and Interactive 3D Worlds from Words or Pixels Video World Models with Long-term Spatial Memory

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T12:26:02.233410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:26:02.233410Z digest=sha256:255cf885d01a280c43ad98b841057677c21665826e8b155c130fc9caab6574b3

Observation f8fc7f98-e26c-420f-a05c-aaef4d954313 · inbound

A Comprehensive Survey on World Models for Embodied AI cites this paper.

A Comprehensive Survey on World Models for Embodied AI Video World Models with Long-term Spatial Memory

Reference 199

Resolution
unresolved
no resolver link, observed 2026-08-04T09:12:52.398588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:12:52.398588Z digest=sha256:a395df5c69daf962d36b9b38849f75438a11d41415d6f3eeaad8fe354db57d57

Observation 37b91f7c-f361-43ea-a426-47bd596a521e · inbound

Vision-Language Memory for Spatial Reasoning cites this paper.

Vision-Language Memory for Spatial Reasoning Video World Models with Long-term Spatial Memory

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-03T20:15:35.555180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:15:35.555180Z digest=sha256:a3724a030aca6110d962c7514f29ed01f0d2ffd35c462ed2a6bf12f6ccba51a5

Observation b8eccc32-8483-4338-8c3c-db7b1b221618 · inbound

CustomX: Unified Character, Action, and Scene Customization in Video World Models cites this paper.

CustomX: Unified Character, Action, and Scene Customization in Video World Models Video World Models with Long-term Spatial Memory

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-03T15:28:59.391766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T15:28:59.391766Z digest=sha256:4d434abf747f77f71ff174cebad3b551a8d8af77b5909e98f20b5dd77471942d

Observation 087dab1d-de5d-4e2e-aded-3bed54e9764c · inbound

Flow Equivariant World Models: Memory for Partially Observed Dynamic Environments cites this paper.

Flow Equivariant World Models: Memory for Partially Observed Dynamic Environments Video World Models with Long-term Spatial Memory

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-03T13:00:58.581479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T13:00:58.581479Z digest=sha256:67f125d4436842e3b846e81880a1f2815f171e665d1c71622208b754724acc90

Observation c3b47843-bc56-4013-a2d5-3b793ce62624 · inbound

UCM: Unified Modeling of Camera Control and Memory with Time-aware Positional Encoding Warping for World Models cites this paper.

UCM: Unified Modeling of Camera Control and Memory with Time-aware Positional Encoding Warping for World Models Video World Models with Long-term Spatial Memory

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-02T20:36:13.618901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:36:13.618901Z digest=sha256:6ed06084f164b4c5d803ddc86315c8414183ef79052a3ff075d5edb85618cca8

Observation 86818377-1329-423c-a336-62f3014d30fd · inbound

ShareVerse: Multi-Agent Consistent Video Generation for Shared World Modeling cites this paper.

ShareVerse: Multi-Agent Consistent Video Generation for Shared World Modeling Video World Models with Long-term Spatial Memory

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T19:20:43.802270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:20:43.802270Z digest=sha256:46076deaa5c4ede0bdb037ae316a21cd08625bf249b013d2ccdb03e75e890d5c

Observation b17f6513-766a-44e5-b364-811e2669f34e · inbound

InSpatio-WorldFM: An Open-Source Real-Time Generative Frame Model cites this paper.

InSpatio-WorldFM: An Open-Source Real-Time Generative Frame Model Video World Models with Long-term Spatial Memory

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:09:59.889971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-15T12:07:17.341611Z digest=sha256:a139a2ec171fc844d5841b3d95e0e3d6d04a486453e86db3858e3130acf181df

Observation db841f41-b426-41cd-9d5c-fd3236bc35a8 · inbound

Evolution of Video Generative Foundations cites this paper.

Evolution of Video Generative Foundations Video World Models with Long-term Spatial Memory

Reference 288

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:05:51.128062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T18:41:38.616611Z digest=sha256:fc594b8a224e7d876d387aabc66db99da5f7ca919a24024e7cbe793a3e53a2f6

Observation b53b137b-73b7-49be-80bd-a27e65817c26 · inbound

Rein3D: Reinforced 3D Indoor Scene Generation with Panoramic Video Diffusion Models cites this paper.

Rein3D: Reinforced 3D Indoor Scene Generation with Panoramic Video Diffusion Models Video World Models with Long-term Spatial Memory

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T10:16:05.837003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T15:34:15.916685Z digest=sha256:c057fa00bf81b0da5465fff6035604107195e8445d6138bc28b23e07b2bff5df

Observation 293929ef-d807-4f29-8824-378a0acf7c45 · inbound

Lyra 2.0: Explorable Generative 3D Worlds cites this paper.

Lyra 2.0: Explorable Generative 3D Worlds Video World Models with Long-term Spatial Memory

Reference 117

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:25:59.826682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T15:30:57.712342Z digest=sha256:2b5c9e7141a90227ba2a0b10a86730ee7c091f75a79ca6085ef17415d0d68009

Observation 9f0559ff-fb52-45f3-b3ff-ccd4b89aff62 · inbound

From Synchrony to Sequence: Exo-to-Ego Generation via Interpolation cites this paper.

From Synchrony to Sequence: Exo-to-Ego Generation via Interpolation Video World Models with Long-term Spatial Memory

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T13:25:26.109906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T13:23:35.738229Z digest=sha256:4c3d468523393ed86f9f28b1e0014a77ac6822d9087861f73705722a287548d7

Observation 2ab0e1d6-3ab7-4024-aa86-8a079fda0a66 · inbound

From Synchrony to Sequence: Exo-to-Ego Generation via Interpolation cites this paper.

From Synchrony to Sequence: Exo-to-Ego Generation via Interpolation Video World Models with Long-term Spatial Memory

Reference 48

Resolution
unresolved
no resolver link, observed 2026-07-12T20:34:35.048469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T20:34:35.048469Z digest=sha256:779c360b77a371d578270949dd938404aedb99d602c56ae6cc1156564839ad48

Observation 7a7fbba3-4c47-49af-93b7-033e624c2f1e · inbound

MultiWorld: Scalable Multi-Agent Multi-View Video World Models cites this paper.

MultiWorld: Scalable Multi-Agent Multi-View Video World Models Video World Models with Long-term Spatial Memory

Reference 54

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T10:09:08.116213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T05:06:11.514186Z digest=sha256:331a5f04a99c272d1b1fb5f051be9cfb358667d0fd4542d8f74a26051d1aa4f8

Observation 7e3da4bb-7bcb-4ef9-a888-b499847b80d1 · inbound

CityRAG: Stepping Into a City via Spatially-Grounded Video Generation cites this paper.

CityRAG: Stepping Into a City via Spatially-Grounded Video Generation Video World Models with Long-term Spatial Memory

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:11:25.889379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T02:10:54.935053Z digest=sha256:d897697c2d93934a563d99c0c4feb9decb30789802f357dd8e9d51292c0d1fe6

Observation a9c1c1d7-7b99-46ac-89b8-1f6c545eb023 · inbound

World-R1: Reinforcing 3D Constraints for Text-to-Video Generation cites this paper.

World-R1: Reinforcing 3D Constraints for Text-to-Video Generation Video World Models with Long-term Spatial Memory

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:46:38.112040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-08T04:26:58.408225Z digest=sha256:df2a860a113422374b9ce0d57eae7d3fb1eb5c0b34f2d29a998f8397974eada0

Observation 5ae0bca7-9524-4a3b-b269-0d929b413b74 · inbound

World-R1: Reinforcing 3D Constraints for Text-to-Video Generation cites this paper.

World-R1: Reinforcing 3D Constraints for Text-to-Video Generation Video World Models with Long-term Spatial Memory

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:49:53.597194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-21T08:47:47.264480Z digest=sha256:d570da9a1a2c044145f9377edd4e31e80123bbf3c15666c37c11282adf9649b1

Observation 0b4c3f49-2130-4177-8b59-5145aa87eea8 · inbound

World-R1: Reinforcing 3D Constraints for Text-to-Video Generation cites this paper.

World-R1: Reinforcing 3D Constraints for Text-to-Video Generation Video World Models with Long-term Spatial Memory

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-25T06:45:25.816518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-25T06:44:23.038517Z digest=sha256:f2e0ac8adee8370edf2e7b75b42ba43383fe786914d06db383d2b278d523e2b0

Observation 497e5e4d-9604-4574-a84b-e92d64a81062 · inbound

Latent State Design for World Models under Sufficiency Constraints cites this paper.

Latent State Design for World Models under Sufficiency Constraints Video World Models with Long-term Spatial Memory

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:36:04.520976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T15:55:31.825583Z digest=sha256:7027b2925e79be06b7747ea23fac46ffd06d931f182055a66737d3eb4dfb6ba1

Observation caabfd76-5e59-474f-88be-df2628208b22 · inbound

SWIFT: Prompt-Adaptive Memory for Efficient Interactive Long Video Generation cites this paper.

SWIFT: Prompt-Adaptive Memory for Efficient Interactive Long Video Generation Video World Models with Long-term Spatial Memory

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:46:30.406350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-12T04:57:38.215348Z digest=sha256:d31a34dd33d6928e63c0b7d1f57426a1a56963c90ea7eadfb8dd3be2aeb9b7ba

Observation 0ea0c8ec-7d31-4e5c-9199-d47fd242540a · inbound

Warp-as-History: Generalizable Camera-Controlled Video Generation from One Training Video cites this paper.

Warp-as-History: Generalizable Camera-Controlled Video Generation from One Training Video Video World Models with Long-term Spatial Memory

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-15T03:19:43.849395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-15T03:16:07.765247Z digest=sha256:acde39d35beb82310980fd8d6884a280d3c8e91c1ce8abd622487668e963f325

Observation 80449b1e-5714-40b3-8b97-52e5afc20cf7 · inbound

GeoFlow: Enforcing Implicit Geometric Consistency in Video Generation cites this paper.

GeoFlow: Enforcing Implicit Geometric Consistency in Video Generation Video World Models with Long-term Spatial Memory

Reference 78

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T11:08:13.549665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T11:06:09.367559Z digest=sha256:2d351dfd401829cad046e4d716c6eee91d4ffbd9672ea03020c50b1d422d4113

Observation 3d08f979-f9d9-4b1c-b494-d9bc7c4483ee · inbound

Remember to be Curious: Episodic Context and Persistent Worlds for 3D Exploration cites this paper.

Remember to be Curious: Episodic Context and Persistent Worlds for 3D Exploration Video World Models with Long-term Spatial Memory

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-22T06:46:10.702631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-22T06:45:33.135713Z digest=sha256:9e45cae8393113be00352e39f2c3392a947e7f97076b73f9e2be159faf23a85e

Observation 43529a1f-e41e-4287-b71c-95d7e23b4be1 · inbound

E$^3$C: Video Generation with 3D Environmental Memory and Ego-Exo Human Pose Control cites this paper.

E$^3$C: Video Generation with 3D Environmental Memory and Ego-Exo Human Pose Control Video World Models with Long-term Spatial Memory

Reference 64

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T22:34:02.508995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-29T22:24:38.389294Z digest=sha256:4c5b51a78c73f7cdb87b8197a79f72793b3b2f58fa9f1d17f438542465a0abbc

Observation 52cfe877-7da2-4a7f-9825-16c53dc2f7f2 · inbound

What-If World: A Causal Benchmark for General World Models in Embodied Scenarios cites this paper.

What-If World: A Causal Benchmark for General World Models in Embodied Scenarios Video World Models with Long-term Spatial Memory

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:23:50.422893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-29T18:23:22.987086Z digest=sha256:819d6604d2f75bca478c99ed0d255850eeca796130fa95812f0305fe3c23d8dd

Observation 922e74e5-09ec-44e6-9686-847f40159c8b · inbound

Robust Dreamer: Deviation-Aware Latent Gaussian Memory for Action-Controlled AR Video Generation cites this paper.

Robust Dreamer: Deviation-Aware Latent Gaussian Memory for Action-Controlled AR Video Generation Video World Models with Long-term Spatial Memory

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-07-01T19:16:00.415735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-28T22:56:35.530309Z digest=sha256:28b30cfee6bf30163fb1ea491614873729859ccaa518cb95678c835d5cf1d212

Observation f10097b2-981b-430b-a5bf-b36fce02d089 · inbound

DecMem: Towards Minute-Long Consistent World Generation with Decoupled Memory cites this paper.

DecMem: Towards Minute-Long Consistent World Generation with Decoupled Memory Video World Models with Long-term Spatial Memory

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-06-28T22:42:46.802396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-28T22:38:10.473453Z digest=sha256:fded778c4478ea0d3916c33c389917e542f5d9c402785bd6b1e0eb234206ebf9

Observation 52e1ebb3-cb56-4643-9b9b-a37d3b567250 · inbound

Geometry-Aware Implicit Memory for Video World Models cites this paper.

Geometry-Aware Implicit Memory for Video World Models Video World Models with Long-term Spatial Memory

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:26:16.916728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-28T15:27:55.009337Z digest=sha256:a6c3f761bba6e2d481b364bcce64e1437f5e016938fb05e8d5a1333c9c8527a2

Observation 15b1745a-4e0f-4613-b3e3-3a22ccec0846 · inbound

MetaWorld: Scaling Multi-Agent Video World Model from Single-view Video Data cites this paper.

MetaWorld: Scaling Multi-Agent Video World Model from Single-view Video Data Video World Models with Long-term Spatial Memory

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:56:20.230079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-28T14:52:30.406683Z digest=sha256:9d14b946905302b0d056563275eb3a34f3f81d9eef4a95b894ed8a312397ee39

Observation d94f74c9-ad90-4347-9bc7-469fe39bc835 · inbound

Echo-Memory: A Controlled Study of Memory in Action World Models cites this paper.

Echo-Memory: A Controlled Study of Memory in Action World Models Video World Models with Long-term Spatial Memory

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:57:29.674605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T16:58:37.552036Z digest=sha256:579a21e8bba2167542216bc2f0b2c76330380849623780303cb90fea02ea6b62

Observation 28a335ba-f5e4-4223-a2fb-0c33e2a2d615 · inbound

Latent Spatial Memory for Video World Models cites this paper.

Latent Spatial Memory for Video World Models Video World Models with Long-term Spatial Memory

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:07:30.600792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T16:47:42.761342Z digest=sha256:2a09d1babc9b20cc5dfd73c2b2c45254ecd2cae2aabd54385a27d36699c8e028

Observation 9afb05e6-0393-4130-bfc4-23383c6a2c43 · inbound

Envision4D: Envisioning Visual Futures via Feed-forward 4D Gaussian Splatting for Autonomous Driving cites this paper.

Envision4D: Envisioning Visual Futures via Feed-forward 4D Gaussian Splatting for Autonomous Driving Video World Models with Long-term Spatial Memory

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-07-03T04:37:36.928655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T13:48:15.379724Z digest=sha256:44377d36f383ac22cf5ca3ef52a2a46c23ed685b3e35b7b56add38af1586d7bc

Observation 670b0a12-4994-45a9-80a7-f23cd4ca1f41 · inbound

WorldOlympiad: Can Your World Model Survive a Triathlon? cites this paper.

WorldOlympiad: Can Your World Model Survive a Triathlon? Video World Models with Long-term Spatial Memory

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:47:41.312221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T13:05:26.397711Z digest=sha256:0c1a6c1db0531c65d882b0102d2756c7c5101156574fc554706844652c6284e0

Observation 0e30d2f3-0821-4a88-b09d-7ac439255ac0 · inbound

PermaVid: Consistent Video Generation Across Edits via Disentangled Context Memory cites this paper.

PermaVid: Consistent Video Generation Across Edits via Disentangled Context Memory Video World Models with Long-term Spatial Memory

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-03T17:58:47.771173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T03:23:58.455778Z digest=sha256:0a87014b33d8f3a1166f24775655bb27aa84aae8f61e36d3ce8a16cc66230fc3

Observation 1a30d43d-6110-4c61-a366-58d5fe3a315e · inbound

Directing the World: Fast Autoregressive Video Generation with Compositional Human-Camera Control cites this paper.

Directing the World: Fast Autoregressive Video Generation with Compositional Human-Camera Control Video World Models with Long-term Spatial Memory

Reference 61

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T18:43:51.062094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-29T05:03:56.624536Z digest=sha256:14e1007488e352651d15c8953a24e487c7dc4721a4a787a91194717628893aac

Observation e4e4f89e-2fa9-4b30-8169-d9633376bce8 · inbound

MemLearner: Learning to Query Context memory for Video World Models cites this paper.

MemLearner: Learning to Query Context memory for Video World Models Video World Models with Long-term Spatial Memory

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-07-01T10:25:41.838697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-01T05:30:56.140465Z digest=sha256:ab3c8b5d355073fd81959e5c7701e466d740d8b748a9e2dc95e47c371a42b116

Observation c7b926ee-a5e3-415c-b35a-b6ed1e0edff9 · inbound

Pano2World: End-to-End 3D Generation via Unified Multi-View Sequences cites this paper.

Pano2World: End-to-End 3D Generation via Unified Multi-View Sequences Video World Models with Long-term Spatial Memory

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:56:59.195617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-02T13:52:00.598189Z digest=sha256:81e325cc060f99cbd23adabb017827962b40c92989a7bf0ae067f8e7220124b0

Observation b00225a9-7fa8-4d5b-8236-e0d14d1fa078 · inbound

Pano2World: End-to-End 3D Generation via Unified Multi-View Sequences cites this paper.

Pano2World: End-to-End 3D Generation via Unified Multi-View Sequences Video World Models with Long-term Spatial Memory

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-12T09:15:55.868800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:15:55.868800Z digest=sha256:7e9f77d7abb396e4a7ca4ddcdb8d99243c3a2c3da835ec2edea0930d55dfa56a

Observation 862c0820-c10b-4fef-9f83-a2ae97effe78 · inbound

World from Motion: Generative Dynamic Gaussian Reconstruction from Monocular Video cites this paper.

World from Motion: Generative Dynamic Gaussian Reconstruction from Monocular Video Video World Models with Long-term Spatial Memory

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:26:58.560815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-02T13:21:51.983252Z digest=sha256:64c969fb768a9b861a712403fcebdd1498ca941fa2bef9702812524d68294a49

Observation 0645d784-629a-4ccb-8952-4f9105147608 · inbound

PE-Field 4D: Video Generation Models as Canvas cites this paper.

PE-Field 4D: Video Generation Models as Canvas Video World Models with Long-term Spatial Memory

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T22:41:46.096258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:41:46.096258Z digest=sha256:7789db41dc2283c01fcb7ae9dbbd61edcb0e872ba7da8c6b88767988afacce24

Observation 9e44f665-33e6-4ece-8411-de2300489ce0 · inbound

WorldCycle: Self-Verifiable Reinforcement Learning for Long-Horizon Video World Models cites this paper.

WorldCycle: Self-Verifiable Reinforcement Learning for Long-Horizon Video World Models Video World Models with Long-term Spatial Memory

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T12:52:07.557992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T12:52:07.557992Z digest=sha256:b8d454516f3b1f6d7cf472b972f5a0caf673f90e6888f7d9f3ef11921e6dfea3

Observation 414c18ac-e99d-4152-bbcd-d0ddc054637c · inbound

HERA: Historical Evidence Routing Adapter for Physical Prediction in Latent World Models cites this paper.

HERA: Historical Evidence Routing Adapter for Physical Prediction in Latent World Models Video World Models with Long-term Spatial Memory

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-08T11:36:53.778505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T11:36:53.778505Z digest=sha256:a586c65cff64ea9755aed5930a553d6ab95a40645a638cb2ff6c0265b4ab4385

Observation 45019931-55ad-42dd-9581-576a360adc4c · inbound

Addressable Memory for Video World Models cites this paper.

Addressable Memory for Video World Models Video World Models with Long-term Spatial Memory

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-10T05:06:52.002085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:06:52.002085Z digest=sha256:261e11e34cace9b81b4d6dfe928773e628d2ed2a5d829de773f22e340910cbba