Pith. sign in

Paper Citation Record · LEDGER

Latent Spatial Memory for Video World Models

As of 5 August 2026, this Paper Citation Record lists 66 of 66 outbound references and 1 inbound Pith citation observation for arXiv:2606.09828.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.09828 v1

Coverage vector

measured 66 of 66 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-27T16:47:42.761342Z

measured 67 of 67 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T13:05:26.397711Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-03T05:47:41.335193Z

Reference resolution

66 of 66 outbound references displayed

  • verified exact40
  • verified fuzzy0
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch4

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 990bd036-d01b-49f8-918b-06808d48fd38 · outbound

This paper cites Sora.https://openai.com/sora/, 2024.

Latent Spatial Memory for Video World Models Sora.https://openai.com/sora/, 2024

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-27T16:47:42.761342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:47:42.761342Z digest=sha256:85e9c39cfc3ab33cde42b84dd3727c6c6b434e383079882226b5d6fdf75ba9da

Observation 778b2813-94f9-42ee-81b1-74f4c3492313 · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

Latent Spatial Memory for Video World Models Wan: Open and Advanced Large-Scale Video Generative Models

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-03T01:07:30.622711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:47:42.761342Z digest=sha256:375519f281677714b3f3dde341887bb33a6edb4ed5293a841b3980ddfb28b682

Observation 5188ed02-ec1c-4eb1-aecd-fd8487065e69 · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

Latent Spatial Memory for Video World Models HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-03T01:07:30.624495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:47:42.761342Z digest=sha256:9d715a89e22d8b5f5179f569a11297fd08872efff60b63926a58bf9f84a47247

Observation dc0681e3-b80a-4288-bb08-3af59c70f151 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

Latent Spatial Memory for Video World Models CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-03T01:07:30.569922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:47:42.761342Z digest=sha256:c6db49d680ab55916707338b5013178708ad66c87070206ce519661eb00fa12e

Observation 53df465a-52cc-4936-8dc9-309b0f9a44ff · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

Latent Spatial Memory for Video World Models Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-03T01:07:30.581500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:47:42.761342Z digest=sha256:2ad21ee9ced177762a1cc6785e81fd4fecfcd0d9f44bbd71ee4239dd65e60eeb

Observation f1139cde-5575-4f4e-aead-f2b4f8a88a6f · outbound

This paper cites Genie: Generative interactive environments.

Latent Spatial Memory for Video World Models Genie: Generative interactive environments

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-27T16:47:42.761342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:47:42.761342Z digest=sha256:a2cd39ba7df5d3e165cb8a75abd1588a340b45f96214cbc99f486f91eb391cef

Observation 5669ecf6-66a3-4218-840b-879a8f988b8b · outbound

This paper cites Genie 2: A large-scale foundation world model.

Latent Spatial Memory for Video World Models Genie 2: A large-scale foundation world model

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-27T16:47:42.761342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:47:42.761342Z digest=sha256:3d715b9fd6944d0e8365c2ece129dd4aa01c659e6104b6301f104689de3a372d

Observation c5848669-8dfc-46a3-b493-d337bc04e11b · outbound

This paper cites Diffusion Models Are Real-Time Game Engines.

Latent Spatial Memory for Video World Models Diffusion Models Are Real-Time Game Engines

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-03T01:07:30.564268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:47:42.761342Z digest=sha256:61dd3af6f2f73db80fe8c7446a485ebed59d4d543d168e928b0dfd9264790847

Observation 3a7c5216-bf70-4e9f-aad0-5b8c82ab6864 · outbound

This paper cites Diffusion for world modeling: Visual details matter in atari.Advances in Neural Information Processing Systems, 37:58757–58791, 2024.

Latent Spatial Memory for Video World Models Diffusion for world modeling: Visual details matter in atari.Advances in Neural Information Processing Systems, 37:58757–58791, 2024

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-27T16:47:42.761342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:47:42.761342Z digest=sha256:19fa3d55431b46ca52e71ad852a3dd826ca89a81b579767091c169365df8205e

Observation a84ec768-7a5a-4f31-85a3-3e99ae30bd87 · outbound

This paper cites GameGen-X: Interactive Open-world Game Video Generation.

Latent Spatial Memory for Video World Models GameGen-X: Interactive Open-world Game Video Generation

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:07:30.590453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:47:42.761342Z digest=sha256:5ff1fa42b4c23ba02c5e5321e383894bb7547e53391501238b08090f74eaf725

Observation d584c8e0-0bf8-41af-8251-329fcbdf6b20 · outbound

This paper cites Spatia: Video generation with updatable spatial memory.

Latent Spatial Memory for Video World Models Spatia: Video generation with updatable spatial memory

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-27T16:47:42.761342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:47:42.761342Z digest=sha256:2abdc34bff1b01cbb2be99816e9901e713774c437f3b78b655a890539c3d9230

Observation 9527c192-978c-442c-92ff-b4f96a806543 · outbound

This paper cites Voyager: Long-Range and World-Consistent Video Diffusion for Explorable 3D Scene Generation.

Latent Spatial Memory for Video World Models Voyager: Long-Range and World-Consistent Video Diffusion for Explorable 3D Scene Generation

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:07:30.584646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:47:42.761342Z digest=sha256:8935897aab3f68b9881c246f25270b044284ffa98040ef3c1a6401d43e5a2152

Observation f8b9fd73-8324-4d81-b9ac-8f523d07b4fd · outbound

This paper cites WonderWorld: Interactive 3D Scene Generation from a Single Image.

Latent Spatial Memory for Video World Models WonderWorld: Interactive 3D Scene Generation from a Single Image

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T01:07:30.613347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:47:42.761342Z digest=sha256:55a8bd4775811a6233d541b4b4137fa5c46712f1c8f4742bd1db4158153ef297

Observation 48bb8dd4-3d6e-4232-9e8b-f464e34de3c4 · outbound

This paper cites Wonderjourney: Going from anywhere to everywhere.

Latent Spatial Memory for Video World Models Wonderjourney: Going from anywhere to everywhere

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-27T16:47:42.761342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:47:42.761342Z digest=sha256:ad1c9de34b1b5e82bb37986c9d77927bb4f53d0d8a950c94c0bb9e8808e9b8d2

Observation 1511a3e9-fe03-4caf-b227-f0322b9ad9d8 · outbound

This paper cites LiveWorld: Simulating out-of-sight dynamics in generative video world models.

Latent Spatial Memory for Video World Models LiveWorld: Simulating out-of-sight dynamics in generative video world models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:07:30.581469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:47:42.761342Z digest=sha256:c90704fa7a452a7a37cf92c332c15f33948b182e98b75411ee9776e7867c4aea

Observation 96c59b79-4bf3-47d2-89e9-346d8137dd47 · outbound

This paper cites DriveGen3D: Boosting Feed-Forward Driving Scene Generation with Efficient Video Diffusion.

Latent Spatial Memory for Video World Models DriveGen3D: Boosting Feed-Forward Driving Scene Generation with Efficient Video Diffusion

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-03T01:07:30.617423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:47:42.761342Z digest=sha256:91f3778ae89d1286d2ae146c10f036992490f34d3403a72b046c3b4f4e6c9aad

Observation e5a1b005-f31c-4f2d-9ea5-0caac68328d6 · outbound

This paper cites World-R1: Reinforcing 3D Constraints for Text-to-Video Generation.

Latent Spatial Memory for Video World Models World-R1: Reinforcing 3D Constraints for Text-to-Video Generation

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-07-03T01:07:30.597836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:47:42.761342Z digest=sha256:2505b3e2c30165248751314203b4efa1ecf4e778ba228db9e141001d9d877092

Observation 28a335ba-f5e4-4223-a2fb-0c33e2a2d615 · outbound

This paper cites Video World Models with Long-term Spatial Memory.

Latent Spatial Memory for Video World Models Video World Models with Long-term Spatial Memory

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:07:30.600792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:47:42.761342Z digest=sha256:ce51e94422db318146e6231062965b64a1fae21f30c566d22a4866fdc6e6aaa5

Observation 74f807bd-b84f-47b0-b409-9ad097205e3f · outbound

This paper cites Captain safari: A world engine.

Latent Spatial Memory for Video World Models Captain safari: A world engine

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:07:30.602248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:47:42.761342Z digest=sha256:041aaebbb380fc9a46983eb8c71268f990db7c0435c65436f8e9206958961dfa

Observation a40b37c4-07ca-4a6f-9820-585ec5219d0e · outbound

This paper cites Depth anything 3: Recovering the visual space from any views.International Conference on Learning Representations (ICLR), 2026.

Latent Spatial Memory for Video World Models Depth anything 3: Recovering the visual space from any views.International Conference on Learning Representations (ICLR), 2026

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-27T16:47:42.761342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:47:42.761342Z digest=sha256:06e13cc51f8abc75502268d57bd09d56bcebbc6ba85bd74e2f8ea2d80aceb104

Observation 51f5dbb6-521c-4134-bc84-c74b9f6dbe44 · outbound

This paper cites Feed-Forward 3D Scene Modeling: A Problem-Driven Perspective.

Latent Spatial Memory for Video World Models Feed-Forward 3D Scene Modeling: A Problem-Driven Perspective

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-07-03T01:07:30.592988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:47:42.761342Z digest=sha256:b640af4f2adc05ed4378e16fac928d0e3f67524bca2509c4a14bb8313abdf838

Observation c4abd168-1871-4c9f-8fe5-a7938ee64550 · outbound

This paper cites Zpressor: Bottleneck-aware compression for scalable feed-forward 3dgs.

Latent Spatial Memory for Video World Models Zpressor: Bottleneck-aware compression for scalable feed-forward 3dgs

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-27T16:47:42.761342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:47:42.761342Z digest=sha256:20c517696dbe4f53cfd0e250d9f50fa1178e7f30cc52d54259784bf3c96ed3ef

Observation 398d010d-89df-4676-8c2a-021bcae6b72b · outbound

This paper cites VolSplat: Rethinking Feed-Forward 3D Gaussian Splatting with Voxel-Aligned Prediction.

Latent Spatial Memory for Video World Models VolSplat: Rethinking Feed-Forward 3D Gaussian Splatting with Voxel-Aligned Prediction

Reference 23

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T01:07:30.604666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:47:42.761342Z digest=sha256:bf500ecb80170d194b4f6af60581655f31abbee5e9bb6e965985a78f45062701

Observation f5822075-6c29-4512-8788-2162d5328b31 · outbound

This paper cites TriSplat: Simulation-Ready Feed-Forward 3D Scene Reconstruction.

Latent Spatial Memory for Video World Models TriSplat: Simulation-Ready Feed-Forward 3D Scene Reconstruction

Reference 24

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T01:07:30.616533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:47:42.761342Z digest=sha256:9dcc18351641a4c89e971eeba05a80cd3cb6a49268c8ae2c041aea180eb169e5

Observation 37ab7e0b-6690-4838-8540-9e76a0516d54 · outbound

This paper cites Adding conditional control to text-to-image diffusion models, 2023.

Latent Spatial Memory for Video World Models Adding conditional control to text-to-image diffusion models, 2023

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-27T16:47:42.761342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:47:42.761342Z digest=sha256:105d9082f3cbfa93fa5b1821488093dbf8e0feb1593666ee3d689f9ecc95d06a

Observation 86eed827-9d06-4a7a-9120-82f462e6a728 · outbound

This paper cites SAM 3: Segment Anything with Concepts.

Latent Spatial Memory for Video World Models SAM 3: Segment Anything with Concepts

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-07-03T01:07:30.595055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:47:42.761342Z digest=sha256:d4fbec4215dcc12030e1a78e558999206c834062264a73898db24b1260da7dac

Observation ed9153cc-8143-4d96-b3ff-9757bbe304c9 · outbound

This paper cites Worldscore: A unified evaluation benchmark for world generation.

Latent Spatial Memory for Video World Models Worldscore: A unified evaluation benchmark for world generation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-06-27T16:47:42.761342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:47:42.761342Z digest=sha256:b69cc7704e953f50321e449f9454f3e75e8ca40c7d5c5544138dc4cf33819267

Observation 6bde1f3b-1b51-4638-845c-0b5e94121b4a · outbound

This paper cites Worldscore: A unified evaluation benchmark for world generation.

Latent Spatial Memory for Video World Models Worldscore: A unified evaluation benchmark for world generation

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:07:30.609100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:47:42.761342Z digest=sha256:031985eae79c181b3a6d862bbbef17e85bb6dfa581a9d6e7ec2f0a34e90956d7

Observation fe347f30-0f91-429f-a311-bc8840580baa · outbound

This paper cites Stereo Magnification: Learning View Synthesis using Multiplane Images.

Latent Spatial Memory for Video World Models Stereo Magnification: Learning View Synthesis using Multiplane Images

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-07-03T01:07:30.572676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:47:42.761342Z digest=sha256:2ea24fec1919bc8fae21c52ee4db291501aaa1699125939958ac25a96e07fcb7

Observation bde3d998-f1d2-439e-81b0-b4ea513ff99d · outbound

This paper cites Flow Matching for Generative Modeling.

Latent Spatial Memory for Video World Models Flow Matching for Generative Modeling

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-07-03T01:07:30.567183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:47:42.761342Z digest=sha256:78e15a9f9609597836d5169a4ea98fefb54bd536dc4cf227a917316942f4a2d1

Observation 25654b58-d96e-4671-8652-aa968b561f92 · outbound

This paper cites Scaling rectified flow transformers for high-resolution image synthesis.

Latent Spatial Memory for Video World Models Scaling rectified flow transformers for high-resolution image synthesis

Reference 31

Resolution
unresolved
no resolver link, observed 2026-06-27T16:47:42.761342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:47:42.761342Z digest=sha256:2abbd6d2d574675ddaec5c71e633a1e604bc2c2dadac99604df29365a7b5cb18

Observation 772aef66-2622-46bf-bc52-8472bcf01f00 · outbound

This paper cites AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning.

Latent Spatial Memory for Video World Models AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-07-03T01:07:30.572510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:47:42.761342Z digest=sha256:c7b9659c1e5c0681847067f12fb11fa4fe1ee9c417a15a0dc8317e6e4c6a2f32

Observation 2c620fc0-0f46-4464-9f37-b7aac75ab450 · outbound

This paper cites Videocrafter2: Overcoming data limitations for high-quality video diffusion models.

Latent Spatial Memory for Video World Models Videocrafter2: Overcoming data limitations for high-quality video diffusion models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-27T16:47:42.761342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:47:42.761342Z digest=sha256:38baac3c95559ff4e027ba8fc52fcf704de07812a8b2e6cfaee1e921cc0048c4

Observation d40fa7fe-c78f-4370-b058-7b51833433ea · outbound

This paper cites Easyanimate: A high-performance long video generation method based on transformer architecture.

Latent Spatial Memory for Video World Models Easyanimate: A high-performance long video generation method based on transformer architecture

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:07:30.520717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:47:42.761342Z digest=sha256:4c530417c8f77e9a97737baacfe6394abc2ea1f2089ccde394666994de6357bd

Observation 941cf920-cc9c-4ca0-a84e-c8418149e6d3 · outbound

This paper cites Allegro: Open the Black Box of Commercial-Level Video Generation Model.

Latent Spatial Memory for Video World Models Allegro: Open the Black Box of Commercial-Level Video Generation Model

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:07:30.587853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:47:42.761342Z digest=sha256:5f00595ede22507057518d7a19a1455f5c7fb8950dba5ed96dd83a80cda0710d

Observation 39441c36-f316-4f64-a979-41fe82e66c8a · outbound

This paper cites Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models.

Latent Spatial Memory for Video World Models Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:07:30.541086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:47:42.761342Z digest=sha256:69c2da481fdbfc65b5e0cc2f74a28228348be58a4f3fa15e22df0b1f67c81678

Observation 4160614a-c9df-4a8d-9c4c-0b1bbe7606f6 · outbound

This paper cites LTX-Video: Realtime Video Latent Diffusion.

Latent Spatial Memory for Video World Models LTX-Video: Realtime Video Latent Diffusion

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-07-03T01:07:30.544795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:47:42.761342Z digest=sha256:361ccd48048bc5c8d48e1346b136e467bcf59265bdb64d4017e57d6d759966a3

Observation e699419d-b1f0-45fe-aa22-d3e4ae3e8815 · outbound

This paper cites Scalable diffusion models with transformers.

Latent Spatial Memory for Video World Models Scalable diffusion models with transformers

Reference 38

Resolution
unresolved
no resolver link, observed 2026-06-27T16:47:42.761342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:47:42.761342Z digest=sha256:cb3d185ec954371c007bde309827f773211715524c2c4edf09e4d852fcf0cb7b

Observation 4cfd1af3-9154-4fd6-ad6f-5754a72e1ca7 · outbound

This paper cites Diffusion forcing: Next-token prediction meets full-sequence diffusion.Advances in Neural Information Processing Systems, 37:24081–24125, 2024.

Latent Spatial Memory for Video World Models Diffusion forcing: Next-token prediction meets full-sequence diffusion.Advances in Neural Information Processing Systems, 37:24081–24125, 2024

Reference 39

Resolution
unresolved
no resolver link, observed 2026-06-27T16:47:42.761342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:47:42.761342Z digest=sha256:4350951a9e8daf8f4008cd7a09fb706c477b9a81a8bdf8e1cd6e668006f5eb69

Observation 377a522c-dfb8-4440-87e6-72d18c27457d · outbound

This paper cites StreamingT2V: Consistent, Dynamic, and Extendable Long Video Generation from Text.

Latent Spatial Memory for Video World Models StreamingT2V: Consistent, Dynamic, and Extendable Long Video Generation from Text

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:07:30.542544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:47:42.761342Z digest=sha256:2146866779c7c756ededb13d702e906719fe855ee73e84e88c84044d50525016

Observation fabb33da-7002-44be-8caa-def0bde54458 · outbound

This paper cites Long-Context Autoregressive Video Modeling with Next-Frame Prediction.

Latent Spatial Memory for Video World Models Long-Context Autoregressive Video Modeling with Next-Frame Prediction

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-07-03T01:07:30.539762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:47:42.761342Z digest=sha256:2c37ef43ecfce4d09d5c0c7433cb69a17454c5d08826aa29e17e0d7f7c274c99

Observation 187ec094-66b8-4342-93d5-403816e3d79f · outbound

This paper cites SkyReels-V2: Infinite-length Film Generative Model.

Latent Spatial Memory for Video World Models SkyReels-V2: Infinite-length Film Generative Model

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-07-03T01:07:30.518085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:47:42.761342Z digest=sha256:04e539cf01793b3cc946f343d08dff89496942a4a78963b25b820a2fc63bc832

Observation 0e731683-2faf-4f13-b001-80924ee332b3 · outbound

This paper cites Progressive Autoregressive Video Diffusion Models.

Latent Spatial Memory for Video World Models Progressive Autoregressive Video Diffusion Models

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:07:30.543766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:47:42.761342Z digest=sha256:e5e696d5dfcb62a5e0092796dd93acc5cc37f618b53a77349f4f098976b92163

Observation 35027349-0290-4749-8a70-06ceedc5da22 · outbound

This paper cites CameraCtrl: Enabling Camera Control for Text-to-Video Generation.

Latent Spatial Memory for Video World Models CameraCtrl: Enabling Camera Control for Text-to-Video Generation

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-07-03T01:07:30.549908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:47:42.761342Z digest=sha256:cf167ed8ce887ca6a839deb91fb3eb34bf435bb90341a7e0e4f3dfe3a1d07e8c

Observation dc8ed5b7-74df-4ee6-a82c-351fb092dda7 · outbound

This paper cites CameraCtrl II: Dynamic Scene Exploration via Camera-controlled Video Diffusion Models.

Latent Spatial Memory for Video World Models CameraCtrl II: Dynamic Scene Exploration via Camera-controlled Video Diffusion Models

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:07:30.531837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:47:42.761342Z digest=sha256:a053d44b8b3bcafa2a6fe645f56e855b50dcc31e739c7b81c7ab638aec25f638

Observation 49510953-c288-47e1-b79f-ce600acd1d4f · outbound

This paper cites I2VControl-Camera: Precise Video Camera Control with Adjustable Motion Strength.

Latent Spatial Memory for Video World Models I2VControl-Camera: Precise Video Camera Control with Adjustable Motion Strength

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:07:30.504853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:47:42.761342Z digest=sha256:f2a8692443747fc4e5cd35f1865f200e6b7b7ab1a7dbd04fb30136324d7f4c3a

Observation ca2910c8-f94a-405a-a6a1-3a3d187a12b4 · outbound

This paper cites Panflow: Decoupled motion control for panoramic video generation.

Latent Spatial Memory for Video World Models Panflow: Decoupled motion control for panoramic video generation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-06-27T16:47:42.761342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:47:42.761342Z digest=sha256:a81c4c131ef9e5bf557455b7c94a279f59ac7c0cee91cd3a4718282920242430

Observation 6ae2c67d-b19a-4f0a-9fe2-ed851ec43425 · outbound

This paper cites ViewCrafter: Taming Video Diffusion Models for High-fidelity Novel View Synthesis.

Latent Spatial Memory for Video World Models ViewCrafter: Taming Video Diffusion Models for High-fidelity Novel View Synthesis

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-07-03T01:07:30.547536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:47:42.761342Z digest=sha256:0734e3ce298c8105b7270a3842389f0c24a6ca6043ff1e428daa488977a56eeb

Observation 3596fd8e-bc4f-46ed-b64f-1f6c770255ec · outbound

This paper cites Stable Virtual Camera: Generative View Synthesis with Diffusion Models.

Latent Spatial Memory for Video World Models Stable Virtual Camera: Generative View Synthesis with Diffusion Models

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T01:07:30.619192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:47:42.761342Z digest=sha256:0578cf0e55dcbf9164663afe2cd3b255ac8b48e63c4b83c3e9b49c5b194d82a2

Observation 87069452-d0db-4868-afda-6a7dd23a95c0 · outbound

This paper cites Diffusion as Shader: 3D-aware Video Diffusion for Versatile Video Generation Control.

Latent Spatial Memory for Video World Models Diffusion as Shader: 3D-aware Video Diffusion for Versatile Video Generation Control

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:07:30.525430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:47:42.761342Z digest=sha256:eebc648713c856910d39e824b1c45965bef50b0c33087b09875a5835bb13c88d

Observation e817054a-d947-4f19-840f-cf74bb56677e · outbound

This paper cites Gen3c: 3d-informed world- consistent video generation with precise camera control.

Latent Spatial Memory for Video World Models Gen3c: 3d-informed world- consistent video generation with precise camera control

Reference 51

Resolution
unresolved
no resolver link, observed 2026-06-27T16:47:42.761342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:47:42.761342Z digest=sha256:e2470e49634466e871b9ff499612a40f08f63d6ea334baa9b1c9bdcd284d7df9

Observation 1e768aa3-a099-42f7-b544-72c934469877 · outbound

This paper cites OmniCam: Unified Multimodal Video Generation via Camera Control.

Latent Spatial Memory for Video World Models OmniCam: Unified Multimodal Video Generation via Camera Control

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:07:30.522986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:47:42.761342Z digest=sha256:2ffd92aa8dba9fec3d48918bfd10640ce083bc87141ef1f1e3e10dae3deb30e6

Observation fafee4b7-b81e-4000-9a0c-9b359056c179 · outbound

This paper cites Invisible stitch: Gener- ating smooth 3d scenes with depth inpainting.

Latent Spatial Memory for Video World Models Invisible stitch: Gener- ating smooth 3d scenes with depth inpainting

Reference 53

Resolution
unresolved
no resolver link, observed 2026-06-27T16:47:42.761342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:47:42.761342Z digest=sha256:e09f3e2dbed2402b930a8e4538455461f446da9b5b9105712daa72f2b9661890

Observation 61e12937-678e-464e-afcd-2c530b8ceb1b · outbound

This paper cites FlexWorld: Progressively Expanding 3D Scenes for Flexiable-View Synthesis.

Latent Spatial Memory for Video World Models FlexWorld: Progressively Expanding 3D Scenes for Flexiable-View Synthesis

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:07:30.605978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:47:42.761342Z digest=sha256:a99ddd6766266e319c7dcfdbf976d7120b6df5dcd6cd9216dd09d05f36d76b3a

Observation 99edf4a0-a104-4d72-b421-d62804aad845 · outbound

This paper cites Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval.

Latent Spatial Memory for Video World Models Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:07:30.554029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:47:42.761342Z digest=sha256:6899da65fe0d12848ee23d142d776aecc6c03674d89529ae291780ab08398dbd

Observation 9143008a-ce83-41c2-a166-c05eb9be6cba · outbound

This paper cites arXiv preprint arXiv:2504.12369 , year=.

Latent Spatial Memory for Video World Models arXiv preprint arXiv:2504.12369 , year=

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:07:30.558394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:47:42.761342Z digest=sha256:33c38a7c1623b638a81494504b3b532f33c4d5d400f81ec2efacf1305fdac01d

Observation 9abdcd1c-0cdf-44fa-acaa-293c453449d1 · outbound

This paper cites VMem: Consistent Interactive Video Scene Generation with Surfel-Indexed View Memory.

Latent Spatial Memory for Video World Models VMem: Consistent Interactive Video Scene Generation with Surfel-Indexed View Memory

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:07:30.579093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:47:42.761342Z digest=sha256:010f6b0bb7a18c802f6db3bad344ff37496ea924968e28d29a548bb06a056cc0

Observation 72b7fa6e-ac9f-46c7-88a4-1e3697de5f75 · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063, 2024.

Latent Spatial Memory for Video World Models Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063, 2024

Reference 58

Resolution
unresolved
no resolver link, observed 2026-06-27T16:47:42.761342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:47:42.761342Z digest=sha256:ece924c8ec6426bb18595b5bf0be5d3a75d9ab7a6a9d33887b9243096dc83a70

Observation 755c2684-9da0-48fa-85b6-2e309f7730ab · outbound

This paper cites Qwen3 Technical Report.

Latent Spatial Memory for Video World Models Qwen3 Technical Report

Reference 59

Resolution
verified exact
local_arxiv, observed 2026-07-03T01:07:30.611581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:47:42.761342Z digest=sha256:75179cb8a5599914d2fff006df897189b6fa6a460427b871c5587951c79f2bdd

Observation 0c52b5aa-e4d9-4bcb-b748-f4b85bdbb253 · outbound

This paper cites Lora: Low-rank adaptation of large language models.

Latent Spatial Memory for Video World Models Lora: Low-rank adaptation of large language models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-06-27T16:47:42.761342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:47:42.761342Z digest=sha256:05494bd63da044fd7c12c26c0b41708c04a05bc34975152ba3a0c7baa7f92cd7

Observation 3f5ae6b8-2c0f-41fd-a41b-0878963dc1fa · outbound

This paper cites Flashworld: High-quality 3d scene generation within seconds, 2025.

Latent Spatial Memory for Video World Models Flashworld: High-quality 3d scene generation within seconds, 2025

Reference 61

Resolution
unresolved
no resolver link, observed 2026-06-27T16:47:42.761342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:47:42.761342Z digest=sha256:fc789708f5b496c718df6510ee2551cac1212f5520c8e3f7560a7a51c393f7cb

Observation 7e81dedf-a950-40b1-887b-18fc933ec81d · outbound

This paper cites LucidDreamer: Domain-free Generation of 3D Gaussian Splatting Scenes.

Latent Spatial Memory for Video World Models LucidDreamer: Domain-free Generation of 3D Gaussian Splatting Scenes

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:07:30.520332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:47:42.761342Z digest=sha256:fb2083d0ade080a5184c5369613255a2098c3b7d84deb735e7bc709255721400

Observation fdea7b39-f900-4423-9899-f08518edf7b8 · outbound

This paper cites ViPE: Video Pose Engine for 3D Geometric Perception.

Latent Spatial Memory for Video World Models ViPE: Video Pose Engine for 3D Geometric Perception

Reference 63

Resolution
verified exact
local_arxiv, observed 2026-07-03T01:07:30.512664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:47:42.761342Z digest=sha256:dc2808794e0fe1ec96c886b575cdf838c0e492cf36c02dacbb9042229230d0da

Observation ed825fd4-ae90-4e6f-b90c-f023deb7e3f6 · outbound

This paper cites MapAnything: Universal Feed-Forward Metric 3D Reconstruction.

Latent Spatial Memory for Video World Models MapAnything: Universal Feed-Forward Metric 3D Reconstruction

Reference 64

Resolution
verified exact
local_arxiv, observed 2026-07-03T01:07:30.534308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:47:42.761342Z digest=sha256:5c0edcd50091efa419a3ddc4690e3f5dac17f3b40d29d9f089c006817f181e63

Observation 53544648-4499-4378-bab4-66804870607f · outbound

This paper cites UniDepth: Universal monocular metric depth estimation.

Latent Spatial Memory for Video World Models UniDepth: Universal monocular metric depth estimation

Reference 65

Resolution
unresolved
no resolver link, observed 2026-06-27T16:47:42.761342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:47:42.761342Z digest=sha256:75c124c336f374f5d9de7d6fef198ffd7a7537434d7ce036a739a6ca614350bb

Observation 1faff4ad-7a8f-40bf-991e-dafb66ab7587 · outbound

This paper cites hole rate.

Latent Spatial Memory for Video World Models hole rate

Reference 66

Resolution
unresolved
no resolver link, observed 2026-06-27T16:47:42.761342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:47:42.761342Z digest=sha256:b4423c9d8c1c8d28b3445270129f24628b216473f028e64fa7c5097060a1d80c

Pith citing papers

Observation 230f8b24-10d8-44b2-944d-405378e52366 · inbound

WorldOlympiad: Can Your World Model Survive a Triathlon? cites this paper.

WorldOlympiad: Can Your World Model Survive a Triathlon? Latent Spatial Memory for Video World Models

Reference 38

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T05:47:41.336483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T13:05:26.397711Z digest=sha256:97e480315d30a9de379c559a094fd77c3d3f40f3a6e23c2aba63c1258242881d