Pith. sign in

Paper Citation Record · LEDGER

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models

As of 5 August 2026, this Paper Citation Record lists 100 of 101 outbound references and 4 inbound Pith citation observations for arXiv:2606.00793.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.00793 v2

Coverage vector

measured 100 of 101 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-28T19:04:58.672464Z

measured 104 of 104 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T02:52:21.297924Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T07:14:45.272447Z

Reference resolution

100 of 101 outbound references displayed

  • verified exact48
  • verified fuzzy0
  • unresolved50
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c72e8d42-85c0-460b-9c30-e022930d004a · outbound

This paper cites ai, Hansi Teng, Hongyu Jia, Lei Sun, Lingzhi Li, Maolin Li, Mingqiu Tang, Shuai Han, Tianning Zhang, W.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models ai, Hansi Teng, Hongyu Jia, Lei Sun, Lingzhi Li, Maolin Li, Mingqiu Tang, Shuai Han, Tianning Zhang, W

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-28T19:04:58.672464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:b51c726019bf74ae47e184a65a6733c0d114d713a08f1c11cff5b320210dabaf

Observation 54ed3b2c-1a34-47af-878e-67a2193ab28b · outbound

This paper cites World simulation with video foundation models for physical ai, 2025.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models World simulation with video foundation models for physical ai, 2025

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-28T19:04:58.672464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:76018faccc7d877fae8f7b7a3bca6c9905f3fd2c67fce262cb910b2502986193

Observation c8676be0-69a6-4ce5-9dec-85d06c6113e3 · outbound

This paper cites Diffusion for world modeling: Visual details matter in atari.Advancesin Neural Information Processing Systems, 37:58757–58791, 2024.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models Diffusion for world modeling: Visual details matter in atari.Advancesin Neural Information Processing Systems, 37:58757–58791, 2024

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-28T19:04:58.672464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:5f3c42906e79affa1831ccb8760c857c7fc01c135d729560c7fb40b8b903da04

Observation ebc63a4a-24ec-414f-930d-dd2ad5f0ae72 · outbound

This paper cites Physics-aware-videos.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models Physics-aware-videos

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-28T19:04:58.672464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:3a775c5741c3b84d11174c99f6642b752d45b106906dc00d7e9b319da19ae326

Observation e4d959b3-5dfd-44bf-bb4a-7cf134a0a956 · outbound

This paper cites Qwen3-VL Technical Report.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models Qwen3-VL Technical Report

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-06-28T19:22:35.202367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:b524f5d58bafde1761674064dfe1ae1b01ff7e8f5a51f3e0f7c8594a51505633

Observation 12a423cc-2705-4888-8b7a-45329c48af5c · outbound

This paper cites VideoPhy: Evaluating Physical Commonsense for Video Generation.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models VideoPhy: Evaluating Physical Commonsense for Video Generation

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-06-28T19:22:35.204746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:a343f22d60db0b7d5cdab4b53736e74f290aef6f1cb2cef6c2e330986f045112

Observation f9a28301-9eb9-484d-b53d-9d4644c28bc4 · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-06-28T19:22:35.199395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:8d5acab6160848d6778ab749ce1edbb9614de222c91a30f43876dcac9795a4bb

Observation 7b6de40a-3dd3-451b-b082-86da2e4d1892 · outbound

This paper cites Genie: Generative interactive environments.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models Genie: Generative interactive environments

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-28T19:04:58.672464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:147784bb70e4b9ad298ba5c6ffa2974e78de5e815228e66874d3698d08d65df5

Observation 16bd7734-fe63-4d9d-b935-97eea23836b0 · outbound

This paper cites Seed2.0 model card: Towards intelligence frontier for real-world complexity.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models Seed2.0 model card: Towards intelligence frontier for real-world complexity

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-28T19:04:58.672464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:c8bf8c3ec5e9209ba8bae3e6719dd9187695fd13856043ef5c306f2a205a271c

Observation 2fe1020b-e50e-47ab-b28b-3bb15f26ffa2 · outbound

This paper cites Diffusion forcing: Next-token prediction meets full-sequence diffusion.Advancesin Neural Information Processing Systems, 37:24081–24125, 2024.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models Diffusion forcing: Next-token prediction meets full-sequence diffusion.Advancesin Neural Information Processing Systems, 37:24081–24125, 2024

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-28T19:04:58.672464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:0c7b0ed4902e4508fd3c1b38de532670097ecaa21f668ce3c856dbfa512c27db

Observation cdc21271-a501-4138-98a9-41bd7f9abd19 · outbound

This paper cites Skyreels-v2: Infinite-length film generative model, 2025.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models Skyreels-v2: Infinite-length film generative model, 2025

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-28T19:04:58.672464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:2d3d5890289afc3a20821bdfe7203b46ec9db79cccdb18dc518ac7a7ee4d28ff

Observation 257bb731-5de5-4ced-977a-58182c4a0fb9 · outbound

This paper cites VideoCrafter1: Open Diffusion Models for High-Quality Video Generation.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models VideoCrafter1: Open Diffusion Models for High-Quality Video Generation

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-06-28T19:22:35.146424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:3662ec993ce4a6e51dddaafae383f6aa28a9345885ec97ebb0de5bad9caea0ee

Observation 21af5c0e-d3c4-4ad0-b53c-f0055add531d · outbound

This paper cites VRAG: Learning World Models for Interactive Video Generation.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models VRAG: Learning World Models for Interactive Video Generation

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-06-28T19:22:35.179693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:c578fcc53f50f766f80ce0307c59c7d2702bddd8a4954b998d2bc4d8450f098f

Observation 633ff394-e41e-4034-8bb0-260d87b2c592 · outbound

This paper cites Adversarial Video Generation on Complex Datasets.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models Adversarial Video Generation on Complex Datasets

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T19:22:35.164769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:7d20fa681875c562e563607632e28120fad949f47042b199c188c410d15e1874

Observation 8138f958-5592-4ec6-8927-fe2b4b3f67e5 · outbound

This paper cites One-minute video generation with test-time training.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models One-minute video generation with test-time training

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-28T19:04:58.672464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:be0283d3817b92aff0f4965fb87c846295c6a3f2b79ba4da8b00db98010faf05

Observation 4c2928de-4a7e-45df-abbc-aaae3c2da9e2 · outbound

This paper cites Arcface: Additive angular margin loss for deep face recognition.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models Arcface: Additive angular margin loss for deep face recognition

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-28T19:04:58.672464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:67a5e8a9d8d135d61e499faf4bfce4b93320b90916011c83ea236f89defafa85

Observation f67022ff-e741-463a-a385-57a93cc4e70d · outbound

This paper cites Worldscore: A unified evaluation benchmark for world generation.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models Worldscore: A unified evaluation benchmark for world generation

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:22:35.133872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:f09722be7ea40c9091139d1f0d63a96969336c01820fb38930c3560ca9e53531

Observation 60a787f0-fb25-4c30-a95c-4a994c9b8946 · outbound

This paper cites Vista: A generalizable driving world model with high fidelity and versatile controllability.Advancesin Neural Information Processing Systems, 37:91560–91596, 2024.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models Vista: A generalizable driving world model with high fidelity and versatile controllability.Advancesin Neural Information Processing Systems, 37:91560–91596, 2024

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-28T19:04:58.672464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:302361751763ad4ac30db4a29225ca2ca170ab3cc3535aec53473765cd4330af

Observation 7c3bdd96-ce1f-45ef-a9ea-36ba8eb7a4ef · outbound

This paper cites Veo-3 technical report.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models Veo-3 technical report

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-28T19:04:58.672464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:2a2625cdcb57990b6dd00595a36ecf2b5fb01d18ed73ddc9378e0b245d7f89a5

Observation 38fa18db-11b1-4c5f-aad9-934815fd949d · outbound

This paper cites Long-Context Autoregressive Video Modeling with Next-Frame Prediction.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models Long-Context Autoregressive Video Modeling with Next-Frame Prediction

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-06-28T19:22:35.142836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:7724b58365caed4aa117200d60e1d74c04dc1295913cef37ca8d9b3bd1730dcf

Observation fc290460-3cc0-4b3e-ac19-f4548e4b75d2 · outbound

This paper cites T2VPhysBench: A First-Principles Benchmark for Physical Consistency in Text-to-Video Generation.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models T2VPhysBench: A First-Principles Benchmark for Physical Consistency in Text-to-Video Generation

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:22:35.134374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:48c4b6b6436af3a38744b2105c5d082d38fabbecd88df36d707af37eddde48f5

Observation 09061cf4-72c1-401e-946d-cade91cb4cfe · outbound

This paper cites AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-06-28T19:22:35.191867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:c65659732009f7dd3d5997fd6058c13f52eba31c5eab90d77e0807dd3d3cf69d

Observation 764769eb-340f-4b94-808b-d1b2007f7bfd · outbound

This paper cites Recurrent world models facilitate policy evolution.Advances in neural information processing systems, 31, 2018.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models Recurrent world models facilitate policy evolution.Advances in neural information processing systems, 31, 2018

Reference 23

Resolution
unresolved
no resolver link, observed 2026-06-28T19:04:58.672464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:729110b85113e892f150449781faf3f1a7bb079c0f09018dbf200eed8959b675

Observation 83e6f961-3392-4b7d-91c1-8770fc42e248 · outbound

This paper cites Mastering Diverse Domains through World Models.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models Mastering Diverse Domains through World Models

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-06-28T19:22:35.084706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:c0f8277b7ded8d7f3f93963457b452ac65a9b42ee0ffcbe7b65da880451b8216

Observation 8c24f42d-eb75-46ba-9a42-ae4f45f2ff5e · outbound

This paper cites Video-bench: Human-aligned video generation benchmark.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models Video-bench: Human-aligned video generation benchmark

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-28T19:04:58.672464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:42208f171a48d6c10f73562bc030b880b3b9fe1250da60cba7d3591271960a55

Observation 41af21c2-018f-4cb7-b79b-b9c0443d7f31 · outbound

This paper cites Matrix-game 2.0: An open-source real-time and streaming interactive world model.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models Matrix-game 2.0: An open-source real-time and streaming interactive world model

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-06-28T19:22:35.118293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:5101d606d01f51cb786f428b0eb32a30908df8bd3f302f48d78e73551b5eadec

Observation 0c9c9042-5fb7-478b-b0d7-e3bf63784b5b · outbound

This paper cites Streamingt2v: Consistent, dynamic, and extendable long video generation from text.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models Streamingt2v: Consistent, dynamic, and extendable long video generation from text

Reference 27

Resolution
unresolved
no resolver link, observed 2026-06-28T19:04:58.672464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:a8a174fdd824b3a7954d358cfc5f896f2b8432ec2c53e7e252d587589e43c000

Observation 43d98785-7df8-4d6d-acef-ee284e7debaf · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilibrium.Advancesin neural information processing systems, 30, 2017.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models Gans trained by a two time-scale update rule converge to a local nash equilibrium.Advancesin neural information processing systems, 30, 2017

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-28T19:04:58.672464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:e295307b26e6d2fc1bf862aabe4174e71f4dd439365475f648abaf420a0f824e

Observation b66a9f7b-2952-4be3-b6c6-f1e680fd9a49 · outbound

This paper cites Imagen Video: High Definition Video Generation with Diffusion Models.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models Imagen Video: High Definition Video Generation with Diffusion Models

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-06-28T19:22:35.174312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:becbe39a5d657f8d4c6e55b3c0d7a106be82eb44f3d8fbe03151910fa8ab6bb5

Observation 9a7eea10-f7d4-40bf-9748-d549ad06834b · outbound

This paper cites Video diffusion models.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models Video diffusion models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-06-28T19:04:58.672464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:c96afc5864fd2da810569938577682a16e3d00bc307f1c73f60638c478fbf92a

Observation 9a996e82-2f43-4fa0-897d-e79b1de8188f · outbound

This paper cites Gaia-1: A generative world model for autonomous driving, 2023.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models Gaia-1: A generative world model for autonomous driving, 2023

Reference 31

Resolution
unresolved
no resolver link, observed 2026-06-28T19:04:58.672464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:6603a995ddd73d6a940ceb775547edddca53b80cdff7b5ad6b2847910d813301

Observation e2c9c63c-bce7-49cc-a03a-cd830eefdff6 · outbound

This paper cites Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-06-28T19:22:35.181671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:f0a6530f29fff11a723961b20e7f067c88c91e1ad637345b20c9c8f24c317cdf

Observation 669241d5-c5ba-4586-85f7-5a9d4550914c · outbound

This paper cites VBench: Comprehensive benchmark suite for video generative models.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models VBench: Comprehensive benchmark suite for video generative models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-28T19:04:58.672464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:62c45f4a171c9eff3be901cea92b6ffecc485a8997c2fb12240e5fa51821e74b

Observation 606c04e7-ad21-4544-9efe-fa22c5f4352d · outbound

This paper cites VBench++: Comprehensive and versatile benchmark suite for video generative models.IEEE Transactionson Pattern Analysis and Machine Intelligence, 2025.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models VBench++: Comprehensive and versatile benchmark suite for video generative models.IEEE Transactionson Pattern Analysis and Machine Intelligence, 2025

Reference 34

Resolution
unresolved
no resolver link, observed 2026-06-28T19:04:58.672464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:3d31ac081dc48866a65540a6798495e227eafbfed71244f0b9674e39fc34c16e

Observation f8c2e3b4-5d8c-45bb-8f15-ce4dc0bce61a · outbound

This paper cites Hy-world 1.5: A systematic framework for interactive world modeling with real-time latency and geometric consistency.arXiv preprint, 2025.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models Hy-world 1.5: A systematic framework for interactive world modeling with real-time latency and geometric consistency.arXiv preprint, 2025

Reference 35

Resolution
unresolved
no resolver link, observed 2026-06-28T19:04:58.672464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:ab9a2f30227993d3719ecbe1f93bdb19fa56a29123c5a5a81991228f6efa0ee5

Observation dd945270-0c51-474a-b372-fcfdf49e526c · outbound

This paper cites Openclip.Zenodo, 2021.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models Openclip.Zenodo, 2021

Reference 36

Resolution
unresolved
no resolver link, observed 2026-06-28T19:04:58.672464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:decec6a7647c7e411de718692d05b89f698dc1e9f766939477e95d068308bcf2

Observation 1089c305-d6d6-4d29-9c1d-3ecb1328a8a3 · outbound

This paper cites T2vbench: Benchmarking temporal dynamics for text-to-video generation.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models T2vbench: Benchmarking temporal dynamics for text-to-video generation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-06-28T19:04:58.672464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:ea62e54210b04a68ca2bec867cbcb8e3dcd506feb29575d31cbeb91611d7dbb9

Observation da1cfd43-2976-450c-b00f-6d8ceb25b45d · outbound

This paper cites Memflow: Flowing adaptive memory for consistent and efficient long video narratives.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models Memflow: Flowing adaptive memory for consistent and efficient long video narratives

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:22:35.162161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:88949223261ca65ccb69e21cd89da223afb754c9bc5f4d0837bd818840617f65

Observation 8c90ed9c-56a3-44d3-a344-4c72359bda0e · outbound

This paper cites Pyramidal flow matching for efficient video generative modeling.arXiv preprint arXiv:2410.05954.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models Pyramidal flow matching for efficient video generative modeling.arXiv preprint arXiv:2410.05954

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:22:35.155508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:e919dac4f11f273fbb4a6f2d9440c07021e8b2b737b69775c77d89fe007fb33a

Observation 1e9ddf07-25a9-41ef-a262-7a384c3baaba · outbound

This paper cites Fifo-diffusion: Generating infinite videos from text without training.Advancesin Neural Information Processing Systems, 37:89834–89868, 2024.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models Fifo-diffusion: Generating infinite videos from text without training.Advancesin Neural Information Processing Systems, 37:89834–89868, 2024

Reference 40

Resolution
unresolved
no resolver link, observed 2026-06-28T19:04:58.672464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:0b607e4c740c3e4ebbff3b31c35626cd79bbb4d8b30a17a08bf352ae44dc530d

Observation 91620372-dfdc-4fda-afd3-4bcc5bf479bd · outbound

This paper cites Tanks and temples: Benchmarking large-scale scene reconstruction.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models Tanks and temples: Benchmarking large-scale scene reconstruction

Reference 41

Resolution
unresolved
no resolver link, observed 2026-06-28T19:04:58.672464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:25b8de5f5ad2e263de1300c9766f076ebc0a53f08cb95a2252cafb0cc5777d87

Observation a191da2c-078e-47ee-9fc3-e168d586d647 · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-06-28T19:22:35.169682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:15afba87ecec28de689d94030bd786a21f54bfbacd717ac97d44e0c5e79ab28f

Observation 8de8deec-bb39-4177-84c5-e71291cd3ad6 · outbound

This paper cites Gonzalez, Ion Stoica, Song Han, and Yao Lu.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models Gonzalez, Ion Stoica, Song Han, and Yao Lu

Reference 43

Resolution
unresolved
no resolver link, observed 2026-06-28T19:04:58.672464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:1d7d8dba59c0c5d1377ad62b77a6c39c2402d09b8be0209e63e3628df210ab11

Observation a360f3ab-e8d4-4a1f-beb7-0de58ec56bc2 · outbound

This paper cites OpenHumanVid: A Large-Scale High-Quality Dataset for Enhancing Human-Centric Video Generation.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models OpenHumanVid: A Large-Scale High-Quality Dataset for Enhancing Human-Centric Video Generation

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:22:35.174977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:6ab3ff822345182570497b3c00a72b2351818e5c638d64e78976fc93106493f9

Observation d6f2e002-cc34-418e-98b0-887ca4534632 · outbound

This paper cites Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-06-28T19:22:35.191121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:9406df66a9f0bd07999ed5a7467c9f85572925f1e825d23133bf45433e659e32

Observation a3496d47-9283-499d-ac68-1053101e5008 · outbound

This paper cites Depth Anything 3: Recovering the Visual Space from Any Views.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models Depth Anything 3: Recovering the Visual Space from Any Views

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-06-28T19:22:35.184272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:557ede2642411102e4d2644caaf9354c11e72ea5ec84d33da2e169aaeb8338cb

Observation ba7c79e8-aebc-4bd4-b86c-7a4fae9883c1 · outbound

This paper cites Dl3dv-10k: A large-scale scene dataset for deep learning-based 3d vision.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models Dl3dv-10k: A large-scale scene dataset for deep learning-based 3d vision

Reference 47

Resolution
unresolved
no resolver link, observed 2026-06-28T19:04:58.672464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:babf57f0a4241a886f813eee8d40cf0615a5f616d68b1440c6d23a64742e946f

Observation bb850652-8307-4b74-9fa8-5d2751189378 · outbound

This paper cites Evalcrafter: Benchmarking and evaluating large video generation models.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models Evalcrafter: Benchmarking and evaluating large video generation models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-06-28T19:04:58.672464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:1dd33d90489e9fe375025788c9e9eca21b3658d7fd643d0208c4ef6fcdd50347

Observation 3f3362f9-8e0f-4e4c-a47f-edf54e66a604 · outbound

This paper cites Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-06-28T19:22:35.120276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:967a7be4d0aed7413c9d3ae0bc1c664b6942981200f383a9292c8d4572ed1e01

Observation a17cf3e6-dbba-40c9-88e9-15211af52095 · outbound

This paper cites Fetv: A benchmark for fine-grained evaluation of open-domain text-to-video generation.Advancesin Neural Information Processing Systems, 36:62352–62387, 2023.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models Fetv: A benchmark for fine-grained evaluation of open-domain text-to-video generation.Advancesin Neural Information Processing Systems, 36:62352–62387, 2023

Reference 50

Resolution
unresolved
no resolver link, observed 2026-06-28T19:04:58.672464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:cc686c6b79f32297c88100a1ba63b913fa4a07d9933753237f5b493abcb9782d

Observation 0a054d5f-7fd2-4718-a358-6e0370d23062 · outbound

This paper cites Out of sight, out of mind? evaluating state evolution in video world models, 2026.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models Out of sight, out of mind? evaluating state evolution in video world models, 2026

Reference 51

Resolution
unresolved
no resolver link, observed 2026-06-28T19:04:58.672464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:9dd06a396ebecd58869d51f93aaf2fc62e2d8258b400632df55bb88f2a4c3027

Observation 70bfef34-4972-4621-a4c7-ea80756d8813 · outbound

This paper cites Yume-1.5: A text-controlled interactive world generation model.arXiv preprint arXiv:2512.22096.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models Yume-1.5: A text-controlled interactive world generation model.arXiv preprint arXiv:2512.22096

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:22:35.143652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:7d1f36d888febb894af3ad10367f6484aa4e9119606a06665a9301e4523ea077

Observation de6011d9-542d-4cab-9555-71330b6bde08 · outbound

This paper cites Yume: An Interactive World Generation Model.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models Yume: An Interactive World Generation Model

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:22:35.161751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:4bfc8518cfea734ea4a389d91684025671a1e6d59a6d976f01ce330fc19d39c2

Observation 75d96a76-7321-439e-b54e-68a7a75b22a0 · outbound

This paper cites Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-06-28T19:22:35.182074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:4363ae3e28074bf73bb57060c7224f8351f404d1d5a128f75b0630160b0b5617

Observation d711bbb5-370d-4bdd-a173-52ac616c4363 · outbound

This paper cites Transformers are Sample-Efficient World Models.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models Transformers are Sample-Efficient World Models

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:22:35.159472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:366d79183d21bc7c94c264a73854a37d4cb9533ca007794516e7d3578980e796

Observation 22d4fa73-7cf4-4e23-ba2c-85e9da63c1dd · outbound

This paper cites Sora2.https://openai.com/zh-Hans-CN/index/sora-2/, 2025.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models Sora2.https://openai.com/zh-Hans-CN/index/sora-2/, 2025

Reference 56

Resolution
unresolved
no resolver link, observed 2026-06-28T19:04:58.672464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:88ba5f400391906e3b24cdf8a458f56c734ec125aa04c4e98e1b2344f0c94099

Observation f4daa6cf-2470-4db6-99c2-c99c8a4814ec · outbound

This paper cites an unresolved cited work.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models Unresolved cited work

Reference 57

Resolution
unresolved
no resolver link, observed 2026-06-28T19:04:58.672464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:80eb8a27f2c1acdf4286270991e13e98ef164cc03b999e779e70de71d73f2a45

Observation 9839d058-e001-40c9-a8fd-7a98adc93fbf · outbound

This paper cites Scalable diffusion models with transformers.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models Scalable diffusion models with transformers

Reference 58

Resolution
unresolved
no resolver link, observed 2026-06-28T19:04:58.672464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:bf97bc8d8cf9b2e4878501aae327cb53f83559e5ffb07dda2d701a03b1b36f6c

Observation bf1d6a3e-859c-4098-95dd-8d9b78a35af3 · outbound

This paper cites FreeNoise: Tuning-Free Longer Video Diffusion via Noise Rescheduling.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models FreeNoise: Tuning-Free Longer Video Diffusion via Noise Rescheduling

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:22:35.136808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:749c430ad6c2ed42034e13afb7caae2a12c0ee8a91234bef859bfe631046cb89

Observation 7ba08a8b-85a5-449c-a962-08722445bf96 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models SAM 2: Segment Anything in Images and Videos

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-06-28T19:22:35.156713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:ef1a77b313276f47920e91188b4bb123eb1c55daf1aa47dc444e72dece8c72ad

Observation d35cf775-3a33-4231-870a-7538a3cf4c37 · outbound

This paper cites Improved techniques for training gans.Advancesin neural information processing systems, 29, 2016.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models Improved techniques for training gans.Advancesin neural information processing systems, 29, 2016

Reference 61

Resolution
unresolved
no resolver link, observed 2026-06-28T19:04:58.672464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:09563862f10b9efbfedd0b495aaa543bd721ff506af7fd023a235c42b97bcd54

Observation 62c1d652-7efc-4a67-bd4b-d2758edcffe3 · outbound

This paper cites an unresolved cited work.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models Unresolved cited work

Reference 62

Resolution
unresolved
no resolver link, observed 2026-06-28T19:04:58.672464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:af5ac988e64254eca857104d51814fad1719455cc0c318c3be13b0efb515392e

Observation 4cd2b49b-d99d-4d2b-b307-bec094267bdd · outbound

This paper cites Make-A-Video: Text-to-Video Generation without Text-Video Data.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models Make-A-Video: Text-to-Video Generation without Text-Video Data

Reference 63

Resolution
verified exact
local_arxiv, observed 2026-06-28T19:22:35.115850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:abee1d41900bbe2aaa50913fb7f473734047efd5430587caef985d35553cb742

Observation 1c4a48fa-fa2c-42e8-997c-7dcc4ca23b62 · outbound

This paper cites Matrix-game 3.0: Real-time and streaming interactive world model with long-horizon memory.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models Matrix-game 3.0: Real-time and streaming interactive world model with long-horizon memory

Reference 64

Resolution
unresolved
no resolver link, observed 2026-06-28T19:04:58.672464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:29134116be9f14630537511f3de287efa65787947622616fa674016d320682ea

Observation 00240fd7-57a3-4379-ae0f-fe9cce3d13cc · outbound

This paper cites History-Guided Video Diffusion.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models History-Guided Video Diffusion

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-06-28T19:22:35.125350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:ab98b01000d3c4690c2b88b63782ec54ba0ec07e2bd7bfd9980b1081e52da844

Observation 6d00d965-cec9-4291-b05f-e1d6653e44a2 · outbound

This paper cites T2v-compbench: A comprehensive benchmark for compositional text-to-video generation.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models T2v-compbench: A comprehensive benchmark for compositional text-to-video generation

Reference 66

Resolution
unresolved
no resolver link, observed 2026-06-28T19:04:58.672464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:2227c78e7421b63ccd38c90c9e3fc1d9bb430d7456fccbb52a0f77faef51ac83

Observation 55a91075-c61b-4823-83ab-861c7a1d21cb · outbound

This paper cites Longcat-video technical report, 2025.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models Longcat-video technical report, 2025

Reference 67

Resolution
unresolved
no resolver link, observed 2026-06-28T19:04:58.672464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:50dd6013255778f8f9e4e870d30bd948b5899dd4d9dea8e3297a3677f8d40df2

Observation 63e186c1-0317-4c65-827b-426b9de957c0 · outbound

This paper cites Advancing Open-source World Models.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models Advancing Open-source World Models

Reference 68

Resolution
verified exact
local_arxiv, observed 2026-06-28T19:22:35.172016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:1ef3cee434febc14b48d52cc3cfdb87f1636e7622a3bab202bc028e55b012b93

Observation 587808fa-ae8d-4071-a8d3-aea3e12bef02 · outbound

This paper cites Mocogan: Decomposing motion and content for video generation.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models Mocogan: Decomposing motion and content for video generation

Reference 69

Resolution
unresolved
no resolver link, observed 2026-06-28T19:04:58.672464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:0c4b888488951f3fc5541a4d6c39aa13ed16f86c55f006998ee9455f11e5eedc

Observation f1264607-66fa-4fbc-aa1a-1ed2c39b24a0 · outbound

This paper cites UniFOLM-WMA-0: A world-model-action (wma) framework under UniFOLM family, 2025.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models UniFOLM-WMA-0: A world-model-action (wma) framework under UniFOLM family, 2025

Reference 70

Resolution
unresolved
no resolver link, observed 2026-06-28T19:04:58.672464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:11e9303319b767b8dbfef061400ed7ea9782f34c008d1f50dd446e716d686194

Observation 18df11a5-8698-469f-ae9d-039594313bad · outbound

This paper cites Towards Accurate Generative Models of Video: A New Metric & Challenges.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models Towards Accurate Generative Models of Video: A New Metric & Challenges

Reference 71

Resolution
verified exact
local_arxiv, observed 2026-06-28T19:22:35.176706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:421727b5fb218e17793f7233fdee1bf9c90bc1f9bd9e16606dce7044fa2a8c9f

Observation 080dd672-301c-432c-beff-fb72cba38d53 · outbound

This paper cites How close are world models to the physical world?arXiv preprint arXiv:2501.xxxxx, 2025.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models How close are world models to the physical world?arXiv preprint arXiv:2501.xxxxx, 2025

Reference 72

Resolution
unresolved
no resolver link, observed 2026-06-28T19:04:58.672464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:29799b519a34b40a0cacf5983c2873d96165dd56bb805a0932b2cf3cb88cf5e6

Observation 946167d5-d460-4ba5-ab16-6331f16e8872 · outbound

This paper cites Phenaki: Variable Length Video Generation From Open Domain Textual Description.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models Phenaki: Variable Length Video Generation From Open Domain Textual Description

Reference 73

Resolution
verified exact
local_arxiv, observed 2026-06-28T19:22:35.046561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:ad801271d92b694dba1b0d0097bcf6e605e4bac5405a63e1ad6bab1bc837c727

Observation 8dbc399c-0761-46ee-a89a-69b611b86003 · outbound

This paper cites Generating videos with scene dynamics.Advances in neural information processing systems, 29, 2016.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models Generating videos with scene dynamics.Advances in neural information processing systems, 29, 2016

Reference 74

Resolution
unresolved
no resolver link, observed 2026-06-28T19:04:58.672464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:33606df54eb1194db04d58f1d7956cf7ee89386822d9bedae61a7400ac029e62

Observation 5325b6c0-740c-4b8a-9831-3ebdedc48950 · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models Wan: Open and Advanced Large-Scale Video Generative Models

Reference 75

Resolution
verified exact
local_arxiv, observed 2026-06-28T19:22:35.169842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:1f28419e48050002f5c637437a92ad27c7ece66ec9a9a644d75419a8dae97d7d

Observation 3e2224ea-c61a-490d-88c8-980f2bb8fb53 · outbound

This paper cites Gen-L-Video: Multi-Text to Long Video Generation via Temporal Co-Denoising.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models Gen-L-Video: Multi-Text to Long Video Generation via Temporal Co-Denoising

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:22:35.077252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:4f75041136a24288a2fd33bc618b8207b65a4b44795c3301310e6a717b6238ff

Observation 0c04f754-c2c4-49d2-9fbb-b7f5a5cc91d2 · outbound

This paper cites Spatialvid: A large-scale video dataset with spatial annotations.arXiv preprint arXiv:2509.09676, 2025a.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models Spatialvid: A large-scale video dataset with spatial annotations.arXiv preprint arXiv:2509.09676, 2025a

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:22:35.140390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:5005a138d4b598a3f2300cf134d394d905da21d4bb69420e8f87d117ed22872b

Observation 24f75957-786b-4dfe-93ef-5129a7a5472c · outbound

This paper cites DriveDreamer: Towards Real-world-driven World Models for Autonomous Driving.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models DriveDreamer: Towards Real-world-driven World Models for Autonomous Driving

Reference 78

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:22:35.117848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:05a3845ec25c3b168f1b03fe46a513994da85f2deae42a6ef52236dd03214ec3

Observation 58f0d558-e674-41c1-b0fd-59594131ca80 · outbound

This paper cites Matrix-Game 3.0: Real-Time and Streaming Interactive World Model with Long-Horizon Memory.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models Matrix-Game 3.0: Real-Time and Streaming Interactive World Model with Long-Horizon Memory

Reference 79

Resolution
verified exact
local_arxiv, observed 2026-06-28T19:22:35.065286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:a3c579999748413703c647340ae66fbdcaa7377348ce9f99d29779ea4c8ca1da

Observation 6e2c0d80-0a6c-4434-8fa2-75abfb114abe · outbound

This paper cites GODIVA: Generating Open-DomaIn Videos from nAtural Descriptions.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models GODIVA: Generating Open-DomaIn Videos from nAtural Descriptions

Reference 80

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:22:35.179257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:3b8c967beb3249136b7dc59e333989c00aaf5fb093f95c2ebf188a5b162457df

Observation 219bff66-5671-45f5-b59f-eddbdef9677b · outbound

This paper cites Omni-WorldBench: Towards a comprehensive interaction-centric evaluation for world models.arXiv preprint arXiv:2603.22212, 2026.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models Omni-WorldBench: Towards a comprehensive interaction-centric evaluation for world models.arXiv preprint arXiv:2603.22212, 2026

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:22:35.196990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:236b00f6e5137e4a9f3526570742b31cb9090979b7377880c44ecbf5005c8a77

Observation 593a04dc-d92c-487a-9099-3a322d71a3d9 · outbound

This paper cites Infinite-World: Scaling interactive world models to 1000-frame horizons via pose-free hierarchical memory.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models Infinite-World: Scaling interactive world models to 1000-frame horizons via pose-free hierarchical memory

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:22:35.184436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:1aac70ddeaa35114d2c2875451b7c2c9f2d1432e6a23f00454c16a2cabf62bd3

Observation bac2e25f-0ffe-4d7f-8264-4068164159a7 · outbound

This paper cites arXiv preprint arXiv:2504.12369 , year=.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models arXiv preprint arXiv:2504.12369 , year=

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:22:35.079970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:5995734421e82d2b835bd3259b66bc2747b4b703171d26eab7b73552ad71f512

Observation ba32537c-dce8-4497-acb5-c23187971e83 · outbound

This paper cites WorldMark: A Unified Benchmark Suite for Interactive Video World Models.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models WorldMark: A Unified Benchmark Suite for Interactive Video World Models

Reference 84

Resolution
verified exact
local_arxiv, observed 2026-06-28T19:22:35.145661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:77aa27c83466feae97adb0e25ba1b21420e56d02ae889b41c49ea1c4c809bb91

Observation 03bbca4f-6a72-43a4-bd42-0d171f8181fd · outbound

This paper cites VideoGPT: Video Generation using VQ-VAE and Transformers.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models VideoGPT: Video Generation using VQ-VAE and Transformers

Reference 85

Resolution
verified exact
local_arxiv, observed 2026-06-28T19:22:35.152501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:b199865be0993b39817fa31643f605cc318a1a4ff289ae06b1df11e44231727b

Observation b5e95b31-779b-44b3-8f37-832344776485 · outbound

This paper cites Learning Interactive Real-World Simulators.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models Learning Interactive Real-World Simulators

Reference 86

Resolution
verified exact
local_arxiv, observed 2026-06-28T19:22:35.079542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:f746470ca05e41679483a8b7eb86ef0a676a054830a777131f2d2152a1005bf7

Observation 63c32d4d-a49e-46bc-a541-9aa5de1f03bd · outbound

This paper cites Longlive: Real-time interactive long video generation.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models Longlive: Real-time interactive long video generation

Reference 87

Resolution
unresolved
no resolver link, observed 2026-06-28T19:04:58.672464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:e8d6b836b17bc52da84aea7f78bc7c34cbb1e333e14d12ec5762c7c961e0fbca

Observation 14c646c6-285b-4e6b-8955-8d5684e6a612 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 88

Resolution
verified exact
local_arxiv, observed 2026-06-28T19:22:35.194339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:58780a935fac3d8b00f8e861d1b0915af4747b23ea5a36128c8f4447c06080d6

Observation e33b51ff-059e-45b9-8a74-dd3809420bc0 · outbound

This paper cites Yan: Foundational interactive video generation, 2025.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models Yan: Foundational interactive video generation, 2025

Reference 89

Resolution
unresolved
no resolver link, observed 2026-06-28T19:04:58.672464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:90a460f7eae6d2d5ff29df21189e1b9e3f48249f7ec3e2fec05d18501985d383

Observation 3d2eb81f-c57c-4451-be05-7da8abd9cdfc · outbound

This paper cites Mind: Benchmarking memory consistency and action control in world models, 2026.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models Mind: Benchmarking memory consistency and action control in world models, 2026

Reference 90

Resolution
unresolved
no resolver link, observed 2026-06-28T19:04:58.672464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:ebbbe632852445c85e08b818aa09d12823d636ace2d2452712ebe0469a381514

Observation d47625e7-5a5e-4634-a605-177a2d0549ea · outbound

This paper cites Infinity-RoPE: Action-controllable infinite video generation emerges from autoregressive self- rollout.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models Infinity-RoPE: Action-controllable infinite video generation emerges from autoregressive self- rollout

Reference 91

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:22:35.167229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:4170a926ab147c98511c649d37a0b00ae1963926fe4d61ce6368360b28c913fa

Observation 93333b36-21f1-4840-8716-1ab11677e31c · outbound

This paper cites Nuwa-xl: Diffusion over diffusion for extremely long video generation.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models Nuwa-xl: Diffusion over diffusion for extremely long video generation

Reference 92

Resolution
unresolved
no resolver link, observed 2026-06-28T19:04:58.672464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:580f188698bf684d0483d266a06fde61247b1b8bb9fb424053ed06594289a090

Observation f99d5eb1-efaa-4e26-addf-cef51050d486 · outbound

This paper cites Magvit: Masked generative video transformer.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models Magvit: Masked generative video transformer

Reference 93

Resolution
unresolved
no resolver link, observed 2026-06-28T19:04:58.672464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:23414d1d21940ca4ca7bdd5abf48392c6628223fb7bd816252e8b4f8d7311520

Observation 811b914f-efc3-4dc9-b2d4-a68e2f82d472 · outbound

This paper cites Chronomagic-bench: A benchmark for metamorphic evaluation of text-to-time-lapse video generation.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models Chronomagic-bench: A benchmark for metamorphic evaluation of text-to-time-lapse video generation

Reference 94

Resolution
unresolved
no resolver link, observed 2026-06-28T19:04:58.672464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:af54b2976292221ee6f5642a8a5ad99f04e72e189504d218c7fece65ca9c739b

Observation 3d8681f0-17d0-43f5-af2f-3cb9ea5aac46 · outbound

This paper cites Improved distribution matching distillation for fast image synthesis.Advances in neural information processing systems, 37:47455–47487, 2024a.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models Improved distribution matching distillation for fast image synthesis.Advances in neural information processing systems, 37:47455–47487, 2024a

Reference 95

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T19:22:35.164500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:64ba0379156f23bcedf334393a72ac2bdbdf4a403578521d34a04b1fe03cd121

Observation 6c3134f0-8a4e-45be-a330-8cc504dc63b1 · outbound

This paper cites Packing input frame context in next-frame prediction models for video generation.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models Packing input frame context in next-frame prediction models for video generation

Reference 96

Resolution
unresolved
no resolver link, observed 2026-06-28T19:04:58.672464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:85a972b150d755c689b5508c3dd98f92c90fee98906959ab3ae06d867cf3fbad

Observation 4f87a064-6d35-4cfc-b1b4-9456da40bfc9 · outbound

This paper cites The unreasonable effectiveness of deep features as a perceptual metric.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models The unreasonable effectiveness of deep features as a perceptual metric

Reference 97

Resolution
unresolved
no resolver link, observed 2026-06-28T19:04:58.672464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:77d756f5de2246a8303ac057795c99cabebc243495c3b7e4b99d3d3b3840d69c

Observation f39305a2-4a6b-41ef-ad91-b22585cceff1 · outbound

This paper cites VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 98

Resolution
verified exact
local_arxiv, observed 2026-06-28T19:22:35.113216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:579efda9c873b0ebc2c9d1238f888d83285cb1cb5b9c2ee34478b3838d1389e3

Observation 3e601acf-e865-4f6b-8696-b9a30703d0a6 · outbound

This paper cites LoCoT2V-Bench: Benchmarking Long-Form and Complex Text-to-Video Generation.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models LoCoT2V-Bench: Benchmarking Long-Form and Complex Text-to-Video Generation

Reference 99

Resolution
verified exact
local_arxiv, observed 2026-06-28T19:22:35.151377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:1d23bcd13b64f82c2cd1cc55f40ba0cd006bdf56f2d1296e996220d2419051b2

Observation 4eb2d6c1-03ef-4251-88a7-152c73e06aed · outbound

This paper cites Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video Generation.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video Generation

Reference 100

Resolution
verified exact
local_arxiv, observed 2026-06-28T19:22:35.177383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:a65a0582e011ec96801df87f14a48acf0a2fe30ed7144de12b951a8ad512e126

Pith citing papers

Observation 70403776-aa78-497c-80e5-560eb357732b · inbound

Current World Models Lack a Persistent State Core cites this paper.

Current World Models Lack a Persistent State Core MBench: A Comprehensive Benchmark on Memory Capability for Video World Models

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-07-04T03:49:30.965350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T17:33:41.461245Z digest=sha256:5e86eeb997975f2e049ec96a6e2201849a51188e93e13980829b249612856f88

Observation 5a1d00f9-c940-4488-bcce-40c3f1c4b11b · inbound

A Definition and Roadmap for World Models cites this paper.

A Definition and Roadmap for World Models MBench: A Comprehensive Benchmark on Memory Capability for Video World Models

Reference 279

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T07:14:45.274747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-07-08T07:10:33.826140Z digest=sha256:69bd61dd278b8acd1b657255d80be2abcafdd7947f248c4e76d7c021386cf63d

Observation a357498d-905d-409f-bfbb-9cb9e3313398 · inbound

From Pixels to States: Rethinking Interactive World Models as Game Engines cites this paper.

From Pixels to States: Rethinking Interactive World Models as Game Engines MBench: A Comprehensive Benchmark on Memory Capability for Video World Models

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-02T02:52:21.297924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:52:21.297924Z digest=sha256:b5114fa5f41fe89f4013490fa516358402521fe440a807034b312456eac62c93

Observation 1a1b349b-7eca-424a-afc7-de60fb2d1b2a · inbound

Persistent Computational State: A Session-Centric Runtime for Generative World Models cites this paper.

Persistent Computational State: A Session-Centric Runtime for Generative World Models MBench: A Comprehensive Benchmark on Memory Capability for Video World Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T07:46:24.802500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:46:24.802500Z digest=sha256:e3247244fe167bc84ec0f17560ea65d19bb25ff0f23326a5dac02c489a28d067