Pith. sign in

Paper Citation Record · LEDGER

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

As of 5 August 2026, this Paper Citation Record lists 96 of 96 outbound references and 100 inbound Pith citation observations for arXiv:2503.21755.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.21755 v2

Coverage vector

measured 96 of 96 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-14T18:42:02.940250Z

measured 196 of 196 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 100 of 126 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T10:54:23.761559Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

96 of 96 outbound references displayed

  • verified exact47
  • verified fuzzy45
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch4

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation ba914254-4d15-4d69-bcf3-ff00c43de820 · outbound

This paper cites MagicEdit: High-Fidelity and Temporally Coherent Video Editing.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness MagicEdit: High-Fidelity and Temporally Coherent Video Editing

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:42:03.305992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:704c235c8865682814500948709e965a0b794aa3b2dc8ecc54fab468281f3bdc

Observation 0aae27db-8c0b-4314-8129-07b998544b1a · outbound

This paper cites StableVideo: Text-driven Consistency-aware Diffusion Video Editing.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness StableVideo: Text-driven Consistency-aware Diffusion Video Editing

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:42:03.120227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:27395a085982565aba4110aacfeb855f2bee05571027f7e1ddc2dd3a1fbc98de

Observation 9bf14d16-edfa-4e48-baca-279c2dc30c34 · outbound

This paper cites TokenFlow: Consistent Diffusion Features for Consistent Video Editing.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness TokenFlow: Consistent Diffusion Features for Consistent Video Editing

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:17:47.115702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:2f41bd6ea7021cbff9ef53f082b12cd2285b85fb544b06c618c3f1b5c36d7447

Observation 61d61183-d2ab-43bf-abf8-6e901ed0d23a · outbound

This paper cites INVE: Interactive Neural Video Editing.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness INVE: Interactive Neural Video Editing

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:42:03.263828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:a3c032d02d3f9e25605c04872f90b0f7d57a953c5bc4855099a9482cbec54a8c

Observation 8af9ce0c-0fa1-452a-a43f-ec97a91e515c · outbound

This paper cites VidEdit: Zero-Shot and Spatially Aware Text-Driven Video Editing.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness VidEdit: Zero-Shot and Spatially Aware Text-Driven Video Editing

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:42:03.035650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:e4c6d4c88788ab6aad82df3fb63a3478ddcd982f84da2bfbae02a44095f6519f

Observation 2bd06f04-49e6-4f92-ad6f-eca2bbb6fb48 · outbound

This paper cites Video-P2P: Video Editing with Cross-attention Control.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness Video-P2P: Video Editing with Cross-attention Control

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:42:03.074904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:a9129a46f1a7846232097ec8b6135da33b2ca1fd1ea0e9fa80b78239c545f62a

Observation ab550ae9-dc24-4609-87ac-398428111115 · outbound

This paper cites Towards Consistent Video Editing with Text-to-Image Diffusion Models.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness Towards Consistent Video Editing with Text-to-Image Diffusion Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:42:03.101089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:03cfc740e846726a84718ab878f0816a8f760ad11d9652863bfadf4fe1a93184

Observation 99fe037b-73ba-46aa-bdaa-403fa1eb9ea2 · outbound

This paper cites ControlVideo: Conditional Control for One-shot Text-driven Video Editing and Beyond.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness ControlVideo: Conditional Control for One-shot Text-driven Video Editing and Beyond

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:42:03.106987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:2441ce0b3d18e7b2fc4a912de17c2737edc116a722b10b13098cd4f9455c0f8f

Observation 212609fc-a2bc-44b4-ad08-ba8489fb9c1c · outbound

This paper cites Zero-Shot Video Editing Using Off-The-Shelf Image Diffusion Models.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness Zero-Shot Video Editing Using Off-The-Shelf Image Diffusion Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:42:03.113128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:f408e3fce43202e19bc36b31bc5e3fc889575bf2b0539baae09663bfc122899c

Observation b085ec1d-a1b1-43cc-802f-efce63c829e3 · outbound

This paper cites Pix2video: Video editing using image diffusion.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness Pix2video: Video editing using image diffusion

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:42:03.351996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:cf119fc0eafe497fb9267982e3033195d9c87ac8d9c58ad9b2fcbea3660341a2

Observation d813f99a-6758-46c2-ac02-cc2dd0f7cb0d · outbound

This paper cites FateZero: Fusing Attentions for Zero-shot Text-based Video Editing.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness FateZero: Fusing Attentions for Zero-shot Text-based Video Editing

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:42:03.167787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:1166ebdafb4fac1268fc54dcff191489ae300560d26eb5f627b0d5627dbfcf5e

Observation 1408eb54-f942-4301-8c58-00397d631360 · outbound

This paper cites Shape-aware Text-driven Layered Video Editing.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness Shape-aware Text-driven Layered Video Editing

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T18:42:03.185592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:e612bda999b59a24f48086747ed2709c8a885cceb6745e801a65efc5d3c48d74

Observation 2fd84b6c-4f86-4e94-a3b7-33b696640cd4 · outbound

This paper cites Make-A-Protagonist: Generic Video Editing with An Ensemble of Experts.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness Make-A-Protagonist: Generic Video Editing with An Ensemble of Experts

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T18:42:03.194287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:b4a56638bd144af961e7a09f4ef2eb637085b87fd6482dbbae1b1605f5d00d49

Observation 4c4f5aa0-4469-43ce-af46-85814478e3aa · outbound

This paper cites Videograin: Modulating space- time attention for multi-grained video editing.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness Videograin: Modulating space- time attention for multi-grained video editing

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:42:03.356861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:3f1c3370442310dfd9df408490245edd4e821341030b62682c5deff35eb436ad

Observation af84aaa9-a05c-479a-a5fe-f3c8e93fa322 · outbound

This paper cites Multi-Concept Customization of Text-to-Image Diffusion.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness Multi-Concept Customization of Text-to-Image Diffusion

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:42:03.210335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:f4094eb99014ed4ec5686ad1c8f103196496578273c833bcd049ad189d8be82e

Observation 33d2f420-8534-4bd1-95e2-d9d537fdae1c · outbound

This paper cites DreamPose: Fashion Image-to-Video Synthesis via Stable Diffusion.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness DreamPose: Fashion Image-to-Video Synthesis via Stable Diffusion

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:42:03.244661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:c512a0d3b8df68a7ef74a1e03d3cf0f26bccdf27ae13946fe305e658e462c060

Observation 67b1625f-6bbf-4966-a5db-e6638d3b1ead · outbound

This paper cites Animate-A-Story: Storytelling with Retrieval-Augmented Video Generation.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness Animate-A-Story: Storytelling with Retrieval-Augmented Video Generation

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:42:03.257455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:33f1ef21c612c78f9ed8a1151ff0fdfb22b16bfe2a48b95a77d70cf387bd727d

Observation 673183fc-06e6-420a-90a9-397abf4acf99 · outbound

This paper cites Animatediff: Animate your personalized text-to-image diffusion models without specific tuning.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness Animatediff: Animate your personalized text-to-image diffusion models without specific tuning

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:42:03.361780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:78c74ba541d753b86372ffc3f70f41590997c889354f098d81f8d821722e7516

Observation 3ff369ee-bffb-4214-b2d0-05b85f3042a8 · outbound

This paper cites DynamiCrafter: Animating Open-domain Images with Video Diffusion Priors.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness DynamiCrafter: Animating Open-domain Images with Video Diffusion Priors

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:42:03.281084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:565571f2e959a0fec8a8dc860fd8492a57bcf0e56a3f0c45022061a734716c05

Observation b2a5a481-821d-47ca-b442-25c02053e973 · outbound

This paper cites Cosmos World Foundation Model Platform for Physical AI.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness Cosmos World Foundation Model Platform for Physical AI

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-14T18:42:03.311760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:a123688fbb731b0a0d0902791455b166697bf381b828e753d17c892e146e8f63

Observation da2ed417-d0c8-4e28-a494-b6b97da79a99 · outbound

This paper cites LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:42:03.319741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:49b4094b2307797c38f77e66be7fcdd6e24527d1dc07b750540fb2e38f7bdaee

Observation 36fd5551-48b9-4d7f-a4ad-066365df1f3d · outbound

This paper cites ModelScope Text-to-Video Technical Report.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness ModelScope Text-to-Video Technical Report

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-14T18:42:03.330403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:9eea586c9b9398f459e41fd655f8173102e39e02aca71131327f85a80b8771a0

Observation 5fc17b73-e264-4628-a563-8c92539999b5 · outbound

This paper cites VideoCrafter1: Open Diffusion Models for High-Quality Video Generation.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness VideoCrafter1: Open Diffusion Models for High-Quality Video Generation

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:40:44.284627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:8850205e873b072206d188be93f34aa78fa645adf42c365505c0c295ae556b08

Observation fa345237-6193-4a18-8fac-8eff8ecf5c23 · outbound

This paper cites CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-14T18:42:03.029148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:71c385f1fe14b9554845123efe529a07e678465f9e13b572d3cf59e39fbff3ea

Observation 3dd40a92-8d94-43ce-8535-af053855d6fc · outbound

This paper cites Vbench: Comprehensive benchmark suite for video generative models.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness Vbench: Comprehensive benchmark suite for video generative models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:42:03.366650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:0730fdae849528400d519f8947c3497db97c90a4b8b1bb457352b71c79987a90

Observation 7700a839-f3fd-4584-aebc-0edf89d9011e · outbound

This paper cites VBench++: Comprehensive and Versatile Benchmark Suite for Video Generative Models.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness VBench++: Comprehensive and Versatile Benchmark Suite for Video Generative Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:42:03.062549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:0b61bd91b5848441fe2c3dfd20e0b108993b50e4285d35ed5d8195514f74a8d9

Observation c43040c1-b322-4984-907b-e424e4912422 · outbound

This paper cites Evalcrafter: Benchmarking and evaluating large video generation models.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness Evalcrafter: Benchmarking and evaluating large video generation models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:42:03.371390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:e8c942cc2e6ff93920174b5021224c504a960b777fe5ed78edfcf87802a00ab5

Observation 1f354b25-c8cb-4a5f-852c-aa238d63624f · outbound

This paper cites [Online].

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness [Online]

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:42:03.375721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:97816bf6a0dd1205e841a2cffb25405794503868a54243d51560f2d219a4a623

Observation 9958637a-01d0-4454-8a0d-5e5de704abbf · outbound

This paper cites Team, “Kling,” Accessed December 9, 2024 [Online] https://klingai.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness Team, “Kling,” Accessed December 9, 2024 [Online] https://klingai

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:42:03.380044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:0e85e6ad8f3004b0277e5b4a4e6b4631bd1e7b1c9a421499cbb6384f781bf611

Observation 273c13ec-928d-4bb1-8e26-c6b7785661cf · outbound

This paper cites com/research/introducing-gen-3-alpha.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness com/research/introducing-gen-3-alpha

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:42:03.384201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:b6ad30b10500da780edf9d235ad7a7bc769a09851273ed7b8284e180360b6211

Observation fa4130e4-d2a2-4b68-8c5d-e3c698504879 · outbound

This paper cites Hunyuanvideo: A systematic framework for large video generative models.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness Hunyuanvideo: A systematic framework for large video generative models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:42:03.388926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:c9a81793b9ab0ca2f7594ecff3ae7999aee371e178bd1d3628887b57f696906b

Observation d1a5381e-9c04-487d-8b89-2f8dab47d150 · outbound

This paper cites Team, “Veo2,” Accessed December 18, 2024 [Online] https: //deepmind.google/technologies/veo/veo-2/.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness Team, “Veo2,” Accessed December 18, 2024 [Online] https: //deepmind.google/technologies/veo/veo-2/

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:42:03.393570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:c62b087be81967b5a54f7fd777683d4cb26067ed81ff872f379910607cf12a9f

Observation 58c11fcd-da1a-4903-9a15-dbe16d061edc · outbound

This paper cites Deep unsupervised learning using nonequilibrium thermodynamics.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness Deep unsupervised learning using nonequilibrium thermodynamics

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:42:03.397839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:2cb641cbd1b875db8c495cd0bf25aa156527eb0fa5f76965b1ab4e45745b7f4c

Observation 020fc850-bc8d-43e2-a2aa-39c23398f2a6 · outbound

This paper cites Score-based generative modeling through stochastic differ- ential equations.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness Score-based generative modeling through stochastic differ- ential equations

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:42:03.402472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:fd615de610644ecb30fb4a5f9ef3be1f0a5241590b0dba2332886eca53a032a5

Observation 626b02fb-8c49-498f-a941-0d56457318cf · outbound

This paper cites Denoising diffusion probabilistic models.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness Denoising diffusion probabilistic models

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:42:03.406763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:b100a4c4a5281881393131eb33a26c72c87a16d38c7231fb4880efd0595081c9

Observation d295fc6d-51d4-4670-a1e0-49097dfa3b5c · outbound

This paper cites Denoising diffusion implicit models.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness Denoising diffusion implicit models

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:42:03.411177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:bc5c26d163a94ffb6bdd595eeddb3e1327c7c67f2253ca0e72e52e5ca011037f

Observation 4672fd26-301d-4972-afa2-f243e495c65f · outbound

This paper cites Adding Conditional Control to Text-to-Image Diffusion Models.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness Adding Conditional Control to Text-to-Image Diffusion Models

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:43:11.169114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:a51e54b84c2a8f7cefdc5f4787cdd9723293cea4ef829d4f9cc9d264d27578fd

Observation bd8ea729-be5e-477f-8d2f-4fe0a3743754 · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-14T18:42:03.230991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:0d7c2963c9bf21d213e6659409a2fedac2fad02cdab1474fdb4432171a4f1459

Observation b67829bd-60fe-4e5c-a384-8bcecd2ad8c3 · outbound

This paper cites Scaling rectified flow transformers for high-resolution image synthesis.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness Scaling rectified flow transformers for high-resolution image synthesis

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:42:03.415291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:9ce88a3f125e8f6c6fea389e44138c6a5bdc70649ff2a5a1471ad0f721d5f819

Observation c6088f87-ea1e-44e9-b06b-9bbc4f115da1 · outbound

This paper cites T2I-Adapter: Learning Adapters to Dig out More Controllable Ability for Text-to-Image Diffusion Models.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness T2I-Adapter: Learning Adapters to Dig out More Controllable Ability for Text-to-Image Diffusion Models

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:47:50.630300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:cf3d8899ce0fd2be3c74e801428ced263d0dc70f49dc91cce64a59c713c70260

Observation 2947b7a6-d793-4a72-8739-008030b7d37a · outbound

This paper cites Collaborative diffusion for multi-modal face generation and editing.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness Collaborative diffusion for multi-modal face generation and editing

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:42:03.421002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:4c3af301d133b6689644b88030151d75725ceeb7bd3b5db5f5154a7be7921420

Observation fe9c68e6-615f-45dd-859b-60d0484661b6 · outbound

This paper cites CogView: Mastering text-to-image generation via transformers.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness CogView: Mastering text-to-image generation via transformers

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:42:03.424880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:d708311af25a19bfd898a69dd6b209a024bc9d489b0f6c188d1b4091d45187fa

Observation 15ad4a1f-5cf1-4142-b688-5826b408c6fd · outbound

This paper cites Cogview2: Faster and better text-to-image generation via hierarchical transformers.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness Cogview2: Faster and better text-to-image generation via hierarchical transformers

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:42:03.429684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:72b2fb77bf6789a30cfd4f9638b022d6f2920bf4c00e0d64916fa53a0a7adddc

Observation 72de7d59-0577-4115-b29c-f8158decc8ee · outbound

This paper cites Imagen Video: High Definition Video Generation with Diffusion Models.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness Imagen Video: High Definition Video Generation with Diffusion Models

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-14T18:42:03.293731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:10406cf25f7014f9e9aaf1e6e21feb93c358d2311a41bd709eef00ebe04de9b2

Observation 5773b77c-aa12-4d88-940c-bdb27c410734 · outbound

This paper cites Auto-Encoding Variational Bayes.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness Auto-Encoding Variational Bayes

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-14T18:42:03.299662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:7746093430847cd2d6a65fa27a6767a2613fd68cac000e45ee6ee5484852051b

Observation 18c38134-9789-4c5a-8e63-9133e55bcc95 · outbound

This paper cites Neural discrete representation learning.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness Neural discrete representation learning

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:42:03.434179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:f8e674bf37ff21fef3f67e21683d9121feec7e65cf8c544f6fc011137e3a842c

Observation 2ecde4b6-6175-4201-8b6b-cc300ee384ad · outbound

This paper cites Taming transformers for high- resolution image synthesis.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness Taming transformers for high- resolution image synthesis

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:42:03.439291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:710a4c9a67f2bc4d5cd0ed655ba09b4249169cd3961de924b731bd7edf2e2f74

Observation 0bb84220-036c-4ded-a13e-2e85e919627d · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-05-14T18:42:03.325190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:dcab665b41469f925060bdaa9a01ceee9b5090c49c92639d0a44d0e96c755aed

Observation cac0a3e3-4a3d-4053-a6f0-906e2846d551 · outbound

This paper cites Magvit: Masked generative video transformer.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness Magvit: Masked generative video transformer

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:42:03.443625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:03100a7d2bfbdc4ff0f9e1d311f65b4fa90589db5bd415a515aa55444105e6b1

Observation 953fb594-95b7-46bd-b6dc-7358a049a127 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-05-14T18:42:03.336558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:a12629e19129dc387ff2ac38d8f0a7b26d9f8df8fea222169800b556425f312c

Observation fad64694-d0e8-418a-a68a-9a8d2dc671e2 · outbound

This paper cites Scalable Diffusion Models with Transformers.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness Scalable Diffusion Models with Transformers

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-05-14T18:42:03.342098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:0e5a26e2f1b8dbc89ed952c7ebde5762c2de26b66e1fa66e794a10b37838bf81

Observation 4ed68ca9-affb-4168-95b9-7da6861273de · outbound

This paper cites Videofusion: Decomposed diffusion models for high-quality video generation.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness Videofusion: Decomposed diffusion models for high-quality video generation

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:42:03.447912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:0186ee8f7089107a9d8de6e9e6ed34c20449f790916f5aab2718d1303135d687

Observation 2116d735-d45d-4b37-9f38-714c5765e739 · outbound

This paper cites Latent Video Diffusion Models for High-Fidelity Long Video Generation.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness Latent Video Diffusion Models for High-Fidelity Long Video Generation

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-15T04:27:43.523852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:964d82a7aec4cded2361e173425f67122eb49bb26d39f10b442bd851dba8ca17

Observation a94c373d-5d63-47f6-9fca-05f9b0f2495e · outbound

This paper cites MagicVideo: Efficient Video Generation With Latent Diffusion Models.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness MagicVideo: Efficient Video Generation With Latent Diffusion Models

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:47:57.649874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:7a3e4b79a7ed92d796cc160abfcf30ad544a4f1700844ea00b1a35ec8ac1783a

Observation 79daf691-163a-4b8d-a35a-e57b70bf965e · outbound

This paper cites Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:42:03.022934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:05adc1d417c4876360be6329fd6b9856ffb33ab4aa15e1d55b3a4142c7e1acef

Observation 7afdc277-4732-46c2-8d0c-17d14610ad6e · outbound

This paper cites Preserve your own correlation: A noise prior for video diffusion models.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness Preserve your own correlation: A noise prior for video diffusion models

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:42:03.452601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:fd75256aaea65b8b82e0854e1e39751d05ac7576968d0a97dfe0f26a591f270e

Observation 631ac26c-9748-4ee5-8489-1f0ced631733 · outbound

This paper cites Align your latents: High-resolution video synthesis with latent diffusion models.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness Align your latents: High-resolution video synthesis with latent diffusion models

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:42:03.457294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:47daa898d2db6070296399dafbd51624aa6300342eab05f658509b637587a89e

Observation 7f7ad92f-57d9-425b-b565-19e92af455aa · outbound

This paper cites Text2Video-Zero: Text-to-Image Diffusion Models are Zero-Shot Video Generators.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness Text2Video-Zero: Text-to-Image Diffusion Models are Zero-Shot Video Generators

Reference 58

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T18:42:03.042078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:059332c91208333dc3614153832c1c77bb9f0c2d0f959b01ebe796c83a71d1ae

Observation f035c9d1-27e7-4124-846a-0f4a687a0fc8 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 59

Resolution
verified exact
local_arxiv, observed 2026-05-14T18:42:03.048515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:89e7a908788b1e152570a3d3d538ff74d252fa96dbae6d08e7ccefd6e5c3b2e7

Observation ccc8dbbc-c5fa-4857-9c86-a53272c1d1af · outbound

This paper cites Movie Gen: A Cast of Media Foundation Models.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness Movie Gen: A Cast of Media Foundation Models

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-05-14T18:42:03.055974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:9ba92dd20244891cb01f26f061f92e3e9cf43457910aabe385aec8a40cc28af8

Observation 8be8bbcb-7323-4df3-99c8-eeb53d52e18b · outbound

This paper cites Wan: Open and advanced large-scale video generative models.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness Wan: Open and advanced large-scale video generative models

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:42:03.461626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:782e4750de4ec90c7ed4ba4d51720fe35109dc2ff93e3499706befeb4d6f88e6

Observation 31dcae6c-86f1-42e0-880c-09d71cdcfaf0 · outbound

This paper cites Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-19T08:02:24.560574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:9787c60fb3bb59e161e4b6ba45172e4a571689db036584f842175bd6ee9c56ce

Observation 5508daed-cfb5-4eb8-be50-6b1fe5a4610f · outbound

This paper cites Team, “Minmax,” Accessed August 31, 2024 [Online] https: //hailuoai.com/.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness Team, “Minmax,” Accessed August 31, 2024 [Online] https: //hailuoai.com/

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:42:03.465947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:94e1d0af46ef1ebe4d782b145e10bef6ee1840c485e67a89185b083c9de3c725

Observation bc835403-83a4-48e5-873d-630c80b32ee4 · outbound

This paper cites Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:42:03.081613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:d64e538e24ada0cb3c0beabd68db19dccc08a9388a0986952aed74d339f813d5

Observation a29b4045-5aca-4cfb-909a-de0895d41532 · outbound

This paper cites RepVideo: Rethinking Cross-Layer Representation for Video Generation.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness RepVideo: Rethinking Cross-Layer Representation for Video Generation

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:42:03.088487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:79c84b5b2ccc90d5fbcd320ac16cd91a03ea5a0ee46e3717b3c52a42931b8462

Observation 920a82e0-214c-4678-a179-f6e5abbeb42f · outbound

This paper cites Open-Sora 2.0: Training a Commercial-Level Video Generation Model in $200k.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness Open-Sora 2.0: Training a Commercial-Level Video Generation Model in $200k

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:09:39.930389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:efe5aaa4102f819ce0e1459677118fb02157f31a5f244e0886cc25b1ebae8ecd

Observation 36a0d02e-62d3-4cb4-96d8-31ed49942f54 · outbound

This paper cites GANs trained by a two time-scale update rule converge to a local nash equilibrium.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness GANs trained by a two time-scale update rule converge to a local nash equilibrium

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:42:03.470397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:0423cccfec2764a5d3551793fa311855796a249facc6a0aaba3fcdf43447971d

Observation 9e9d2415-b372-4502-8d10-9216a1ec7b63 · outbound

This paper cites Improved techniques for training gans.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness Improved techniques for training gans

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:42:03.474841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:8dcb881c62791d13e8a5cdd334744339f017da01e32f89f424782f8af02f2f03

Observation f1a6d1c4-2072-4c5b-8318-f4640ea123f5 · outbound

This paper cites FVD: A new metric for video generation.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness FVD: A new metric for video generation

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:42:03.479242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:cc29405a02d2eaac9b6adb3c0f5df00a91503a9b9cbfc66e80939f9e5e3e88c6

Observation 5992c6f6-e435-4c0f-9796-34fc64b654a9 · outbound

This paper cites Fetv: A benchmark for fine-grained evaluation of open-domain text-to-video generation.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness Fetv: A benchmark for fine-grained evaluation of open-domain text-to-video generation

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:42:03.483812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:4bf3097962b3af6af31b36f01063dd066383080d63d14a84dde6c67fa599c73d

Observation adc5faa4-c36c-44e8-b045-06c54c93d8c4 · outbound

This paper cites Evaluation Agent: Efficient and Promptable Evaluation Framework for Visual Generative Models.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness Evaluation Agent: Efficient and Promptable Evaluation Framework for Visual Generative Models

Reference 71

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T18:42:03.126020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:f8b9e0c9d1492b3b1c508bb5b0e7f648f7cfd3653efe9385b5409c9afca09b54

Observation e2e6eac7-7da8-447c-bb74-dd796e7ece3c · outbound

This paper cites Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:40:00.221207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:d1190dc805ff22be551aba3a8d7633a8d45606521b5809460626fc806ca673e1

Observation c3e7287d-1a9d-48af-9ee4-ebd163dfbe9e · outbound

This paper cites T2V-CompBench: A Comprehensive Benchmark for Compositional Text-to-video Generation.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness T2V-CompBench: A Comprehensive Benchmark for Compositional Text-to-video Generation

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:42:03.138921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:2ab441fa70d07e6b61db66e25676ef0db81719a848118b72772c8e098299df53

Observation 736f7829-6ae8-437e-a342-9e09d3c5c1a5 · outbound

This paper cites Is Your World Simulator a Good Story Presenter? A Consecutive Events-Based Benchmark for Future Long Video Generation.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness Is Your World Simulator a Good Story Presenter? A Consecutive Events-Based Benchmark for Future Long Video Generation

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:42:03.145783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:28688c287f98ab738895cf246f412e4e587e3d05a4a011fe8fd698c0849b90c7

Observation 28953e7a-4738-41e5-8848-730399223951 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 75

Resolution
verified exact
local_arxiv, observed 2026-05-14T18:42:03.151308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:9e71616b282ef4eeed6806977e11611a84218806ab232dbbf82099567578abc7

Observation b9edc57b-42a4-46a4-bd98-c10ff3a6f1dc · outbound

This paper cites Qwen2.5 Technical Report.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness Qwen2.5 Technical Report

Reference 76

Resolution
verified exact
local_arxiv, observed 2026-05-14T18:42:03.157051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:c06c700dabc6802bc2f382abbf410b14a42967b785e825fedb95c9c6cffdb1ae

Observation c843e642-3527-4423-8485-b7d014e60095 · outbound

This paper cites Simmim: A simple framework for masked image modeling.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness Simmim: A simple framework for masked image modeling

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:42:03.488586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:86bb58fc4abb2409441c16e08c30407ece6fa5a25a50add2cae2a1c6371ad308

Observation 19a1c183-6010-4650-bea0-f95e76e3d8c9 · outbound

This paper cites Yolo- world: Real-time open-vocabulary object detection.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness Yolo- world: Real-time open-vocabulary object detection

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:42:03.492963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:2d588f265219426dabac4fe3361d0ca7168808707d6d3137408ffa9049b5b1fa

Observation aa281663-f70a-4c83-ada8-bb9c7d645aff · outbound

This paper cites Humanrefiner: Benchmarking abnormal human generation and refining with coarse-to-fine pose-reversible guidance.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness Humanrefiner: Benchmarking abnormal human generation and refining with coarse-to-fine pose-reversible guidance

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:42:03.497553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:86ee2de7faf8c7e219a2be4e8db4a3e3bb893707cb1344b8c58895683d3d44e3

Observation 14e98149-ded4-4d1f-b8c1-4d3197d7c68f · outbound

This paper cites Arcface: Additive angular margin loss for deep face recognition.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness Arcface: Additive angular margin loss for deep face recognition

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:42:03.501878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:ac35f2455925c82082dc579f2ae902e15553f34ff9d651631516a7369524cc78

Observation 5499d47b-2e78-4aee-afd7-cc054eaed578 · outbound

This paper cites Retinaface: Single-shot multi-level face localisation in the wild.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness Retinaface: Single-shot multi-level face localisation in the wild

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:42:03.506340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:1f0a7305bce822b281f05145568416e3e060c444ed73aa0ebe08d55e98f03ba2

Observation 43d002d4-3b19-499c-ab73-b073569f8172 · outbound

This paper cites Very deep convolutional networks for large-scale image recognition.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness Very deep convolutional networks for large-scale image recognition

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:42:03.510549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:eeb36864475dcbd38e5e3de22d4bd4c934da7d6efbd622ffaa4a460bf095361e

Observation 2f400fc0-205e-42a1-a6f6-f7fd6170ef87 · outbound

This paper cites A Neural Algorithm of Artistic Style.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness A Neural Algorithm of Artistic Style

Reference 83

Resolution
verified exact
local_arxiv, observed 2026-05-14T18:42:03.224471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:ab1355c8f07b7d20e367d57caf7d04fef7f9d284576315d49072ee220dcc0e00

Observation 032b2326-a66f-4d5b-997c-a3ff1055e39a · outbound

This paper cites Cotracker: It is better to track together.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness Cotracker: It is better to track together

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:42:03.514988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:bc87979967539da255220ca079d9814b6fc6313a951c1defa0b8ae2db0dbecd2

Observation c6e732e8-35ef-492b-82e6-61fb5ae2c0d9 · outbound

This paper cites Sora Generates Videos with Stunning Geometrical Consistency.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness Sora Generates Videos with Stunning Geometrical Consistency

Reference 85

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:42:03.237706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:a37023065a595f43767c8933903fb642aafab0adba3d442c59e13b4edff18f7f

Observation 3f80b2d5-fbe4-4959-8f47-8201f690eb89 · outbound

This paper cites Sift-the scale invariant feature transform.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness Sift-the scale invariant feature transform

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:42:03.519714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:4e7ddc39b1bb17c02ae355236fd741992216bb1fe1c78ebda93d5d0b14751db7

Observation e7cb0722-131d-4622-80b1-ac6d69f54b70 · outbound

This paper cites Fast approximate nearest neighbors with automatic algorithm configuration.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness Fast approximate nearest neighbors with automatic algorithm configuration

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:42:03.523517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:0d122c63bfb61afc4271f910d87c1c444b05f68c0bd813b31c50c3f8a7830ed0

Observation 03e0f3e9-a660-407d-8101-a251e15fd2c2 · outbound

This paper cites Random sample consensus: a paradigm for model fitting with applications to image analysis and automated cartography.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness Random sample consensus: a paradigm for model fitting with applications to image analysis and automated cartography

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:42:03.527608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:f6b6dfa0268fdac57b7adacd6bde995dde8863acb66f002722e3753634f8f2e9

Observation 41289fae-0ad5-47f7-978b-e5ac333fd3de · outbound

This paper cites Raft: Recurrent all-pairs field transforms for optical flow.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness Raft: Recurrent all-pairs field transforms for optical flow

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:42:03.531919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:3ad4a4f5aab02dc53ce25727be32200020631a627adb263af2862a3739494d18

Observation 7a86cda7-44f5-4ef5-9ed8-d0697f2581ae · outbound

This paper cites Qwen2.5-VL Technical Report.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness Qwen2.5-VL Technical Report

Reference 90

Resolution
verified exact
local_arxiv, observed 2026-05-14T18:42:03.269357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:e743411dcb65742a749063245e840de813bf06e58cd6840fc6ac1df5e7c87220

Observation f47bb99f-d1ac-4019-9fa0-df0c51e5d463 · outbound

This paper cites Hunyuan-Large: An Open-Source MoE Model with 52 Billion Activated Parameters by Tencent.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness Hunyuan-Large: An Open-Source MoE Model with 52 Billion Activated Parameters by Tencent

Reference 91

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:42:03.275424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:1411017460684373f78a967113168186c402c576fef3fb4fb6d0ebbc798e68f9

Observation cf11a842-9673-4075-8df7-be6e953242eb · outbound

This paper cites Thinking in space: How multimodal large language models see, remember, and recall spaces.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness Thinking in space: How multimodal large language models see, remember, and recall spaces

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:42:03.536170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:e38be1dd8e667256661d347b3f159d0aebacad037fd47b52b3280ea4c01a0d3b

Observation 085d88d7-c439-492c-bdea-529675975e09 · outbound

This paper cites Shotbench: Expert-level cinematic understanding in vision-language models.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness Shotbench: Expert-level cinematic understanding in vision-language models

Reference 93

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:42:03.287472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:847fe8e233be2f1994cf32936640efc25003faeb6ff386ed4aadf1f0af08ba18

Observation c45d38c6-c3bb-4f31-95e0-3c3354ccb47d · outbound

This paper cites Synthetic vision: Training vision-language models to understand physics.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness Synthetic vision: Training vision-language models to understand physics

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:42:03.541135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:59cacb3a9ea5a7f3b164fd7d8f55c34856b32c9f2b422a9da6302a99cdc4d144

Observation 45b627b5-d178-46e2-ab38-308587535c17 · outbound

This paper cites Unibench: Visual reasoning requires rethinking vision- language beyond scaling.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness Unibench: Visual reasoning requires rethinking vision- language beyond scaling

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:42:03.546455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:a10d0f7b3932ed1ef1887e48abcf78a67f6ba4b33754b25fd65f453d1011015a

Observation 03147edd-8351-4874-8f66-55bf7d83f6f5 · outbound

This paper cites Yinan He is currently a Research Engineer at Shanghai AI Laboratory, where he is a member of the OpenGVLab.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness Yinan He is currently a Research Engineer at Shanghai AI Laboratory, where he is a member of the OpenGVLab

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:42:03.347038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:9a0ae223e7d497104fa3aa11546d3e363db61c9f7ec9e19e2114a480f651b964

Pith citing papers

Observation 48d2e080-75c0-4e5b-958c-96f9626c6c8c · inbound

Physics-Driven Spatiotemporal Modeling for AI-Generated Video Detection cites this paper.

Physics-Driven Spatiotemporal Modeling for AI-Generated Video Detection VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T10:54:23.761559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:54:23.761559Z digest=sha256:e88c5e13a99e3f0afb92017cc117a50e67916bb0a632ca25d36ad8fa6166346a

Observation d1bb3569-13ef-44bc-8d71-64c311d94929 · inbound

PEDRA: Evaluating the Realism of Pedestrian Dynamics in Video Generation cites this paper.

PEDRA: Evaluating the Realism of Pedestrian Dynamics in Video Generation VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:46.348667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:46.348667Z digest=sha256:4278b26a3a9f8de74c6f4f7294f36968a6a35890e75c3a33baedeab730895338

Observation 9c840c7c-9009-4f58-ad61-19a053859894 · inbound

Diff4Splat: Controllable 4D Scene Generation with Latent Dynamic Reconstruction Models cites this paper.

Diff4Splat: Controllable 4D Scene Generation with Latent Dynamic Reconstruction Models VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 111

Resolution
verified exact
local_arxiv, observed 2026-05-18T02:00:39.460436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T01:59:23.928725Z digest=sha256:f6ae96a06a792551fd64e846d5bf7307e3c277db863d426580c5ea8726ed8448

Observation 6a715269-7317-4d48-bd85-48f9045caa8f · inbound

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation cites this paper.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:11.222875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:11.222875Z digest=sha256:c57ceb8384a8c24c87e4fcda79728dcb928844ba3f000d4f3f348841c570c073

Observation 00b4eaac-2f57-44cb-a055-d6fb0a4a26cb · inbound

Generative Action Tell-Tales: Assessing Human Motion in Synthesized Videos cites this paper.

Generative Action Tell-Tales: Assessing Human Motion in Synthesized Videos VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-03T19:11:53.511931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:11:53.511931Z digest=sha256:0c67be233ae0127790160bbfba14503b03fe7be94c264af77f5ab0301e955883

Observation c209ae9e-d936-4cf3-93c8-9c84627ae244 · inbound

PhyDetEx: Detecting and Explaining the Physical Plausibility of T2V Models cites this paper.

PhyDetEx: Detecting and Explaining the Physical Plausibility of T2V Models VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-05-21T18:00:27.270911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T17:57:57.263574Z digest=sha256:d59516e5619530db4f3781943f45d9784f078f7926949d570650bcf0c31eb864

Observation d8d81511-a1ce-42fd-94cd-f5c679fe3751 · inbound

MIND-V: Hierarchical World Model for Long-Horizon Robotic Manipulation with RL-based Physical Alignment cites this paper.

MIND-V: Hierarchical World Model for Long-Horizon Robotic Manipulation with RL-based Physical Alignment VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-03T18:09:53.091830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:09:53.091830Z digest=sha256:018fc0021be785844149637890e5ad6b381b7fe93a093ac45d4f708f2ef29b0d

Observation df4841dc-dc21-463a-a21f-44ac8b575ebd · inbound

VABench: A Comprehensive Benchmark for Audio-Video Generation cites this paper.

VABench: A Comprehensive Benchmark for Audio-Video Generation VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 59

Resolution
verified exact
local_arxiv, observed 2026-05-17T00:08:43.888647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:03:45.576961Z digest=sha256:1fbf7cfca0e2b247c396afdc15196a2d1c9b3c32aedbc32872a5633b85f1f853

Observation fef3fb4c-309c-4151-8b6a-ab6615a855d1 · inbound

WorldLens: Full-Spectrum Evaluations of Driving World Models in Real World cites this paper.

WorldLens: Full-Spectrum Evaluations of Driving World Models in Real World VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 143

Resolution
unresolved
no resolver link, observed 2026-08-03T17:02:43.826638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:02:43.826638Z digest=sha256:22e9db3435218ff1464b1f868b07b0f90978af5a883c2862d41d4a1e45c7cd7d

Observation e739784d-cd95-4afd-bfc7-7d2021d20d51 · inbound

TinyHistory: Lightweight Video History Embeddings via Two-Stage Context Learning cites this paper.

TinyHistory: Lightweight Video History Embeddings via Two-Stage Context Learning VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-03T13:35:50.055528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:35:50.055528Z digest=sha256:4eedb239cf1c4b735e0e588806cbd0b83115d982eedefb037d82bbd3a8e60bcb

Observation e92d8b5a-297f-4a86-83d0-e80cb53d7e77 · inbound

CoMoVi: Co-Generation of 3D Human Motions and Realistic Videos cites this paper.

CoMoVi: Co-Generation of 3D Human Motions and Realistic Videos VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 114

Resolution
verified exact
local_arxiv, observed 2026-05-16T13:47:57.401677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T13:43:26.460480Z digest=sha256:dd74ac56670280c2eedf597f5f8a54eac871a720fe4544ec4c219f29d9be49be

Observation 291f92dc-54b8-4450-9a32-6c44f1bc1797 · inbound

Rolling Sink: Bridging Limited-Horizon Training and Open-Ended Testing in Autoregressive Video Diffusion cites this paper.

Rolling Sink: Bridging Limited-Horizon Training and Open-Ended Testing in Autoregressive Video Diffusion VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 113

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T07:07:29.949973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T07:02:38.876518Z digest=sha256:d20d162f2dbae95c047b60cff8ddf24ed8a2e2b2f31efd0c8a5193c359a82ebc

Observation f521e21b-80fc-4bec-99cc-e2b38ad47eda · inbound

EchoTorrent: Towards Swift, Sustained, and Streaming Multi-Modal Video Generation cites this paper.

EchoTorrent: Towards Swift, Sustained, and Streaming Multi-Modal Video Generation VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-15T22:20:22.274966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T22:20:16.320171Z digest=sha256:680497d40767c6252d94357c6089991ed5b8445d9c696c71d25eb141f2f1d148

Observation 7c89d9fb-7672-4ac1-9d8b-545a91e91710 · inbound

Are Multimodal LLMs Ready for Surveillance? A Reality Check on Zero-Shot Anomaly Detection in the Wild cites this paper.

Are Multimodal LLMs Ready for Surveillance? A Reality Check on Zero-Shot Anomaly Detection in the Wild VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-05-21T12:24:10.643726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T12:23:02.046757Z digest=sha256:44b7999779d8c66c1cfdbfa2b05f8e858c53c8d6c15e74cfff0c95258ee6010d

Observation 87711a17-bc1d-406e-8bf7-43bd3b0a8a0e · inbound

What if? Emulative Simulation with World Models for Situated Reasoning cites this paper.

What if? Emulative Simulation with World Models for Situated Reasoning VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 125

Resolution
unresolved
no resolver link, observed 2026-07-15T13:51:30.008232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T13:51:30.008232Z digest=sha256:13a0b4da6978cbbcdc6ba12db7dc707ac4c3c6dc3216ac1c934c1e0255e60ba8

Observation cf20a088-0f91-4618-990f-ca1a9eef1480 · inbound

Event-Driven Video Generation cites this paper.

Event-Driven Video Generation VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-14T22:56:23.337858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T22:56:23.337858Z digest=sha256:a6f90e8dd1cf7fc56b88bb45d463b62319b71898161ccd2c3ecb5e7ae6c60abe

Observation 77ffc277-090a-4635-97fb-05ac0de0481d · inbound

LatSearch: Latent Reward-Guided Search for Faster Inference-Time Scaling in Video Diffusion cites this paper.

LatSearch: Latent Reward-Guided Search for Faster Inference-Time Scaling in Video Diffusion VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 62

Resolution
unresolved
no resolver link, observed 2026-07-14T21:09:36.040879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T21:09:36.040879Z digest=sha256:6696bd4c45040c0f0526056e71f7284d50495d90704aa660d6aa15d6cf2799ed

Observation ed668bc2-4087-4903-919d-98214c2caa99 · inbound

Demystifying Video Reasoning cites this paper.

Demystifying Video Reasoning VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 78

Resolution
unresolved
no resolver link, observed 2026-07-13T23:27:11.006580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:27:11.006580Z digest=sha256:2addfbd89728c0e811bf579d714ef24569010be5b756710bfebc97d008cdf43d

Observation 0871f3ab-e617-4dea-8835-4fdf9977c74b · inbound

Demystifying Video Reasoning cites this paper.

Demystifying Video Reasoning VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-03T02:34:01.381315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:34:01.381315Z digest=sha256:7277401877d9c39f15950fd9c9e0c85f520ce2fe832f4e159e0f428c026f2d91

Observation 719b23bb-341b-4d62-a191-f0ba63c89016 · inbound

Attention Sparsity is Input-Stable: Training-Free Sparse Attention for Video Generation via Offline Sparsity Profiling and Online QK Co-Clustering cites this paper.

Attention Sparsity is Input-Stable: Training-Free Sparse Attention for Video Generation via Offline Sparsity Profiling and Online QK Co-Clustering VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-15T08:59:53.186042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T08:55:52.757489Z digest=sha256:f4b275f6eb08ece0d89b76db8a8cfa7c722f17cd480b92df0b22711fb215ca19

Observation f6141929-8344-4cf3-bb2f-bbf2f6333b8a · inbound

Evolution of Video Generative Foundations cites this paper.

Evolution of Video Generative Foundations VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 167

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:42:03.548540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:41:38.616611Z digest=sha256:607e645754a8cae8c856dacbbc13380cbf9626a7f4c00f3af09f2456ddc7d7f5

Observation 51c9e71b-2bd5-420c-87da-8a95373434c8 · inbound

Grounded Forcing: Bridging Time-Independent Semantics and Proximal Dynamics in Autoregressive Video Synthesis cites this paper.

Grounded Forcing: Bridging Time-Independent Semantics and Proximal Dynamics in Autoregressive Video Synthesis VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:42:03.548540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:32:01.665435Z digest=sha256:a156e7a111aeff898c7681450e0f13dddc6e20300f9b2a569d67d4e785e9d12b

Observation 07c10dc0-b2ce-4af9-a18d-efdbe2818d71 · inbound

ImVideoEdit: Image-learning Video Editing via 2D Spatial Difference Attention Blocks cites this paper.

ImVideoEdit: Image-learning Video Editing via 2D Spatial Difference Attention Blocks VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:42:03.548540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T17:49:33.822264Z digest=sha256:c4a5658c74a91162a5835c888d07c475d043242d5bf9ff1de86fb99fd59e0975

Observation 45f169c0-6fbc-4050-ba06-3d967b51bb53 · inbound

Phantom: Physics-Infused Video Generation via Joint Modeling of Visual and Latent Physical Dynamics cites this paper.

Phantom: Physics-Infused Video Generation via Joint Modeling of Visual and Latent Physical Dynamics VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 49

Resolution
malformed identifier
arxiv_id, observed 2026-05-14T18:42:03.548540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:15:37.338442Z digest=sha256:4b636921843fabac0667df8553cc1e64ff2b0e585f014fb66dc15504dfa5da75

Observation 5c8ae224-0e31-44d3-8c07-69f7d1fb8bfa · inbound

Phantom: Physics-Infused Video Generation via Joint Modeling of Visual and Latent Physical Dynamics cites this paper.

Phantom: Physics-Infused Video Generation via Joint Modeling of Visual and Latent Physical Dynamics VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 49

Resolution
malformed identifier
local_arxiv, observed 2026-05-21T09:19:56.631167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T09:17:53.731701Z digest=sha256:ba27b72c76bb5b0181abc8c1fe95b5d6e145d3028eac8857280f57e15b44875a

Observation 35875d69-0581-4bc9-b2e4-4589cc23fa50 · inbound

PhysInOne: Visual Physics Learning and Reasoning in One Suite cites this paper.

PhysInOne: Visual Physics Learning and Reasoning in One Suite VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 102

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:42:03.548540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T16:39:48.066744Z digest=sha256:bcdea9b3e210edcdc5a7d1652a3ac4e12d18a35a422b45d48bc97352b3b3766d

Observation c44fc420-4c93-4dcc-9607-6c66d6cea954 · inbound

VGA-Bench: A Unified Benchmark and Multi-Model Framework for Video Aesthetics and Generation Quality Evaluation cites this paper.

VGA-Bench: A Unified Benchmark and Multi-Model Framework for Video Aesthetics and Generation Quality Evaluation VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:42:03.548540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:44:02.709823Z digest=sha256:c411f1e1b3786640b3157ee1e7dce085267e8a1f6d62475a55b6ebf4274e41e2

Observation 0cc9bd07-3732-41a5-a679-84c3b5c05826 · inbound

AnimationBench: Are Video Models Good at Character-Centric Animation? cites this paper.

AnimationBench: Are Video Models Good at Character-Centric Animation? VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:42:03.548540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T11:47:27.131584Z digest=sha256:7e713592c5b8cb8c3032a95b94648b6b7ef8f58a2662e58bb84ddf2f47efc3e0

Observation 95534f3d-8e80-45d9-87cf-2f02e42d7709 · inbound

Long-CODE: Isolating Pure Long-Context as an Orthogonal Dimension in Video Evaluation cites this paper.

Long-CODE: Isolating Pure Long-Context as an Orthogonal Dimension in Video Evaluation VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:42:03.548540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T05:55:51.789570Z digest=sha256:09105ac452c899412fe483f87ca6186104268876c1fa1907a4dba1b41e8b2ee5

Observation 93dd91ef-3576-4bfa-baf2-7a34cf65be63 · inbound

RoboWM-Bench: A Benchmark for Evaluating World Models in Robotic Manipulation cites this paper.

RoboWM-Bench: A Benchmark for Evaluating World Models in Robotic Manipulation VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 60

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T18:42:03.548540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T03:02:26.084185Z digest=sha256:122e2b29060275b55e017b6d8428fd0accc756879677a35c0666f4d6cf9c0f55

Observation 76fecfd6-f0b7-40e0-8667-20abb4805852 · inbound

RoboWM-Bench: A Benchmark for Evaluating World Models in Robotic Manipulation cites this paper.

RoboWM-Bench: A Benchmark for Evaluating World Models in Robotic Manipulation VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 60

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T06:19:50.182670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T06:15:32.881140Z digest=sha256:4d6df5179b75dc45f0b08b1ae7ef6293e4547ae176640053a2a48d63c60400a7

Observation c8a34807-950b-4a13-ae38-70acb8307eed · inbound

Seeing Fast and Slow: Learning the Flow of Time in Videos cites this paper.

Seeing Fast and Slow: Learning the Flow of Time in Videos VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:42:03.548540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-09T21:55:17.377887Z digest=sha256:15e7ec198e14bb4727d214bf8f1458e9b099b73c60e2d574a4bbb1ccadde8ceb

Observation 9905d9fe-7616-4981-8e01-82be8b333db6 · inbound

Hallo-Live: Real-Time Streaming Joint Audio-Video Avatar Generation with Asynchronous Dual-Stream and Human-Centric Preference Distillation cites this paper.

Hallo-Live: Real-Time Streaming Joint Audio-Video Avatar Generation with Asynchronous Dual-Stream and Human-Centric Preference Distillation VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:42:03.548540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T06:56:19.795651Z digest=sha256:7ad850c822e538767d1176ccca35886fed6e79f9b2e99d32833aaeefd943ba2e

Observation 05d7f6f9-2afb-4618-b5ef-45ec5da75177 · inbound

HuM-Eval: A Coarse-to-Fine Framework for Human-Centric Video Evaluation cites this paper.

HuM-Eval: A Coarse-to-Fine Framework for Human-Centric Video Evaluation VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:42:03.548540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:53:42.069330Z digest=sha256:90f5d10cf75a7c44c3427132eaac1d6dc2542deb2b29fd441ccc283c3d42a61b

Observation 2d882fe4-0c96-445e-9794-b98a06955be2 · inbound

Mamoda2.5: Enhancing Unified Multimodal Model with DiT-MoE cites this paper.

Mamoda2.5: Enhancing Unified Multimodal Model with DiT-MoE VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:42:03.548540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T18:26:58.696936Z digest=sha256:877a534c5dd5c83dc0ac0f2908ea86bf73f5279694b6ca62d938ac59ec8e7bef

Observation f9dd60d1-2874-4631-b7c6-2a210c7980fc · inbound

WorldJen: An End-to-End Multi-Dimensional Benchmark for Generative Video Models cites this paper.

WorldJen: An End-to-End Multi-Dimensional Benchmark for Generative Video Models VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:42:03.548540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T04:13:26.433133Z digest=sha256:16238686ad910b1fc87d07eefab416b0246b22c1b6da8b021a76b32209f455f1

Observation c94f50d6-ba34-4fcc-ae69-cb75797ae008 · inbound

WorldJen: An End-to-End Multi-Dimensional Benchmark for Generative Video Models cites this paper.

WorldJen: An End-to-End Multi-Dimensional Benchmark for Generative Video Models VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:42:03.548540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T19:08:07.487384Z digest=sha256:8ff76d29c5ded0908c372ad75aab494bee91996ca85acc3a8ee126f6aceca22e

Observation e726c2a2-2b93-4152-b367-6d2966d3c0bf · inbound

LoViF 2026 The First Challenge on Holistic Quality Assessment for 4D World Model (PhyScore) cites this paper.

LoViF 2026 The First Challenge on Holistic Quality Assessment for 4D World Model (PhyScore) VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:42:03.548540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T16:20:40.024442Z digest=sha256:33989f28f39ff5da8079b9c5276f296666a671ea6a90075c0f491d9d8970676c

Observation 89a46d18-ad7e-4eec-9f40-2485b79f3f43 · inbound

Diffusion-APO: Trajectory-Aware Direct Preference Alignment for Video Diffusion Transformers cites this paper.

Diffusion-APO: Trajectory-Aware Direct Preference Alignment for Video Diffusion Transformers VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:42:03.548540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:58:25.291448Z digest=sha256:ed3399f3d071851d52e3c7f03ed140119e513770aa7ca515abdd7487f7a6e7db

Observation 4be55b5e-4e0b-4820-9815-f4f825f0de85 · inbound

SARA: Semantically Adaptive Relational Alignment for Video Diffusion Models cites this paper.

SARA: Semantically Adaptive Relational Alignment for Video Diffusion Models VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:42:03.548540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T02:21:52.861714Z digest=sha256:0779f482519e15518a65a26b1afdefa0dc7349e009bdf1829ac4ca7f4471f14f

Observation 72c4ceeb-1209-4b66-9595-832b79380c7b · inbound

SARA: Semantically Adaptive Relational Alignment for Video Diffusion Models cites this paper.

SARA: Semantically Adaptive Relational Alignment for Video Diffusion Models VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-06-30T23:15:07.972570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T23:13:31.195562Z digest=sha256:dccd59bf5845570f978bc69edb78686a2062cf2c762228657f778487a5cceafe

Observation 68692374-cc1f-4b67-bb2d-865ebdda8e73 · inbound

WorldReasonBench: Human-Aligned Stress Testing of Video Generators as Future World-State Predictors cites this paper.

WorldReasonBench: Human-Aligned Stress Testing of Video Generators as Future World-State Predictors VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:42:03.548540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T03:39:01.254916Z digest=sha256:3304cd76d533416e9d3eaf9c64e2eb9c3f17fd4ed681d400307b700d6e921c94

Observation 876ce071-c6a5-494f-b23c-0f48aed884a6 · inbound

PhyGround: Benchmarking Physical Reasoning in Generative World Models cites this paper.

PhyGround: Benchmarking Physical Reasoning in Generative World Models VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:42:03.548540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T05:17:30.010064Z digest=sha256:77afef838d9e129dfd9cad43cefd3459b637044f9a59f92d8a7bc7ed73f75d35

Observation fd1cf23c-94e7-4639-a279-ccc890647146 · inbound

World Action Models: The Next Frontier in Embodied AI cites this paper.

World Action Models: The Next Frontier in Embodied AI VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 214

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:42:03.548540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T05:01:16.802019Z digest=sha256:eaced405371c24f1154f399a6bd9782e91c3b15b4f804c233cedb09d50122588

Observation d8452ca0-bc95-4925-9e73-3c92d414f3de · inbound

PhyMotion: Structured 3D Motion Reward for Physics-Grounded Human Video Generation cites this paper.

PhyMotion: Structured 3D Motion Reward for Physics-Grounded Human Video Generation VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-15T02:49:41.457390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T02:49:21.291716Z digest=sha256:6c2b7075268e40388bc4a82380ebdb836c529bde5765cbc571a6819836704d66

Observation 08d7cb65-e235-4e37-9e86-d1d2457a4bf4 · inbound

KVPO: ODE-Native GRPO for Autoregressive Video Alignment via KV Semantic Exploration cites this paper.

KVPO: ODE-Native GRPO for Autoregressive Video Alignment via KV Semantic Exploration VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 32

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T02:38:34.447281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T02:38:14.180433Z digest=sha256:49ed794fdfabc43bfcea9d25d88ed29deec0f8b874a3b43e85814b0304e2856b

Observation fa154dfa-102c-42cc-9229-b9c7fcbc646c · inbound

Delta Forcing: Trust Region Steering for Interactive Autoregressive Video Generation cites this paper.

Delta Forcing: Trust Region Steering for Interactive Autoregressive Video Generation VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-05-15T01:53:28.794513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:52:14.874049Z digest=sha256:fbcfdd534ef79aa0efd378b37e93f78fd682b6f448eb19b39effdb1b03845f1c

Observation efac5b77-bdd3-4386-bfeb-51c9f34cd9f2 · inbound

Delta Forcing: Trust Region Steering for Interactive Autoregressive Video Generation cites this paper.

Delta Forcing: Trust Region Steering for Interactive Autoregressive Video Generation VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-20T21:59:06.467855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T21:54:33.902256Z digest=sha256:e27fb02aff4c2fda11da614852fc2eed14964f72747dea69f3436ce6a0f82d63

Observation 384a8bc9-d1b1-4831-adee-e6794238fda3 · inbound

Delta Forcing: Trust Region Steering for Interactive Autoregressive Video Generation cites this paper.

Delta Forcing: Trust Region Steering for Interactive Autoregressive Video Generation VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-21T09:14:05.651779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T09:12:35.777810Z digest=sha256:4d4bb0af131e70b220fa70a526b1dacbcb015333393c3c57e85e80f5e0154a2d

Observation 7d9a6d56-cfe5-4c7a-9d10-fc809f2cdf6e · inbound

Delta Forcing: Trust Region Steering for Interactive Autoregressive Video Generation cites this paper.

Delta Forcing: Trust Region Steering for Interactive Autoregressive Video Generation VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-06-30T21:45:05.549385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T21:44:21.427570Z digest=sha256:5222097508a918abb2fa0a61d1a62e05287d50f210f4340f6de415826890d16f

Observation d603e154-5c34-4289-aa7f-ee6f68c3e232 · inbound

Probing into Camera Control of Video Models cites this paper.

Probing into Camera Control of Video Models VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-06-30T21:15:04.297273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T21:11:19.441408Z digest=sha256:49dfcf401a752bfed6000c887d4bc9ca202fd52427beb5c2855fd41ea5835918

Observation 051107ee-7eda-4a79-9b1f-50fba12854d4 · inbound

DriveCtrl: Conditioned Sim-to-Real Driving Video Generation cites this paper.

DriveCtrl: Conditioned Sim-to-Real Driving Video Generation VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-06-30T21:15:04.749003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T21:06:06.548538Z digest=sha256:6db098ea46b52d0e7d404af32e7213cd1686271a5bce8f899a87e5a5aee5584d

Observation b665fa80-5335-4075-97a5-d02a58887526 · inbound

EntityBench: Towards Entity-Consistent Long-Range Multi-Shot Video Generation cites this paper.

EntityBench: Towards Entity-Consistent Long-Range Multi-Shot Video Generation VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-15T03:19:43.927820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T03:16:03.450742Z digest=sha256:7cf1f0e836fde97b1a0393b8fe92ed90268076c4dfb550aca67c8a0c625cb4d3

Observation 36276399-b29d-4e4d-9d0e-bf1848bbbf2d · inbound

Echo-Forcing: A Scene Memory Framework for Interactive Long Video Generation cites this paper.

Echo-Forcing: A Scene Memory Framework for Interactive Long Video Generation VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-05-20T20:03:43.454867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T20:02:45.505404Z digest=sha256:86f6b3feb59b09da60f049227a1df0b7bdd471c5c0be14b4279a76aebf8d8a9a

Observation 13a6eff0-4f16-407b-bc65-e6c49a4ddadd · inbound

PH-Dreamer: A Physics-Driven World Model via Port-Hamiltonian Generative Dynamics cites this paper.

PH-Dreamer: A Physics-Driven World Model via Port-Hamiltonian Generative Dynamics VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-20T12:28:16.664868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-20T12:27:13.931203Z digest=sha256:7b8bab36b2e2b347081ccdd3fe35d92290435707a468ab3b053cc5b81027f6ba

Observation ab5342ea-7643-4e4a-b5e0-9ef3c1297888 · inbound

GeoFlow: Enforcing Implicit Geometric Consistency in Video Generation cites this paper.

GeoFlow: Enforcing Implicit Geometric Consistency in Video Generation VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 93

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T11:08:13.576182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T11:06:09.367559Z digest=sha256:a38ed7092919f65bf572b36e388bd3b3755fa9029ab9506f8463f76716ff44d2

Observation 88cdd745-86b4-4b61-9881-79dc6de571ff · inbound

NEWTON: Agentic Planning for Physically Grounded Video Generation cites this paper.

NEWTON: Agentic Planning for Physically Grounded Video Generation VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-20T10:58:13.768504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T10:56:22.343333Z digest=sha256:d3f20137bc4ed4936ab64670a9ac50993a574a9c7880b331c9f6b45042719de0

Observation f7a7b085-5b71-4cb8-8363-bf3956deddd1 · inbound

Aero-World: Action-Conditioned Aerial Video Generation from Inertial Controls cites this paper.

Aero-World: Action-Conditioned Aerial Video Generation from Inertial Controls VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-20T05:43:05.314784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T05:42:59.996351Z digest=sha256:3d4663415becc054a6258d1d1cd4405676a5afd3260f709f4d42e306250a6ca6

Observation 69412894-ea02-48f8-8557-4a4b19f957fa · inbound

World-Ego Modeling for Long-Horizon Evolution in Hybrid Embodied Tasks cites this paper.

World-Ego Modeling for Long-Horizon Evolution in Hybrid Embodied Tasks VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 78

Resolution
verified exact
local_arxiv, observed 2026-05-20T06:38:05.568899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T06:36:47.265734Z digest=sha256:ef08266ca9c4dd5c519f35e0644b55344e4c8f5712fec0b7b9df0da361acf6bc

Observation 5bcded02-8d64-4f31-af7f-52e916209f62 · inbound

TASTE: A Designer-Annotated Multi-Dimensional Preference Dataset for AI-Generated Graphic Design cites this paper.

TASTE: A Designer-Annotated Multi-Dimensional Preference Dataset for AI-Generated Graphic Design VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-21T05:09:38.435915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T05:09:05.834499Z digest=sha256:3a1f490502e3f507a141be0743c1b745105c2470553ad8ee18150805e6d141d3

Observation e818ee83-4b7a-4287-93dd-a08713b638cd · inbound

TASTE: A Designer-Annotated Multi-Dimensional Preference Dataset for AI-Generated Graphic Design cites this paper.

TASTE: A Designer-Annotated Multi-Dimensional Preference Dataset for AI-Generated Graphic Design VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-07-01T15:05:48.014950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T17:47:47.801278Z digest=sha256:3f5060a957d1f7fd4b73b4b7d7ae955aa6c170f57e84f7b1cc2a7c4d1eecf70d

Observation 2823c6e8-d942-4878-b9bf-803d8a700c40 · inbound

EvalVerse: Pipeline-Aware and Expert-Calibrated Benchmarking for Professional Cinematic Video Generation cites this paper.

EvalVerse: Pipeline-Aware and Expert-Calibrated Benchmarking for Professional Cinematic Video Generation VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-25T04:35:20.647662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-25T04:34:58.052261Z digest=sha256:0a492628679fe59874f71c1ad25b677cf7c88ecec8272ccbf2da9bebba415b97

Observation c4e37179-df03-4788-a0b1-271e57acf0e6 · inbound

CRONOS: Benchmarking Counterfactual Physical Consistency in Video Models cites this paper.

CRONOS: Benchmarking Counterfactual Physical Consistency in Video Models VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 55

Resolution
metadata mismatch
local_arxiv, observed 2026-05-25T04:40:23.301198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-25T04:39:22.400458Z digest=sha256:5968bbcb1bd29df4f23aa3b2264557c657092b8e799bf7fcfb02e974e44fba16

Observation 5de8f7d2-f5c4-429c-b8a7-1d2733c17daa · inbound

{\Phi}-Noise: Training-Free Temporal Video Conditioning via Phase-Based Noise Manipulation cites this paper.

{\Phi}-Noise: Training-Free Temporal Video Conditioning via Phase-Based Noise Manipulation VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 64

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T13:44:40.548924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T13:43:50.161133Z digest=sha256:440c5c805fae5a597231eb52bdaf5a2ffe3318535fe3b3f9e8209363d9fac3ae

Observation 748b889e-3b59-410e-a5db-9cbc03597e8f · inbound

Toward Native Multimodal Modeling: A Roadmap cites this paper.

Toward Native Multimodal Modeling: A Roadmap VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 177

Resolution
verified exact
local_arxiv, observed 2026-06-29T23:04:01.996944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T22:58:38.610609Z digest=sha256:afdfb77f0491ec913139009ab746b8a0856c276801c7b77288eef72571eec761

Observation 1a8e8902-5c98-47f7-9334-02f415bedda7 · inbound

StreamChar: Long-Horizon Streaming Character Audio-Video Generation with Decoupled Orchestration cites this paper.

StreamChar: Long-Horizon Streaming Character Audio-Video Generation with Decoupled Orchestration VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:24:00.833339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T22:15:17.804936Z digest=sha256:90ae76ecfb90d1c0989e555105f247143b44afe905dfbb8e97d0def504e6212d

Observation 324a26da-c816-4261-8dc1-db12d2c7046e · inbound

WBench: A Comprehensive Multi-turn Benchmark for Interactive Video World Model Evaluation cites this paper.

WBench: A Comprehensive Multi-turn Benchmark for Interactive Video World Model Evaluation VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-06-29T23:34:04.893918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T22:57:08.381846Z digest=sha256:245b9d0bbd0b6e54fa3712258db1560b6eced8fd2af948ef65ec83bd657fa73d

Observation db762ebe-603c-4586-a067-5bce8530f9ea · inbound

Are Video Models Zero-Shot Learners and Reasoners in Education? EduVideoBench, A Knowledge-Skills-Attitude Benchmark for Educational Video Generation cites this paper.

Are Video Models Zero-Shot Learners and Reasoners in Education? EduVideoBench, A Knowledge-Skills-Attitude Benchmark for Educational Video Generation VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-06-29T18:23:50.878167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T18:17:52.284353Z digest=sha256:a6659a9217ed87662d9945dad97ad74193bfc98a2a2ac8f0fa6ce090fdf80779

Observation aaddc02a-7f16-4cdc-bf6f-94455700735c · inbound

What-If World: A Causal Benchmark for General World Models in Embodied Scenarios cites this paper.

What-If World: A Causal Benchmark for General World Models in Embodied Scenarios VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 78

Resolution
verified exact
local_arxiv, observed 2026-06-29T18:23:50.420197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T18:23:22.987086Z digest=sha256:6beeba340d6484340d81726e1a355d019786e45375794a10b2ec6a5fe688a4d5

Observation 27571782-f3bb-4227-a2f9-c9d79dad29dc · inbound

MiraBench: Evaluating Action-Conditioned Reliability in Robotic World Models cites this paper.

MiraBench: Evaluating Action-Conditioned Reliability in Robotic World Models VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-06-29T07:53:13.865967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T07:46:28.469913Z digest=sha256:562b39a9cfb98a585f9ef537ddc440055c58222dfb36249aa4e94bcc7ebe5ef0

Observation 77964db8-ffa5-477a-a3c7-3e178f32fe9f · inbound

DirectorBench: Diagnosing Long-Form Video Generation with Personalized Multi-Agent Evaluation cites this paper.

DirectorBench: Diagnosing Long-Form Video Generation with Personalized Multi-Agent Evaluation VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-06-29T08:03:13.607331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T08:00:47.144439Z digest=sha256:b1aec571715e05de36e228552179e1ed08aa5c5ddebc8d51f13f72c9c2a0f2d0

Observation 0e3e906a-261c-4b4b-95e9-2c0414226947 · inbound

YoCausal: How Far is Video Generation from World Model? A Causality Perspective cites this paper.

YoCausal: How Far is Video Generation from World Model? A Causality Perspective VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 132

Resolution
verified exact
local_arxiv, observed 2026-06-29T08:33:15.616945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T08:27:03.674229Z digest=sha256:f0a64f07e48c14b39c29555dcb9ba1f4bdbebf64b2f41499cdbe0693ee396aea

Observation e4d837c3-d93d-4010-bbf0-90d83ba5839e · inbound

TunerDiT: Training-free Progressive Steering of Diffusion Transformer for Multi-Event Video Generation cites this paper.

TunerDiT: Training-free Progressive Steering of Diffusion Transformer for Multi-Event Video Generation VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-06-28T23:12:46.585772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T23:11:58.580004Z digest=sha256:7a8fab937d83dc73e664d774c77e4ee9171da3db1343592a97d8aa9bc6fa2a1d

Observation f39305a2-4a6b-41ef-ad91-b22585cceff1 · inbound

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models cites this paper.

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 98

Resolution
verified exact
local_arxiv, observed 2026-06-28T19:22:35.113216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T19:04:58.672464Z digest=sha256:2e6c90ea628b2e33f32c415ea53f1eab4d96eb13d865f5e6857318ae193511e0

Observation de71d412-7938-449a-8926-ed49498bfff1 · inbound

Echo-Infinity: Learning Evolving Memory for Real-Time Infinite Video Generation cites this paper.

Echo-Infinity: Learning Evolving Memory for Real-Time Infinite Video Generation VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-07-02T11:36:54.897618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T03:27:02.232039Z digest=sha256:3b92ba7b2b2d180840d6909e62864e4d56136df5902d9c21838e65949003c2c3

Observation c0694b18-4a2d-4c96-bc2f-2fdf5d4a00b1 · inbound

Physics-Informed Video Generation via Mixture-of-Experts Latent Alignment cites this paper.

Physics-Informed Video Generation via Mixture-of-Experts Latent Alignment VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-07-02T07:16:44.749104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T07:02:37.291472Z digest=sha256:13baf6b314ed6c4a0074ea53f5f7ff91ab962dbd09c43367d3266e64a625a4e1

Observation de8b3afa-52d2-418c-839a-67364ca93d67 · inbound

V2V-Bench: A Comprehensive Benchmark for Video-to-Video Generation Evaluation cites this paper.

V2V-Bench: A Comprehensive Benchmark for Video-to-Video Generation Evaluation VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-02T12:16:56.695068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T02:20:05.748127Z digest=sha256:5e7f0e0c11b73304b24fbd45c20c7ee8292b8fe7d1f22ead36dec80768ce557a

Observation 07563dcc-115d-4010-a987-114aa518c494 · inbound

Test-Time Scaling in Multimodal Foundation Models: A Comprehensive Survey of Generation and Reasoning cites this paper.

Test-Time Scaling in Multimodal Foundation Models: A Comprehensive Survey of Generation and Reasoning VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 116

Resolution
verified exact
local_arxiv, observed 2026-07-02T21:37:25.416891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-27T19:36:57.231932Z digest=sha256:a8b50d1476c6a166358cb509c0cc96f406f37346dfaf0e416026a440c72091c2

Observation 8a7b593b-41e2-498b-ace9-211bd0ac6387 · inbound

CineDance: Towards Next-Generation Multi-Shot Long-Form Cinematic Audio-Video Generation cites this paper.

CineDance: Towards Next-Generation Multi-Shot Long-Form Cinematic Audio-Video Generation VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 98

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T00:07:28.022371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T17:30:25.371658Z digest=sha256:bc577d98c5784fa959dc93ab2be4e964432f3732b23858df3f3782d14d42d8ad

Observation 1a9097e2-3ca7-439c-ae3c-0bf96d0eeb1b · inbound

Can Image Models Imagine Time? ImageTime: A Novel Benchmark for Probing Visual World Modeling Through Spatiotemporal Consistency cites this paper.

Can Image Models Imagine Time? ImageTime: A Novel Benchmark for Probing Visual World Modeling Through Spatiotemporal Consistency VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-07-03T04:57:38.089097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T13:34:25.079037Z digest=sha256:cd2a34b9410dc7c66fcffa8a9b4d880554735426a71b18ffedca4afe45eb8dda

Observation 0ec81154-947f-4089-a48e-25184b4ea122 · inbound

WorldOlympiad: Can Your World Model Survive a Triathlon? cites this paper.

WorldOlympiad: Can Your World Model Survive a Triathlon? VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-07-03T05:47:41.329969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T13:05:26.397711Z digest=sha256:f7bce8b36cb625ee8513437bb8d4f68ef8fa81f5276f2f53e70eb42ac1823e42

Observation 4dbcfeab-de94-447e-bb06-d890d11d46bd · inbound

OmniDrive: An LLM-Choreographed Multi-Agent World Model with Unified Latent Co-Compression for Multi-View Driving Video Generation cites this paper.

OmniDrive: An LLM-Choreographed Multi-Agent World Model with Unified Latent Co-Compression for Multi-View Driving Video Generation VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 68

Resolution
malformed identifier
local_arxiv, observed 2026-07-03T19:28:52.698300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T01:46:14.430539Z digest=sha256:76fce709cea042f216c72f3b47d2d20e70620b55f8ffa87865005eb407a1b761

Observation 16a955d3-dc51-445b-8f9d-e347351367a4 · inbound

Bridging Creative Intent and Visual Quality: Creator-Driven Recurrent Video Generation with Agentic Feedback Loops cites this paper.

Bridging Creative Intent and Visual Quality: Creator-Driven Recurrent Video Generation with Agentic Feedback Loops VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T23:39:03.326182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T22:02:18.035999Z digest=sha256:abaa6a4620c1bd8dfb59c09eac9cc474c6efb3e0af3f51f04cf0a9579a97f010

Observation 39818612-6086-477c-8d37-1dc8e6664cb4 · inbound

Physics-IQ Verified cites this paper.

Physics-IQ Verified VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:29:15.588672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T21:13:36.568700Z digest=sha256:ef3abdd0a10dc7855e8869490725f76889eeb7f598443bf21a0729424490966d

Observation f5d4044c-b1f6-4307-83b2-219ba02675f4 · inbound

Current World Models Lack a Persistent State Core cites this paper.

Current World Models Lack a Persistent State Core VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-04T03:49:31.001578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T17:33:41.461245Z digest=sha256:b538c8fcf942d96d73fd6271ac18391fb3ca12626edab6309d8a13826a45b41a

Observation e06c768e-ab04-4609-94d1-bd17330571ad · inbound

CoDMD: Copula-aware Distribution Matching Distillation for Fast Video Generation cites this paper.

CoDMD: Copula-aware Distribution Matching Distillation for Fast Video Generation VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-07-04T07:49:38.702544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T12:52:05.465272Z digest=sha256:70a1f82d94fe3d6796fab256905b5321d2da780287ea83a29609b00f3bb11618

Observation 5576326d-9deb-4b3e-a20f-e93d59b39fec · inbound

GeoT2V-Bench: Benchmarking 3D Consistency in Text-to-Video Models via 3D Reconstruction cites this paper.

GeoT2V-Bench: Benchmarking 3D Consistency in Text-to-Video Models via 3D Reconstruction VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-07-04T16:39:57.132385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T00:28:33.734545Z digest=sha256:5fd87140d02e18602c521f95d64a6205a71113472e5c5b32a141bc34d30ada00

Observation 70fab6d5-95c5-4b68-ae6f-5cee8af5347c · inbound

Follow Your Track: Precise Skeleton Animation Controlled by 3D Trajectories cites this paper.

Follow Your Track: Precise Skeleton Animation Controlled by 3D Trajectories VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 51

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T19:40:06.059882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-25T21:12:03.751682Z digest=sha256:c2bf53859e5b526cb34e864e933274372c262651ed298d14fa1b335a0b48c3bf

Observation 7ccc178c-bfdf-4d96-abcc-9ecd4f0e267a · inbound

Disco-LoRA: Disentangled Composition of Content, Style, and Motion for Multi-concept Video Customization cites this paper.

Disco-LoRA: Disentangled Composition of Content, Style, and Motion for Multi-concept Video Customization VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 84

Resolution
verified exact
local_arxiv, observed 2026-07-04T12:49:52.821413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T05:51:29.839498Z digest=sha256:d6504312086857f0abe2e8ecc9aab04ed0c4d15178dd26c98460b75db78273d6

Observation d97b9917-9216-4d53-9db2-d7863335fa32 · inbound

PhysEditWorld: A Large-Scale Dataset Toward Physics-Editable World Models cites this paper.

PhysEditWorld: A Large-Scale Dataset Toward Physics-Editable World Models VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-07-04T12:49:53.206917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T05:46:21.198781Z digest=sha256:3954da62a8dcf5f484c78bd4290d9b93e5cfc7197f0221e51f082f4a2d868253

Observation 6f303530-57df-4530-929a-545d03c29bd5 · inbound

PhysEditWorld: A Large-Scale Dataset Toward Physics-Editable World Models cites this paper.

PhysEditWorld: A Large-Scale Dataset Toward Physics-Editable World Models VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-06-30T12:04:39.241670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T10:19:06.268547Z digest=sha256:f164e3866930529a42547eb155c446ff7ccc613b3de60f07e338b984d25eccee

Observation 4aa509f5-a983-43a7-8e5e-aa12df6ef359 · inbound

RoboGaze: Evaluating Robot World Models via Structured Vision-Language Analysis cites this paper.

RoboGaze: Evaluating Robot World Models via Structured Vision-Language Analysis VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-06-30T10:54:36.522496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T10:49:58.019238Z digest=sha256:a252283018f45ea6b8fde599aaf3b43e6668c7f1253864427586fcfeb16ddaaf

Observation 7789858c-2357-44cf-9630-358abf4ccad8 · inbound

A Physics-Grounded Benchmark for Multi-Agent Dynamics in World Models cites this paper.

A Physics-Grounded Benchmark for Multi-Agent Dynamics in World Models VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 70

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T09:54:34.907927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T09:51:45.272494Z digest=sha256:bb82b1728c9459c095b55fcc0a4ae81b6abe7be7ba579bfc004c3aedae01645b

Observation abef0d2f-f577-4f6d-929d-8a5491541a98 · inbound

A Physics-Grounded Benchmark for Multi-Agent Dynamics in World Models cites this paper.

A Physics-Grounded Benchmark for Multi-Agent Dynamics in World Models VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 70

Resolution
unresolved
no resolver link, observed 2026-07-12T11:18:58.558989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:18:58.558989Z digest=sha256:e593b9f1a8b1a754c07cb11ced1bbc226c5e38a7593457322c60be040de6045f

Observation 079321c7-7923-40f4-b92e-32ad7e4cde6c · inbound

EcoVideo: Entropy-Orchestrated Video Generation Paradigm in Cloud-Edge Dynamics cites this paper.

EcoVideo: Entropy-Orchestrated Video Generation Paradigm in Cloud-Edge Dynamics VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 37

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T07:14:22.114346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T06:15:31.587641Z digest=sha256:edf40042ca011d0e650d590431cfea68e05f3eb0573f9c23faa98b238ae39a19

Observation 3297169d-3b07-4949-b7b4-90f970952944 · inbound

WorldRoamBench: An Open-World Benchmark for Long-Horizon Stability of Interactive World Models cites this paper.

WorldRoamBench: An Open-World Benchmark for Long-Horizon Stability of Interactive World Models VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-01T09:55:41.260286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T06:01:05.571698Z digest=sha256:b57d074977a8ba5f9d524fee14385bc080f16bca7ed75f7a87b119af1fadad51

Observation a9ec7e36-f7e0-440f-b9f0-f417525e6533 · inbound

WorldRoamBench: An Open-World Benchmark for Long-Horizon Stability of Interactive World Models cites this paper.

WorldRoamBench: An Open-World Benchmark for Long-Horizon Stability of Interactive World Models VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-03T21:58:58.308702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-03T21:58:29.499220Z digest=sha256:a7aeeabd7d5c6217a5a1f03b1d31b8e174e677db8f77f1df8d93d6d0cf82a0e4

Observation 4ac97cbb-237a-4999-b55e-0ced07080a90 · inbound

WorldRoamBench: An Open-World Benchmark for Long-Horizon Stability of Interactive World Models cites this paper.

WorldRoamBench: An Open-World Benchmark for Long-Horizon Stability of Interactive World Models VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-12T10:02:29.248139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T10:02:29.248139Z digest=sha256:e56c87dd934877756ec02acc5ef705ed97eac5501146aa80de1da997295c38fe

Observation 66210f22-da9b-480f-8ca8-0d1e4998e44b · inbound

RetailSMV: Exocentric vs. Egocentric Adaptation of Foundation Video World Models in Retail cites this paper.

RetailSMV: Exocentric vs. Egocentric Adaptation of Foundation Video World Models in Retail VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-02T15:37:05.917415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-02T15:33:09.686278Z digest=sha256:8bad39932ce135fda0a98fa6b07946d3dfc3aa99c1939fb73ff34aa7e0cef820

Observation ab1d21b1-ff9d-4138-8a97-836dedc3f9c9 · inbound

Towards Memory-Efficient Autoregressive Video Generation via Instance-Specific Parametric Absorption cites this paper.

Towards Memory-Efficient Autoregressive Video Generation via Instance-Specific Parametric Absorption VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 55

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T14:17:02.608226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-02T14:12:29.528065Z digest=sha256:aca288e0cbef31c4368f33d45be3c4ba0202de0ccb140679a9cf14c6bdffd59c