Pith. sign in

Paper Citation Record · LEDGER

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation

As of 7 August 2026, this Paper Citation Record lists 80 of 80 outbound references and 1 inbound Pith citation observation for arXiv:2605.20183.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.20183 v4

Coverage vector

measured 80 of 80 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-30T18:00:39.556424Z

measured 81 of 81 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T13:41:47.011782Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

80 of 80 outbound references displayed

  • verified exact37
  • verified fuzzy38
  • unresolved3
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2b230e35-e930-4f43-b1b3-7e6da2703009 · outbound

This paper cites an unresolved cited work.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-07-08T06:34:42.264940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:600addec21818f8f3e8efb3860d50979f4f9e23d8f71f2ea6f9cbbafc07e0bdf

Observation 60dcb6f3-4bda-4109-8857-e13451b4e264 · outbound

This paper cites Qwen3-VL Technical Report.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation Qwen3-VL Technical Report

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:04:58.097549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:1ba34b4798e96621ed05029ed64a3f16063b7b4357497b0c7a9c01e70dfe213b

Observation dc38bd91-e344-4e39-9944-cf79927e1327 · outbound

This paper cites BlazePose: On-device Real-time Body Pose tracking.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation BlazePose: On-device Real-time Body Pose tracking

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:04:58.095017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:79eeaf765d76d6ea61400e5d423fdb17e9e807eea654167637394368f244c22b

Observation 21f2cbdd-2c38-495f-a64d-ae57099572eb · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:04:58.094505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:e08dd868020bbafa68e61660924b522f6dd0dc37fe7bdf5d4d415fe282060194

Observation 89bb3661-a6c0-4dc8-835d-be78492277c1 · outbound

This paper cites Video generation models as world simulators.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation Video generation models as world simulators

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T06:34:42.291974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:3ed2e2da840f786b737a9cdaab94ce75df7277071c99a0dbd457dcca90cd13bd

Observation 9b2e3eb9-681d-4a7f-9e8e-18b1b74b17a2 · outbound

This paper cites an unresolved cited work.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-07-08T06:34:42.286190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:00011740f231a192b5d76cf465f087e0ef5c514b1940e9b389d23d0b1591066b

Observation 754ea24a-d19d-4c7e-a89b-286438c28a05 · outbound

This paper cites T2AV-Compass: Towards Unified Evaluation for Text-to-Audio-Video Generation.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation T2AV-Compass: Towards Unified Evaluation for Text-to-Audio-Video Generation

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:04:58.096846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:6d59b0183afd16ac41160938b33084d95be94951784d5942056db4f80a4540e3

Observation 4a119891-80ba-4ef4-ab13-9580042d523f · outbound

This paper cites Talkvid: A large-scale diversified dataset for audio-driven talking head synthesis.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation Talkvid: A large-scale diversified dataset for audio-driven talking head synthesis

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T06:24:43.224310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:bea2f575f6da2783fee79f34a73205d16201d8d822b6fb793345721cff1e4be0

Observation 8df5aad2-007e-4e67-bacc-6409cb6c6e47 · outbound

This paper cites MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T06:24:43.222118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:c319f3d8eef41041c25e9c4ded9fde24baa60a42ea6ae7acb26d067d7f395d34

Observation 16d5c764-e452-4aa6-8e7a-115151217502 · outbound

This paper cites W2v-bert: Combining contrastive learning and masked language modeling for self-supervised speech pre-training.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation W2v-bert: Combining contrastive learning and masked language modeling for self-supervised speech pre-training

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T06:24:43.229203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:e77a5fbdef52e8d419e99e8b0b72599ba2008db14aff3d64f172bbfdf6429b9b

Observation 971b4ae8-8f37-4219-99dc-e5715983224b · outbound

This paper cites Paddleocr-vl-1.5: Towards a multi-task 0.9b vlm for robust in-the-wild document parsing.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation Paddleocr-vl-1.5: Towards a multi-task 0.9b vlm for robust in-the-wild document parsing

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T06:24:43.230936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:949bb576e9d11bf3f70111b915380be1eb67520dc6c13dc7e8cac1353aab2619

Observation 6e9bb927-2ead-4ebf-902d-7b7cb53db59d · outbound

This paper cites PaddleOCR-VL: Boosting Multilingual Document Parsing via a 0.9B Ultra-Compact Vision-Language Model.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation PaddleOCR-VL: Boosting Multilingual Document Parsing via a 0.9B Ultra-Compact Vision-Language Model

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T06:34:42.266730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:4d09c63eb245936e0eec2d74fa7dd0cd3b18ea8314ef3afe5a8b8e4d368b5a62

Observation 9889eef2-f610-432c-ad4f-605c7549969a · outbound

This paper cites Demucs: Deep Extractor for Music Sources with extra unlabeled data remixed.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation Demucs: Deep Extractor for Music Sources with extra unlabeled data remixed

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:04:58.105130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:f4c5ecd12c55870f155deffff635ee7d6773c02c626c8680b1ee2fafb227cff2

Observation 0bf5db51-6e5a-4452-8b96-e804ae2a2cab · outbound

This paper cites Arcface: Additive angular margin loss for deep face recognition.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation Arcface: Additive angular margin loss for deep face recognition

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T06:24:43.218628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:3a327992aa0f377d5f12eae6374a4a2aae794f782a7ce0f0bf16a152c88509db

Observation b428f046-a40d-422f-8123-17153205ff75 · outbound

This paper cites Scaling rectified flow trans- formers for high-resolution image synthesis.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation Scaling rectified flow trans- formers for high-resolution image synthesis

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T06:24:43.220292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:a278d466c9137b6bce33c3f1d9119c73450cd800c62d1b89534ecb14a802e444

Observation f146a850-2e1d-4f43-936c-a14f90ef8ade · outbound

This paper cites Veo 3.1.https://deepmind.google/technologies/veo/.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation Veo 3.1.https://deepmind.google/technologies/veo/

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T06:34:42.287904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:920e4349d372667da8f64c74410b76abead56330b8076ca3b774ca4a7b2c5d56

Observation 4d47b458-493a-4b3d-bd1a-a525431148f8 · outbound

This paper cites Audcast: Audio-driven human video generation by cascaded diffusion transformers.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation Audcast: Audio-driven human video generation by cascaded diffusion transformers

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T06:34:42.284560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:0b7309dfbc0e62d65f27d9a14da44f9e191fe6015b6b04b2da8f25ba33540fe5

Observation ce96b9b5-c5a2-493a-8a5a-dbaa67a6cb56 · outbound

This paper cites Dreamid-omni: Unified framework for controllable human-centric audio-video generation.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation Dreamid-omni: Unified framework for controllable human-centric audio-video generation

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:04:58.102413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:b703bfed6f92e9140a78dfd489176a60725d66d20969dbaeae2e46abc105c9a1

Observation 01c7872a-688d-4aa8-9813-101509c4654e · outbound

This paper cites LTX-Video: Realtime Video Latent Diffusion.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation LTX-Video: Realtime Video Latent Diffusion

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:04:58.052524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:3d83abf6e6f30b5e7fae66ffabed801e46d941b11ddad9a0cba051984c01ee67

Observation f270e092-163c-4e0e-af13-1536d4b9654a · outbound

This paper cites Video-bench: Human-aligned video generation benchmark.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation Video-bench: Human-aligned video generation benchmark

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T06:24:43.215199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:96edf530ae0e05844743396485bf83e562e6b732c179899511909499a2ae540d

Observation 073be41a-85d1-4974-833e-5255aec4d323 · outbound

This paper cites AesRM: Improving Video Aesthetics with Expert-Level Feedback.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation AesRM: Improving Video Aesthetics with Expert-Level Feedback

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:04:58.055663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:2f7fff43d26a2a750cc3f568e9346f877c79153b8c4de5b26cbf254ae18cf4b7

Observation 00575db7-25d9-44bf-a079-cf1226de5126 · outbound

This paper cites HappyHorse 1.0.https://www.happyhorse.cn/, 2026.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation HappyHorse 1.0.https://www.happyhorse.cn/, 2026

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T06:24:43.196554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:f65756c89efa980006245603b6bd933b2a0832838c95ebd1f5412c464b05d850

Observation 041217ff-a865-4c75-8a70-aa976cc281bd · outbound

This paper cites Denoising diffusion probabilistic models.Advances in neural information processing systems, 33:6840–6851.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation Denoising diffusion probabilistic models.Advances in neural information processing systems, 33:6840–6851

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T06:24:43.198478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:efdd8006f371ad3b55e2dbdef16c63ca5ec1c662ffc9ff7b929e26c303a7c029

Observation 68606a4d-3a45-4662-bf95-c21fb80612ec · outbound

This paper cites Video diffusion models.Advances in neural information processing systems, 35:8633–8646.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation Video diffusion models.Advances in neural information processing systems, 35:8633–8646

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T06:24:43.206334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:b319fc2c8494d84550f9afefde89fc97913f6ae590a069bcfc8eddbae0135d93

Observation 302a8e5a-5588-493b-b683-2bd3d6c299ae · outbound

This paper cites VABench: A Comprehensive Benchmark for Audio-Video Generation.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation VABench: A Comprehensive Benchmark for Audio-Video Generation

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:04:58.072872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:ecf0effef12ad42a01c830b6db8976893ae1985cbf5586382d0b1ba30640cc40

Observation 7708f7c2-1e8d-41ae-a40d-437e171f3f57 · outbound

This paper cites Self forcing: Bridging the train-test gap in autoregressive video diffusion.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation Self forcing: Bridging the train-test gap in autoregressive video diffusion

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T06:24:43.216908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:282c6db29d1416ef2661f8c82199de12f910874017dec60bf0d1a643f26a580f

Observation 114684d0-6017-4382-a988-593acce8256e · outbound

This paper cites Vbench: Comprehensive benchmark suite for video generative models.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation Vbench: Comprehensive benchmark suite for video generative models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T06:34:42.290197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:bda1248cccded70f16a05d138f926a1343094a9bf020c9dda5a7e51910a7a351

Observation 99a207e0-4c0a-4edc-ab6d-e306c3b7db01 · outbound

This paper cites Synchformer: Efficient synchronization from sparse cues.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation Synchformer: Efficient synchronization from sparse cues

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T06:24:43.189397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:960b165091421f191fa085e287872323796d0ac52f6404606953efb649581077

Observation 56b8a1c1-72b1-41e0-bcc2-815b993890b6 · outbound

This paper cites All-in-one metrical and functional structure analysis with neigh- borhood attentions on demixed audio.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation All-in-one metrical and functional structure analysis with neigh- borhood attentions on demixed audio

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T06:24:43.184046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:fc0e5c49a781b48b48c2835870bf52df16a252d7a0c47c2b0dbf3ce7465e8cd8

Observation c6e423b3-5615-4b58-90cf-2c9e8003cfed · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:04:58.037994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:8508c2ce40a9f598698419393ddf9356bbf34852e5d82aed5ceb271b3a323cfa

Observation 6dc58720-b1d5-463b-a9c7-54fdc3826dbd · outbound

This paper cites Kling 3.0.https://klingai.com/global/.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation Kling 3.0.https://klingai.com/global/

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T06:24:43.191440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:e49bf2ee950b7d4f856007456da819fbd4d2b5225b5deb7138439457f9e79fd1

Observation 7405bc37-39ad-4398-a4d2-37075cefa508 · outbound

This paper cites LatentSync: Taming Audio-Conditioned Latent Diffusion Models for Lip Sync with SyncNet Supervision.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation LatentSync: Taming Audio-Conditioned Latent Diffusion Models for Lip Sync with SyncNet Supervision

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:04:58.047082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:17d44c6c1b336407de5ff29aa2755cc89f9d2b0cf81e9764ff91cd64cdb15aab

Observation 3d8c1b93-4d58-4921-a35b-577ff1e5ab6f · outbound

This paper cites Lr-asd: Lightweight and robust network for active speaker detection.International Journal of Computer Vision, 133(7):4749–4769.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation Lr-asd: Lightweight and robust network for active speaker detection.International Journal of Computer Vision, 133(7):4749–4769

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T06:24:43.174275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:799a390c9ef65d4136f02e203ce97e8935bf95f68ffcdd2927b07d855128f048

Observation 0b81b2d1-7a63-4814-8ee0-07c2a7036a3e · outbound

This paper cites Aibench: Evaluating visual-logical consistency in academic illustration generation.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation Aibench: Evaluating visual-logical consistency in academic illustration generation

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:04:58.020908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:044f51f373f1296b1f455ddd13f218b6a75e28a7cc3199bb6b139b5b4b63c6bf

Observation 14eceddd-a168-40e3-88c8-ae6c26a3dfba · outbound

This paper cites Javisgpt: A unified multi-modal llm for sounding-video comprehension and generation.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation Javisgpt: A unified multi-modal llm for sounding-video comprehension and generation

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T06:24:43.193107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:6a26deb0ef8324615701cdfc63cbb116cb4a94f6b6d0faf56c36bc622305359e

Observation d438a6e9-944c-41b6-a699-7f2960996f33 · outbound

This paper cites Javisdit++: Unified modeling and optimization for joint audio-video generation.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation Javisdit++: Unified modeling and optimization for joint audio-video generation

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T06:34:42.282889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:b115d02fe9195bfb6d6708a3b0bdfd257c4c7be98f3c3fa803c23cc76d0edcf5

Observation 3780d891-86df-4711-9564-f762964e0800 · outbound

This paper cites Grounding dino: Marrying dino with grounded pre-training for open-set object detection.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation Grounding dino: Marrying dino with grounded pre-training for open-set object detection

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T06:34:42.276184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:733ab7533a81308d029a05158dd3ff344e529879f6880b3e1ba2d69e96d2507c

Observation f302f2bd-8f35-487b-a572-69bfce3f166f · outbound

This paper cites EvalCrafter: Benchmarking and Evaluating Large Video Generation Models.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation EvalCrafter: Benchmarking and Evaluating Large Video Generation Models

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:04:58.029339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:c2954c79ecc60a67b349f145015ddd62be2cd636139a5741a11eb00046728073

Observation a1ab789a-e2ae-46d6-b0c7-840700e0d2e4 · outbound

This paper cites Shotstream: Streaming multi-shot video generation for interactive storytelling.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation Shotstream: Streaming multi-shot video generation for interactive storytelling

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:04:58.046896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:703402eea9ed8991c42cbe3734db877b5a54ae0dddd95a4bb2c3c18762e034d7

Observation 9ec455b3-a782-443c-8102-1e1a1a4d1d7b · outbound

This paper cites Wan-Image: Pushing the Boundaries of Generative Visual Intelligence.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation Wan-Image: Pushing the Boundaries of Generative Visual Intelligence

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:04:58.049614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:646e23771fb6cd5e5b9f9d6e92ed96a747a00b90af48adb195218e904bd78aaa

Observation 2bf79fd5-8cc2-43ae-bdf3-aaa027d17dcc · outbound

This paper cites an unresolved cited work.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-07-08T06:34:42.277803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:419260b50a87b74b9a555a5f5b316b5d4fb7cb29f62261285ca7fca7cbcbb7d9

Observation f7f7f04c-81f1-4bf1-9c6f-a93cf7d3e0cb · outbound

This paper cites Sora 2.https://openai.com/index/sora-2/.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation Sora 2.https://openai.com/index/sora-2/

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T06:34:42.270291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:766ea6dbca2691e8d77d6b76f0202d22694239fd59fcc556b48f5cf3dd627b22

Observation 742ea77e-45c9-4797-a1c0-64838549f617 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation DINOv2: Learning Robust Visual Features without Supervision

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:04:58.052844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:9d49187ac0679d72742c6da531e451c4518aada16c38ad556a1345a50c6d30de

Observation 14d97904-3a22-40ed-920a-980f82fd1886 · outbound

This paper cites Sortformer: A Novel Approach for Permutation-Resolved Speaker Supervision in Speech-to-Text Systems.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation Sortformer: A Novel Approach for Permutation-Resolved Speaker Supervision in Speech-to-Text Systems

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:04:58.049750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:e7b151f172ea4058d1f2be628a5ccf20c50fdc643f0be549192550836505dfac

Observation f210a310-e110-4dc4-a3aa-840a8afb0616 · outbound

This paper cites Movie Gen: A Cast of Media Foundation Models.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation Movie Gen: A Cast of Media Foundation Models

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:04:58.110443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:0f2f81eace5cb52a54a4191ab7caf75de5a4e41add879e45231e56b57e56f02a

Observation 2867909c-fb66-4747-ae02-c697c0be5be9 · outbound

This paper cites Learning transferable visual models from natural language supervision.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation Learning transferable visual models from natural language supervision

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T06:34:42.274373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:c07b4b813abbeeaa5edaaaf1eee302d0521230e6ad2c6e9c9ce97834596f8f1d

Observation 1e1629f0-25b1-47e7-af1f-c73969ca6cdd · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation Robust speech recognition via large-scale weak supervision

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T06:34:42.268466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:88b0af6f0814bdb5a1f4d1be28bc56408eead9742d00d15a1dec3e0966e09c63

Observation 3cd957c4-9bb1-4c81-b13b-ce033e8ec2b7 · outbound

This paper cites Seedance 2.0: Advancing Video Generation for World Complexity.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation Seedance 2.0: Advancing Video Generation for World Complexity

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:04:58.092032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:5453a1c0bcda5289cca5cc4753f648703719ccf2d3fc3e92a59f87803aef5ec5

Observation 3388e2cd-6615-4fff-89af-d755b1db3d19 · outbound

This paper cites Hunyuanvideo-foley: Multimodal diffusion with representation alignment for high- fidelity foley audio generation.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation Hunyuanvideo-foley: Multimodal diffusion with representation alignment for high- fidelity foley audio generation

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T06:34:42.272452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:d5afeec440863ef4cbdfd3042db0b2a399daed5b10315c23348af5e19cc4a601

Observation ece795b4-93cc-4809-93a1-65ceb033aedd · outbound

This paper cites Msvbench: Towards human-level evaluation of multi-shot video generation.arXiv preprint arXiv:2602.23969.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation Msvbench: Towards human-level evaluation of multi-shot video generation.arXiv preprint arXiv:2602.23969

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:04:58.077935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:8e8049ad5386ca948568c538f02e2f35de07ac170fa893e1c4be1c8ef8f5bab2

Observation 35c7c9bc-5571-4b3f-ae55-959150ffd0d4 · outbound

This paper cites Speed by simplicity: A single-stream architecture for fast audio-video generative foundation model.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation Speed by simplicity: A single-stream architecture for fast audio-video generative foundation model

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:04:58.089417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:d0b51d716cf54e6829582f028d50d0391220659854332ffaa1f77be107bdaa5a

Observation 1d2fdd39-7e8f-49fa-9bfd-361a1638cc03 · outbound

This paper cites Make-A-Video: Text-to-Video Generation without Text-Video Data.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation Make-A-Video: Text-to-Video Generation without Text-Video Data

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:04:58.075200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:d035340a646b0fd677743439ef29806564ed9a6d79db4f04eaee33e60c951fb1

Observation da56686f-a0d3-49ff-90bf-e3f0db9ab855 · outbound

This paper cites Measuring Style Similarity in Diffusion Models.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation Measuring Style Similarity in Diffusion Models

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:04:58.067121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:6cec3250be99239aae00b5f04996cf880f929398056cd0325ab25166cda97209

Observation c7e224ab-51cd-4045-93fd-879cbdf319d2 · outbound

This paper cites TransNet V2: An effective deep network architecture for fast shot transition detection.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation TransNet V2: An effective deep network architecture for fast shot transition detection

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:04:58.064900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:b0390d8c2453f562d128e343fd0b6589b55a0558cdd4bac224771f5250f353b6

Observation a50ec9bc-5dd2-4d36-ae12-f4ffdfba8a6a · outbound

This paper cites The proof and measurement of association between two things.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation The proof and measurement of association between two things

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T06:34:42.279485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:66922def5deef3bf06dd706dae89fbf044c0fc4526195f0060ce1c6917d365e8

Observation 46694b87-8d32-4140-bd1b-1406ecfa9b59 · outbound

This paper cites Mova: Towards scalable and synchronized video-audio generation.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation Mova: Towards scalable and synchronized video-audio generation

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:04:58.075300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:d67f710287863d4bb62479560c5a2b1f77a9855e01235f5ad42e1441eb744892

Observation ce04f98c-2084-4f73-ada9-4bab4c7bcf0e · outbound

This paper cites Qwen3.5: Accelerating productivity with native multimodal agents, February 2026.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation Qwen3.5: Accelerating productivity with native multimodal agents, February 2026

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T06:24:43.213330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:6b4b8ecf60ed7dbf26234dfa7da03944e6e479bb982c730792f7dbb9a1077087

Observation 74e7f370-1e20-407a-9f27-a28b0dc5915c · outbound

This paper cites Silero V AD: pre-trained enterprise-grade V oice Activity Detector (V AD), Number Detector and Language Classifier.https://github.com/snakers4/silero-vad.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation Silero V AD: pre-trained enterprise-grade V oice Activity Detector (V AD), Number Detector and Language Classifier.https://github.com/snakers4/silero-vad

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T06:24:43.209041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:2518d0e70e9350eb67d3e61ea77af00c2ca1f69763238e0f0c5f1dd78454b846

Observation 3078c1c7-a593-4f1d-98ac-12f0209e031e · outbound

This paper cites Gemini 3.1 Pro.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation Gemini 3.1 Pro

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T06:24:43.211458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:c4c8eeabc9ae6d99cb0abfc530d4c3d8379cccb02f0aa189f9d068d3801f58a1

Observation 5c007a96-4831-49d4-bb8e-426a4d97b69c · outbound

This paper cites Meta audiobox aesthetics: Unified automatic quality assessment for speech, music, and sound.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation Meta audiobox aesthetics: Unified automatic quality assessment for speech, music, and sound

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T06:34:42.281165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:e43afff56dcee14db4d564b2f3dba1a62c182d8ef5937a9328562cf91606a7ac

Observation e72f9e0c-6812-43c5-a960-c66bad685702 · outbound

This paper cites Wan2.7.https://www.wan27.xyz/.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation Wan2.7.https://www.wan27.xyz/

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T06:24:43.182130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:b8cb3e42ec2b8bfb79b073c58157767a7faf5d31bb89848f6faaecb39c2ae386

Observation ac3dff55-d9b9-469a-a60c-0822fef3f0ac · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation Wan: Open and Advanced Large-Scale Video Generative Models

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:04:58.077837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:6a1c2ab404bcae3cd000100565d2a2efb3a7db0723670d12355e08e58fdfa7d7

Observation 1a2363ae-3ee2-4dfb-828a-5af00680156e · outbound

This paper cites AV-DiT: Efficient Audio-Visual Diffusion Transformer for Joint Audio and Video Generation.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation AV-DiT: Efficient Audio-Visual Diffusion Transformer for Joint Audio and Video Generation

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:04:58.083504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:d9a8f8224feb5f5122f36bcfd9bcc75650b6246c49e6149a5220803bbd4e1426

Observation 472ed063-ba2d-48f4-b396-c244c36ec198 · outbound

This paper cites Japanese Anime Scenes.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation Japanese Anime Scenes

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T06:24:43.177959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:2abbb4900e588bf4919d155281a835c09f42a1f0ccffc52e7a453873c5c5698e

Observation afa99c47-5b05-4a59-b007-1e072b72caf6 · outbound

This paper cites Univbench: Towards unified evaluation for video foundation models.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation Univbench: Towards unified evaluation for video foundation models

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:04:58.067412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:929b9a258226cace9208dee018bb4004f9d925575a179f36c47f58f9dafd544e

Observation ee84af24-707c-475f-ae99-ee32b5aa6994 · outbound

This paper cites Dreamvideo-omni: Omni-motion controlled multi-subject video customization with latent identity reinforcement learning.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation Dreamvideo-omni: Omni-motion controlled multi-subject video customization with latent identity reinforcement learning

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:04:58.092124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:73141dc23fa692fc6f17416471b881c4b636ed204bd42f6575049b1bfde83800

Observation 34354b89-8371-4632-8e5d-f3cb87d41dee · outbound

This paper cites Dreamvideo: Composing your dream videos with cus- tomized subject and motion.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation Dreamvideo: Composing your dream videos with cus- tomized subject and motion

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T06:24:43.180505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:d9c7c52e66eaa5d5473e27e04dab9572d915f082502a331f1a0db092f8bb8489

Observation d1197ce2-70e1-43b7-99e7-e586890d24e1 · outbound

This paper cites Dreamrelation: Relation-centric video customization.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation Dreamrelation: Relation-centric video customization

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T06:24:43.185713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:a8a68eb9f473f7541369b2eb454dfecefade74c5f29d8fcddc2c25fc351eb564

Observation 123c2115-1cb6-411f-8a16-b44a90d228ef · outbound

This paper cites Routing matters in moe: Scaling diffusion transformers with explicit routing guidance.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation Routing matters in moe: Scaling diffusion transformers with explicit routing guidance

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:04:58.024382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:e789a7a8e005f158106d063a459dbcae5fb15f4c2234dfcbdf28ef9dd019a642

Observation c3c0cfa5-9fa3-436b-9ee4-b49b9ff17124 · outbound

This paper cites DreamVideo-2: Zero-Shot Subject-Driven Video Customization with Precise Motion Control.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation DreamVideo-2: Zero-Shot Subject-Driven Video Customization with Precise Motion Control

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:04:58.029625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:183267e6db1cda8d97b22d2e0ec840b2fce5b1bd00e1fd84d3cd538e982c047c

Observation e1490f5b-e558-49c1-8ddb-242b9e46da01 · outbound

This paper cites PhyAVBench: A Challenging Audio Physics-Sensitivity Benchmark for Physically Grounded Text-to-Audio-Video Generation.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation PhyAVBench: A Challenging Audio Physics-Sensitivity Benchmark for Physically Grounded Text-to-Audio-Video Generation

Reference 71

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:04:58.058543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:7a4b7acfbb6b17362462dbb7beb61e24a87e623aa7919d465d52193e0f5766fe

Observation 9aa60b83-0184-4f5d-ac46-5132970b36af · outbound

This paper cites Fireredasr2s: A state-of-the-art industrial-grade all-in-one automatic speech recognition system.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation Fireredasr2s: A state-of-the-art industrial-grade all-in-one automatic speech recognition system

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:04:58.069731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:a6b67c9265de744290c609c5b93c7945cb0308c7680fcedb9d927e28ef87ccdf

Observation 0792d1d1-4dde-499b-a7d9-ee380f321dcb · outbound

This paper cites Longlive: Real-time interactive long video generation.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation Longlive: Real-time interactive long video generation

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T06:24:43.194918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:88194bd15e97d16b3190528d27f89a94c23932f4f4bb25523256836a5c5e1362

Observation 5717e354-4435-4328-84c4-ea7fe842b65e · outbound

This paper cites OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:04:58.012646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:aeb0371533a5a951b6041411de7635da9f2402a37f15a0775193aaf871fb9570

Observation 4528108d-60cc-4d8e-b4c6-64213e478887 · outbound

This paper cites Improved distribution matching distillation for fast image synthesis.Advances in neural information processing systems, 37:47455–47487, 2024a.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation Improved distribution matching distillation for fast image synthesis.Advances in neural information processing systems, 37:47455–47487, 2024a

Reference 75

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T18:04:58.107900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:c12b736c49bff83c549c9a9f5e34c75318d26d798a76b044f49bb3bfccd4acaf

Observation 38385add-b7cf-4a4e-ad03-983f4a143131 · outbound

This paper cites UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:04:58.099423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:261ada6fca629fb46437e020eb40e7ed0fbe33aee8643d7e464074f8824d6591

Observation 23c54114-88bb-4988-85a3-b94ce9e299b7 · outbound

This paper cites MTAVG-Bench: A Diagnostic Benchmark for Multi-Talker Dialogue-Centric Audio-Video Generation.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation MTAVG-Bench: A Diagnostic Benchmark for Multi-Talker Dialogue-Centric Audio-Video Generation

Reference 77

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:04:58.009856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:724429de48fc0e3ede3df61c1a760a14d8518756d8af2d0ef8a58c80b4d6cf32

Observation b84a4872-0c4f-4c41-8d7e-56515dd02d73 · outbound

This paper cites AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation

Reference 78

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:04:58.018320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:1049def0a741be0eb04190c37e4b43f68ff3549e2745cab855d3c10cda6a6be2

Observation d12d440f-07e9-413a-80ad-ba5903af1720 · outbound

This paper cites Causal forcing: Autoregressive diffusion distillation done right for high-quality real-time interactive video generation.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation Causal forcing: Autoregressive diffusion distillation done right for high-quality real-time interactive video generation

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T06:24:43.187597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:d9533b966bb36b38fc3994f373972639d4468496f145f01582dfa11fd6c1c5bf

Observation a52a847a-960b-491c-b25d-5c06a0899540 · outbound

This paper cites Vistorybench: Comprehensive benchmark suite for story visualization.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation Vistorybench: Comprehensive benchmark suite for story visualization

Reference 80

Resolution
malformed identifier
arxiv_id, observed 2026-06-30T18:04:58.027103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:2e1e6581d5c49862503c52f101840ff7e5fecd7d84ac3247da357e53650a370c

Pith citing papers

Observation 08202a3a-1b25-4bc4-b1b9-1b502ccb21f9 · inbound

FilmWorld: Agentic Novel-to-Film Generation through Dynamic Cinematic World Modeling cites this paper.

FilmWorld: Agentic Novel-to-Film Generation through Dynamic Cinematic World Modeling MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T13:41:47.011782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:41:47.011782Z digest=sha256:81e90fc7b5f7d6ccf90785756930cc9588b39508f9724f687baa98856a9eff4a