Pith. sign in

Paper Citation Record · LEDGER

MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video Generation

As of 7 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 0 inbound Pith citation observations for arXiv:2605.28035.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.28035 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-29T12:58:56.555335Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

43 of 43 outbound references displayed

  • verified exact31
  • verified fuzzy0
  • unresolved11
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d996737e-1647-4160-bc43-d7ba78b8ea99 · outbound

This paper cites Ming-Omni: A Unified Multimodal Model for Perception and Generation.

MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video Generation Ming-Omni: A Unified Multimodal Model for Perception and Generation

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:03:26.238741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T12:58:56.555335Z digest=sha256:9487577ce1d58e97de16f69d4185764a4e8538a04bd1978539f428401e77ffed

Observation 3fcfb654-6fd3-4980-9861-2054e302bd18 · outbound

This paper cites T2AV-Compass: Towards Unified Evaluation for Text-to-Audio-Video Generation.

MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video Generation T2AV-Compass: Towards Unified Evaluation for Text-to-Audio-Video Generation

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:03:26.236094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T12:58:56.555335Z digest=sha256:67e1c4c57b7f1c314d12ee342409f02bca3f15045161db1244646e92617aff84

Observation 6d1fa948-c93a-40e5-90b3-d3e1ff661f68 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video Generation VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:03:26.241165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T12:58:56.555335Z digest=sha256:c03740e1a9089562c86304309e4b6d61f16e0cb0de3bf1f472a4c82b02be0972

Observation 8079da02-09ee-4680-acaa-adbe46529f3c · outbound

This paper cites an unresolved cited work.

MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video Generation Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-29T12:58:56.555335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T12:58:56.555335Z digest=sha256:69ba6ca5eeb252c152a450e98ad9d85a2b472db71d6f6a9e6da06d8a1b298719

Observation cc94ef7f-e915-4333-9d95-97c365fd6276 · outbound

This paper cites Seedance 1.0: Exploring the Boundaries of Video Generation Models.

MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video Generation Seedance 1.0: Exploring the Boundaries of Video Generation Models

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:03:26.191814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T12:58:56.555335Z digest=sha256:baffb09661da31ec8f00d896d10cf90b93c7e6e4716241d2cc01e1c67bd21dc0

Observation 5bf2fb70-140e-438e-80db-2cd6529b1a93 · outbound

This paper cites Dreamid-omni: Unified framework for controllable human-centric audio-video generation.

MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video Generation Dreamid-omni: Unified framework for controllable human-centric audio-video generation

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:03:26.231024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T12:58:56.555335Z digest=sha256:1f2aef98357651dca2d2a4d2b34608d8c6f3bf51c32aeae229950a723e1b65c7

Observation 794bee71-30bf-4899-8050-23aa302004a6 · outbound

This paper cites an unresolved cited work.

MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video Generation Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-29T12:58:56.555335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T12:58:56.555335Z digest=sha256:b6eabadecf006844433a900d7967ed113009386d3aecaa9c06ccd21560e4d087

Observation 226454b4-b263-4e3c-9a84-d70393c2e086 · outbound

This paper cites LTX-2: Efficient Joint Audio-Visual Foundation Model.

MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video Generation LTX-2: Efficient Joint Audio-Visual Foundation Model

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:03:26.228415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T12:58:56.555335Z digest=sha256:c0cbfdddafd08b396ad3085abbe93baf349f311296aef35af05a889aa7461e02

Observation 9ed236eb-3aba-48ee-ba01-cef4c48cd8f6 · outbound

This paper cites an unresolved cited work.

MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video Generation Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-29T12:58:56.555335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T12:58:56.555335Z digest=sha256:45d786e275d8b2a8d78e8605dd519a7c8b3802ebbc338b693ccbc90b9462db15

Observation a780d5c7-ca67-449b-a4b1-70a483570ef7 · outbound

This paper cites Harmony: Harmonizing audio and video generation through cross-task synergy.

MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video Generation Harmony: Harmonizing audio and video generation through cross-task synergy

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:03:26.233544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T12:58:56.555335Z digest=sha256:0cf307c2dd3b49717808f34c2f149e1dc2de1a8b9f685bb43402a05b9900b0f2

Observation c14d5e62-6023-4812-bec0-614c49b065f7 · outbound

This paper cites VABench: A Comprehensive Benchmark for Audio-Video Generation.

MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video Generation VABench: A Comprehensive Benchmark for Audio-Video Generation

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:03:26.244286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T12:58:56.555335Z digest=sha256:207dcba9ae83c2943e300398670569ed067c3191cb64536f47d873d5a485cf40

Observation ad7114ac-5690-4e94-bb4a-3ec72a36e83e · outbound

This paper cites FilMaster: Bridging Cinematic Principles and Generative AI for Automated Film Generation.

MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video Generation FilMaster: Bridging Cinematic Principles and Generative AI for Automated Film Generation

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:03:26.223186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T12:58:56.555335Z digest=sha256:196013cc8f9078c5f9163654a5f0033d44e670443d82166939452a55d61b568c

Observation f7845cc7-776a-4722-9511-7b6a38dda052 · outbound

This paper cites an unresolved cited work.

MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video Generation Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-29T12:58:56.555335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T12:58:56.555335Z digest=sha256:e5f800ba4b2c3f3f8992dd69693d3a2cb2d8ae2eaaa76c4e933053b791688dac

Observation 3e802ede-811e-44fd-82a2-6f22cc443071 · outbound

This paper cites Pistachio: Towards Synthetic, Balanced, and Long-Form Video Anomaly Benchmarks.

MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video Generation Pistachio: Towards Synthetic, Balanced, and Long-Form Video Anomaly Benchmarks

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:03:26.219716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T12:58:56.555335Z digest=sha256:37c9ebc3c34dc7f11c530c5d2c91436fd85e2677855749b209ffcc21d3618c09

Observation c306501b-2e4a-429a-9e70-1d971f74f35c · outbound

This paper cites Videohallu: Evaluating and mitigating multi- modal hallucinations on synthetic video understanding.

MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video Generation Videohallu: Evaluating and mitigating multi- modal hallucinations on synthetic video understanding

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:03:26.215950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T12:58:56.555335Z digest=sha256:5e5d1b82726371e0f11b939a8557eba1ce06bf33d3ccf204cdb04ed306cbe013

Observation 2e682433-2b29-4acf-8a57-9cf0fad346ae · outbound

This paper cites VideoDirectorGPT: Consistent Multi-scene Video Generation via LLM-Guided Planning.

MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video Generation VideoDirectorGPT: Consistent Multi-scene Video Generation via LLM-Guided Planning

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:03:26.236751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T12:58:56.555335Z digest=sha256:b9e4f3dc4296255500f6c63e43c22e7e4aa675ca16d59b055841f4e18c9e64c7

Observation 90ea86ce-b87e-400d-8135-a28616103c0e · outbound

This paper cites arXiv preprint arXiv:2512.22905 , year=.

MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video Generation arXiv preprint arXiv:2512.22905 , year=

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:03:26.252283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T12:58:56.555335Z digest=sha256:06876da5f5b4d32b93387f935c6d2d03641b09dce8162df29a993ac20177db5a

Observation 4f470280-bd12-446a-8b70-6c1a34f32946 · outbound

This paper cites JavisDiT: Joint audio-video diffusion transformer with hierarchical spatio-temporal prior synchronization.

MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video Generation JavisDiT: Joint audio-video diffusion transformer with hierarchical spatio-temporal prior synchronization

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:03:26.249648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T12:58:56.555335Z digest=sha256:443178484e2b6d4a99b2b69b620f2bb9b71720874cb899758a4fae1ae8eed181

Observation 80b66b69-d84f-4f50-ae59-f7691fadfb03 · outbound

This paper cites Ilya Loshchilov and Frank Hutter.

MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video Generation Ilya Loshchilov and Frank Hutter

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:03:26.229730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T12:58:56.555335Z digest=sha256:6bafb87dbf868678fb49e96c6baa041d425c11f25fd770a9da93e06b1b68bc4c

Observation 1ed0ee24-4c98-4e50-97f0-f325d22baf69 · outbound

This paper cites Ola: Pushing the Frontiers of Omni-Modal Language Model.

MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video Generation Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:03:26.177260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T12:58:56.555335Z digest=sha256:53a2b2cd3e6fca41bac8dbe1999bb1954557cd646d007c6218a4f87e59e95470

Observation 7b1cb8d3-7c8c-4408-950c-54e13f96b9ae · outbound

This paper cites Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation.

MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video Generation Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:03:26.246924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T12:58:56.555335Z digest=sha256:7262ae004e7bef005a05f6c6a3c41cf8d1b8e455a0aeb81849267a26f4cde62b

Observation 0824869e-f536-44fd-9b9d-6f0dcabdf8bd · outbound

This paper cites an unresolved cited work.

MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video Generation Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-29T12:58:56.555335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T12:58:56.555335Z digest=sha256:1439cf67f269ac5666c0254e6e95696bf6d6c549279276633cc63d58f4ec507b

Observation b9d8c680-87be-4c13-b1dc-10df713092d4 · outbound

This paper cites The Script is All You Need: An Agentic Framework for Long-Horizon Dialogue-to-Cinematic Video Generation.

MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video Generation The Script is All You Need: An Agentic Framework for Long-Horizon Dialogue-to-Cinematic Video Generation

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:03:26.259600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T12:58:56.555335Z digest=sha256:d8960eaddc524abf5edb6547e0dd4cb8755fbd4d8e08b49d404e6a2e4448123b

Observation 2afe3338-8d00-447e-b115-9e3c55508205 · outbound

This paper cites an unresolved cited work.

MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video Generation Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-29T12:58:56.555335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T12:58:56.555335Z digest=sha256:5002d2deba9a2fd18cdc589626e058206144c0a3f8881cee6ae0837bf7ed68d2

Observation 7cd7a621-9897-4692-a21b-cbc81e41cab5 · outbound

This paper cites Msvbench: Towards human-level evaluation of multi-shot video generation.arXiv preprint arXiv:2602.23969.

MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video Generation Msvbench: Towards human-level evaluation of multi-shot video generation.arXiv preprint arXiv:2602.23969

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:03:26.222737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T12:58:56.555335Z digest=sha256:3feab545cb00a5811e797b66e6f68410c69a1d2d62fbbe007effec7668767eff

Observation b9845e67-36c4-437f-8a7c-b6dff77c68fd · outbound

This paper cites an unresolved cited work.

MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video Generation Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-06-29T12:58:56.555335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T12:58:56.555335Z digest=sha256:dc71ba9d892fce508c296f64463ef9555ade41cb274dc7417aa20998dbf47781

Observation eb939407-c196-466e-a1d2-1aa110cc10b8 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video Generation Gemini: A Family of Highly Capable Multimodal Models

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:03:26.257138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T12:58:56.555335Z digest=sha256:a2105f8e7ddb8f330729f77a513f371920a91658de38726b7b2ce90413e981de

Observation 218e493d-2697-403d-b04d-dc34ee948ff2 · outbound

This paper cites Mova: Towards scalable and synchronized video-audio generation.

MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video Generation Mova: Towards scalable and synchronized video-audio generation

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:03:26.198221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T12:58:56.555335Z digest=sha256:fdefc6e62891a514f57cb99e8433e33adf52d31ccb0d00e8269b9cc5da6ad1a6

Observation 98534943-92e7-479d-963b-19192e9337b9 · outbound

This paper cites UniVerse-1: Unified Audio-Video Generation via Stitching of Experts.

MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video Generation UniVerse-1: Unified Audio-Video Generation via Stitching of Experts

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:03:26.184000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T12:58:56.555335Z digest=sha256:a7d14f02ea75a57239679ebf515e86008d5736ee8648615c6db02770331384f0

Observation 69be12ea-83b4-4c50-837b-695bcb86fd61 · outbound

This paper cites CineTechBench: A Benchmark for Cinematographic Technique Understanding and Generation.

MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video Generation CineTechBench: A Benchmark for Cinematographic Technique Understanding and Generation

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:03:26.207486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T12:58:56.555335Z digest=sha256:837e86b7a960a77b8cff812d91cba07fce302ba6a241f1f1fe3530c7c8e198cc

Observation 05bbf75b-5655-4a5e-a6f5-65765e41d172 · outbound

This paper cites an unresolved cited work.

MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video Generation Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-06-29T12:58:56.555335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T12:58:56.555335Z digest=sha256:2335667a059ad0feb7528a01350e646930a46b2e164f587325652faddefe6638

Observation e536ebd5-a872-4879-9f74-ac5354683788 · outbound

This paper cites an unresolved cited work.

MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video Generation Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-06-29T12:58:56.555335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T12:58:56.555335Z digest=sha256:955a2015d82e6f4ee42d880068d526af5bab006c628e640963ec12380e995f17

Observation 30d9c036-7968-40e1-84b9-c98cba14d4b6 · outbound

This paper cites PhyAVBench: A Challenging Audio Physics-Sensitivity Benchmark for Physically Grounded Text-to-Audio-Video Generation.

MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video Generation PhyAVBench: A Challenging Audio Physics-Sensitivity Benchmark for Physically Grounded Text-to-Audio-Video Generation

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:03:26.254615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T12:58:56.555335Z digest=sha256:48d50ebf7b27e9bdce492f18796b992306e77f343ad00be6f9efde2621c1302b

Observation f7c462d1-fa92-40dc-aa23-d4ca554a93cc · outbound

This paper cites Qwen2.5-Omni Technical Report.

MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video Generation Qwen2.5-Omni Technical Report

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:03:26.210047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T12:58:56.555335Z digest=sha256:07b462fc17ef120872557b0d7adf8c645a9467433957c8bec91a3645ae989f09

Observation afbcdca5-3e41-4068-9988-0afc170b8d90 · outbound

This paper cites an unresolved cited work.

MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video Generation Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-06-29T12:58:56.555335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T12:58:56.555335Z digest=sha256:f737c46c56ab72decb8d846420a6de78718558a1c9a8b409d6ad6f6e3f760974

Observation 6d843190-f90b-4f3e-8762-132016adea2c · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video Generation MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:03:26.231759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T12:58:56.555335Z digest=sha256:20e9fb2dc4c11a03a939fbbe362682063198523a18dfd6871eb6b694954f97d6

Observation 2c7fbfed-cd5d-45ec-976f-788ac5a940a5 · outbound

This paper cites Omnivinci: Enhancing architecture and data for omni-modal understanding llm.

MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video Generation Omnivinci: Enhancing architecture and data for omni-modal understanding llm

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:03:26.227450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T12:58:56.555335Z digest=sha256:25fa03ff4acad37cbe1ba96158550a3d68c38bd662390f2a6d0cfb2c209366d7

Observation 5c810940-7eda-44d9-b55c-ab5127162797 · outbound

This paper cites an unresolved cited work.

MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video Generation Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-06-29T12:58:56.555335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T12:58:56.555335Z digest=sha256:6fa7e78cd485527c2ba780ea2aaa3750efcc5bc14463382783c27c6da6525ca3

Observation 04fa848d-a2bd-4fa8-a11f-8c0ac2962f6b · outbound

This paper cites arXiv preprint arXiv:2511.03334 (2025).

MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video Generation arXiv preprint arXiv:2511.03334 (2025)

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:03:26.189340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T12:58:56.555335Z digest=sha256:aaf3ff1c467dacce89fe81b0865ae06f158fd24c25c07a396971e1c55f305b8b

Observation 8d181fff-f4e3-42ea-a15c-87cf0c474c7c · outbound

This paper cites Stage: Storyboard- anchored generation for cinematic multi-shot narrative.

MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video Generation Stage: Storyboard- anchored generation for cinematic multi-shot narrative

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:03:26.239070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T12:58:56.555335Z digest=sha256:0eb51f055e0f8cce074f1f4817059dd3544043ca0bde924291e6cdb6d96adc5f

Observation c3751cfd-fbd3-4d96-a56c-fc2970ea0d74 · outbound

This paper cites MTAVG-Bench: A Diagnostic Benchmark for Multi-Talker Dialogue-Centric Audio-Video Generation.

MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video Generation MTAVG-Bench: A Diagnostic Benchmark for Multi-Talker Dialogue-Centric Audio-Video Generation

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:03:26.225709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T12:58:56.555335Z digest=sha256:506567cee5e287bbdab6037a987ea385dd29fd59dd5f61865522dd9955a413d9

Observation d12b254e-ca36-4bfe-9343-83bba87a0e4f · outbound

This paper cites AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation.

MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video Generation AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:03:26.186447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T12:58:56.555335Z digest=sha256:626a1c52bee22e7137898fa790d4b7830d47dc0acb257b56fe30dfa508fc69e0

Observation a9d52f46-6a57-4a80-ba83-d395486af217 · outbound

This paper cites clip_summary.

MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video Generation clip_summary

Reference 43

Resolution
malformed identifier
no resolver link, observed 2026-06-29T12:58:56.555335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T12:58:56.555335Z digest=sha256:ab0851088edeaa15abac9e39b5012c34d74b18fe4b855e32e07d42b4e01dca01

Pith citing papers

No inbound Pith citation observations are available.