Pith. sign in

Paper Citation Record · LEDGER

AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation

As of 6 August 2026, this Paper Citation Record lists 4 of 4 outbound references and 6 inbound Pith citation observations for arXiv:2604.08540.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.08540 v1

Coverage vector

measured 4 of 4 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T17:14:30.737975Z

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T03:12:38.636242Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-03T00:07:28.033897Z

Reference resolution

4 of 4 outbound references displayed

  • verified exact1
  • verified fuzzy1
  • unresolved0
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 40a9edcf-d16f-4b87-a7ae-9fa34872c127 · outbound

This paper cites Accessed: 2024-02-15.

AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation Accessed: 2024-02-15

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:16:05.652003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T17:14:30.737975Z digest=sha256:5961df0cc840b9bbbc8847a8aba868caa0e8e94a1f3fb3ecfc124f823a9e4b3b

Observation 9000eaa8-df4c-4b49-ac77-9d85d0512ddc · outbound

This paper cites Seedance 1.5 pro: A Native Audio-Visual Joint Generation Foundation Model.

AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation Seedance 1.5 pro: A Native Audio-Visual Joint Generation Foundation Model

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T01:35:37.953991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T17:14:30.737975Z digest=sha256:8b3ec9bad5cbeeffc1c6f926fdbec538302f1ac55dfcfafad65436b053ebf447

Observation 7591e2a6-9cc8-45bc-a135-133d408a960e · outbound

This paper cites • Rationale:Determining ”which voice sounds more natural” is cognitively easier and more consistent via side-by-side comparison than absolute scoring.

AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation • Rationale:Determining ”which voice sounds more natural” is cognitively easier and more consistent via side-by-side comparison than absolute scoring

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T11:24:31.869996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T17:14:30.737975Z digest=sha256:a61526fb235ed530848b2aa2cae0beea74117cc31782231874b92c45b399dc85

Observation aea2f8a1-d7cf-43e9-a49f-357c122fbc55 · outbound

This paper cites A pairwise comparison might result in a ”Tie” if both models produce gibberish, failing to capture the absolute failure.

AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation A pairwise comparison might result in a ”Tie” if both models produce gibberish, failing to capture the absolute failure

Reference 4

Resolution
malformed identifier
raw_fallback, observed 2026-05-17T11:24:31.873429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T17:14:30.737975Z digest=sha256:b6a4888d286f4b3f67c37b44a83f0cb3d396a84171969e267c0296baf030e180

Pith citing papers

Observation c33338a2-d954-4107-b7b0-0dd9b784a88d · inbound

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation cites this paper.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation

Reference 78

Resolution
verified exact
local_arxiv, observed 2026-05-20T05:18:03.141127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T05:17:11.690484Z digest=sha256:eb2346abf170d8ff8e174d0ea1da8113f39d30d72e75c5f69939209e78088376

Observation b84a4872-0c4f-4c41-8d7e-56515dd02d73 · inbound

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation cites this paper.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation

Reference 78

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:04:58.018320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:0e20f0d337004cfa0a68150647c63f6816633b1de6263e2172e6f723195782cd

Observation d79a1bdf-ff4d-40bd-8840-ff993681a2cc · inbound

LongAV-Compass: Towards Unified Evaluation of Minute-Scale Audio-Visual Generation Across T2AV, I2AV, and V2AV cites this paper.

LongAV-Compass: Towards Unified Evaluation of Minute-Scale Audio-Visual Generation Across T2AV, I2AV, and V2AV AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:54:00.796273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T22:52:38.330851Z digest=sha256:6913c7b57bc3b6926afcb4795902a2f24f3f988aca8543c8b8df763c24f2a824

Observation d12b254e-ca36-4bfe-9343-83bba87a0e4f · inbound

MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video Generation cites this paper.

MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video Generation AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:03:26.186447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T12:58:56.555335Z digest=sha256:bc7aa74f4b157037425d49538224b9df2e50bab19ac5aa86aa6b0ecceed9f0ba

Observation 5608c485-f48a-466a-80eb-12584f0683b4 · inbound

CineDance: Towards Next-Generation Multi-Shot Long-Form Cinematic Audio-Video Generation cites this paper.

CineDance: Towards Next-Generation Multi-Shot Long-Form Cinematic Audio-Video Generation AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation

Reference 101

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T00:07:28.035378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T17:30:25.371658Z digest=sha256:0c47d470f78fb411c14d3369d1880d468b55a443eb5135defd2460b6ea3fc959

Observation 8644857c-5477-4bd8-acac-4277519854e5 · inbound

MultiRef-Compass: Towards Comprehensive Evaluation of Multi-Reference-to-Audio-Video Generation cites this paper.

MultiRef-Compass: Towards Comprehensive Evaluation of Multi-Reference-to-Audio-Video Generation AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T03:12:38.636242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:12:38.636242Z digest=sha256:defcacb0b68a6db99908f61a95c99cd2ce6e7427de5afdd9ee46eb07e1e47ad4