Pith. sign in

Paper Citation Record · LEDGER

AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models

As of 22 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 2 inbound Pith citation observations for arXiv:2605.24652.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.24652 v1

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-30T13:35:01.226818Z

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T00:17:53.910165Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-08T00:17:54.057786Z

Reference resolution

37 of 37 outbound references displayed

  • verified exact6
  • verified fuzzy13
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch15

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2e3e5160-e4fe-4abc-89a1-35dd252e9d24 · outbound

This paper cites T2AV-Compass: Towards Unified Evaluation for Text-to-Audio-Video Generation.

AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models T2AV-Compass: Towards Unified Evaluation for Text-to-Audio-Video Generation

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T13:44:41.266655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T13:35:01.226818Z digest=sha256:77a97375c48de4257bd6b2b7f26f3213c3f2b359722682f98f3ac25bfac44104

Observation 8996143b-dd65-46d4-8987-5f3bd63df3b3 · outbound

This paper cites JointAVBench: A Benchmark for Joint Audio-Visual Reasoning Evaluation.

AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models JointAVBench: A Benchmark for Joint Audio-Visual Reasoning Evaluation

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T13:44:41.242904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T13:35:01.226818Z digest=sha256:673ff1f3313f6831b9d06daa085b7bb7cc8fd6aaf7e1af2c1c11f55b6afea0c5

Observation 9c9f1201-4455-474a-99b1-53b841c23dbc · outbound

This paper cites In: 2025 IEEE 37th International Conference on Tools with Artificial Intelligence (ICTAI).

AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models In: 2025 IEEE 37th International Conference on Tools with Artificial Intelligence (ICTAI)

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T01:35:54.184424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T13:35:01.226818Z digest=sha256:688df87cdc789f01a61a1944f857a41ed88da039af32cfbc6d2d7c5a2303849b

Observation ecb2003d-3b8c-4c9f-bf19-2a13a20370de · outbound

This paper cites HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning.

AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T13:44:41.246048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T13:35:01.226818Z digest=sha256:21c7af20456cfd4dbe2beeea311070f66eb286f3e1347c5d4e57812d46803b02

Observation 59c1b992-f5dc-48b5-821b-75e6962f9b72 · outbound

This paper cites Qwen2-Audio Technical Report.

AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models Qwen2-Audio Technical Report

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T13:44:41.240538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T13:35:01.226818Z digest=sha256:de6a651e00e5328fdc7e40266711370dec25102a3eba3ba995ebe23e22128726

Observation f1b68f7e-bbba-447b-ad0c-550d348a1fab · outbound

This paper cites In: Work- shop on Multi-view Lip-reading, ACCV (2016).

AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models In: Work- shop on Multi-view Lip-reading, ACCV (2016)

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T01:35:54.186423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T13:35:01.226818Z digest=sha256:7c5cbd2e7e6e3b5eaf2b34b7a155193c8d6077f5096acbd650ab6f5ddcb7b269

Observation 292d3d76-eead-4c21-884f-0d81ba4e4d7e · outbound

This paper cites IEEE Open Journal of Signal Pro- cessing pp.

AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models IEEE Open Journal of Signal Pro- cessing pp

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:44:40.491286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T13:35:01.226818Z digest=sha256:698eae02b16a0f5241e864df6a52530029b379197a11cf499af47e48728194f6

Observation 2abe84ea-869c-4b58-8472-f95b90ed5640 · outbound

This paper cites In: ICASSP 2023-2023 IEEE Interna- tional Conference on Acoustics, Speech and Signal Processing (ICASSP).

AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models In: ICASSP 2023-2023 IEEE Interna- tional Conference on Acoustics, Speech and Signal Processing (ICASSP)

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T01:35:54.169391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T13:35:01.226818Z digest=sha256:b4fe5508199297b91632c90638f09ee19b94a88687876a7905532f6a102621e4

Observation d24ba60d-a8cc-4542-8fff-df0f0b0b24fd · outbound

This paper cites In: CVPR (2023).

AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models In: CVPR (2023)

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T01:35:54.190151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T13:35:01.226818Z digest=sha256:c5d11938277a777d17fc7fb70be51dba627a2d4729faf22ae03ffd5d8bab2838

Observation 7241cfab-f5bf-4f66-9a12-9227d12dcf2f · outbound

This paper cites an unresolved cited work.

AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-07-09T01:35:54.182453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T13:35:01.226818Z digest=sha256:8acb06d4978243d0fa6b74687e5a8d0e42188af6cd7a12c4a03c440d8d5be8aa

Observation e4624a68-a529-4f14-9881-c38c50d3f1c5 · outbound

This paper cites Dreamid-omni: Unified framework for controllable human-centric audio-video generation.

AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models Dreamid-omni: Unified framework for controllable human-centric audio-video generation

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T13:44:41.264384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T13:35:01.226818Z digest=sha256:26cc27c5346576b88c62c2f0e3bed9965a1a1badcaf58413814003feb245e9ec

Observation f5fbe09d-327d-4aac-b3e9-0b40f1dd756f · outbound

This paper cites In: Proceedings of the IEEE/CVF International Conference on Computer Vision.

AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models In: Proceedings of the IEEE/CVF International Conference on Computer Vision

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T01:35:54.167474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T13:35:01.226818Z digest=sha256:8e498c5efe51b512ae7b959edb9eaba5c7323709e553ac2e424a5fe7b9530406

Observation 3386859a-2bfe-4c4c-ab81-38d77215093d · outbound

This paper cites MMDisCo: Multi-Modal Discriminator-Guided Cooperative Diffusion for Joint Audio and Video Generation.

AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models MMDisCo: Multi-Modal Discriminator-Guided Cooperative Diffusion for Joint Audio and Video Generation

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T13:44:41.233058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T13:35:01.226818Z digest=sha256:f9368f105ec89a83e24992bc823695ee659ae9ca96d36737a7ca48ed44ceb2b6

Observation be1052fb-6d9a-4307-9cc8-bc1c562a24b3 · outbound

This paper cites VABench: A Comprehensive Benchmark for Audio-Video Generation.

AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models VABench: A Comprehensive Benchmark for Audio-Video Generation

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T13:44:41.269079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T13:35:01.226818Z digest=sha256:db1f64b6c4f98ff8dd8418cc8d8818223c391c5cdb9ee39c394728243f3a9560

Observation cc3bf316-8386-48a7-94cc-31e39530d1e3 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (2024).

AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (2024)

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T01:35:54.178651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T13:35:01.226818Z digest=sha256:72ea02b806d532756b23ac14359b8593b9b5954f5760b001fa5001f580e87390

Observation ea6fbebe-35f3-49b4-a253-ef814f737230 · outbound

This paper cites Taming Visually Guided Sound Generation.

AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models Taming Visually Guided Sound Generation

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:44:41.271841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T13:35:01.226818Z digest=sha256:0fb96f03d9009c74ceac400b8618421bb18ee49ca66c5e10c7391520f678024e

Observation b9fa4640-186b-4f4d-bf0e-6044e4b3a10b · outbound

This paper cites LatentSync: Taming Audio-Conditioned Latent Diffusion Models for Lip Sync with SyncNet Supervision.

AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models LatentSync: Taming Audio-Conditioned Latent Diffusion Models for Lip Sync with SyncNet Supervision

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T13:44:41.258619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T13:35:01.226818Z digest=sha256:9a211bac418187de12979d0e585174e95b6b7907aaa77b5582bb4babdf71bdc1

Observation 6900eba7-bfd3-4bba-bd7e-28e4a01cfcee · outbound

This paper cites In: Proceedings of the Computer Vision and Pattern Recognition Conference.

AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models In: Proceedings of the Computer Vision and Pattern Recognition Conference

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T01:35:54.192078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T13:35:01.226818Z digest=sha256:380173576fd79bde3ed2fbd89efa5796490a2049527cfd3153c019b24ae390c8

Observation 2113bfcd-1a7f-4f40-b01a-28e473799c9a · outbound

This paper cites Proceed- ings of the International Conference on Machine Learning pp.

AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models Proceed- ings of the International Conference on Machine Learning pp

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T01:35:54.165503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T13:35:01.226818Z digest=sha256:0fc635c86f94725b76e257f2c6c463c3f60ce25a797dbf5808f80f818dfa7e05

Observation eeac333f-e2bf-4d12-90f3-41074316eb55 · outbound

This paper cites Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation.

AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation

Reference 20

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T13:44:41.258595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T13:35:01.226818Z digest=sha256:07e51fa08bd77cc6ec461d7fd7d696e73c1672461e8c5c55f43e20b3b8681ec0

Observation ebdac6ca-0c7a-4d93-9fec-e08f87560b26 · outbound

This paper cites In: Interspeech (2021),https://api.semanticscholar.org/CorpusID:233296150.

AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models In: Interspeech (2021),https://api.semanticscholar.org/CorpusID:233296150

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T01:35:54.171271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T13:35:01.226818Z digest=sha256:354c80e4916e00bfbea0d717b7b742533c81fa61aa2988466c195c1ea0b287e7

Observation b4276ef5-ac85-495f-beaf-e8240d092c68 · outbound

This paper cites an unresolved cited work.

AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-07-09T01:35:54.176691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T13:35:01.226818Z digest=sha256:f04149f8540f3bd7b9e8fd8e07360abcfe1e924d76c69c67dd75d9f96fb5afc6

Observation e8f4db8c-3b8a-47b4-89b0-3c1004835392 · outbound

This paper cites an unresolved cited work.

AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-07-09T01:35:54.173027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T13:35:01.226818Z digest=sha256:cbda438f9992782f1ea934c2a372359011dea6cf7cf2073434642ce4b3ac28bd

Observation 3426f835-e79d-47a0-8e5c-1480ae042109 · outbound

This paper cites In: International conference on machine learning.

AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models In: International conference on machine learning

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T01:35:54.174785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T13:35:01.226818Z digest=sha256:017b4fa5ec5788a66b008dfac3986a5e78d1b18bcae20b16bfdc253fe32f4b37

Observation 223e80ae-1f4e-4695-b610-02facb2c57f8 · outbound

This paper cites LAION-5B: An open large-scale dataset for training next generation image-text models.

AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models LAION-5B: An open large-scale dataset for training next generation image-text models

Reference 25

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T13:44:41.253324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T13:35:01.226818Z digest=sha256:8e17991e90c8f8c0cab33b39df3f0786ccc021a857a5b3426e6a9832db28bba1

Observation 82d8deff-88c0-487b-9608-c46a012b312e · outbound

This paper cites Seedance 1.5 pro: A Native Audio-Visual Joint Generation Foundation Model.

AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models Seedance 1.5 pro: A Native Audio-Visual Joint Generation Foundation Model

Reference 26

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T13:44:41.253589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T13:35:01.226818Z digest=sha256:7bf46798c4b3c97a13163cd5c1f834b4100266e91f5cd9b7382c467841de7f55

Observation 51e221fe-c2db-4c46-acb7-4321c1340019 · outbound

This paper cites HunyuanVideo-Foley: Multimodal Diffusion with Representation Alignment for High-Fidelity Foley Audio Generation.

AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models HunyuanVideo-Foley: Multimodal Diffusion with Representation Alignment for High-Fidelity Foley Audio Generation

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:44:41.250735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T13:35:01.226818Z digest=sha256:9e34b2f907b0f3a262d00c65ab8c462c5106a6506bbc0d34b00d3bd499583810

Observation 6c1d55db-e97a-4f06-b3ca-42f193b15c9e · outbound

This paper cites Audiobox: Unified Audio Generation with Natural Language Prompts.

AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:44:41.261751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T13:35:01.226818Z digest=sha256:3d395adfb0ccd35b05e79b520d91afdae46444d1b45d1183708292fe9372297d

Observation 200f3fc9-e83f-4985-b3ea-918ae710aed9 · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models Wan: Open and Advanced Large-Scale Video Generative Models

Reference 29

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T13:44:41.274161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T13:35:01.226818Z digest=sha256:0494a68d817c2e460259c6387b61508189e9ccdb5f27c2b8e4d10ec827209ff1

Observation c6d103ef-583b-41b8-88a6-e5e711cd2edb · outbound

This paper cites arXiv preprint arXiv:2601.04151 (2026).

AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models arXiv preprint arXiv:2601.04151 (2026)

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:44:41.242882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T13:35:01.226818Z digest=sha256:607fc12ff2fd8c68354e0bf2ea1351ce2cd3ff004cd7967b5e38a1ab49d8beb7

Observation a34e84b6-6b7c-415e-a663-a463f083e867 · outbound

This paper cites InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation.

AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 31

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T13:44:41.227663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T13:35:01.226818Z digest=sha256:31d4dd02b3788aea73775a70e105591185465d0edd258a45f804d80180e9e6d6

Observation aa13d13b-664e-4639-9772-7c04aa43b297 · outbound

This paper cites In: Proceedings of the IEEE/CVF international conference on computer vision.

AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models In: Proceedings of the IEEE/CVF international conference on computer vision

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T01:35:54.188343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T13:35:01.226818Z digest=sha256:c9c2fda1d8fab75a62ab40545c66ff43c694bc3855cc059435c4fe4fc39790b8

Observation 513fd16c-25df-4e8a-b300-b6af867d39a7 · outbound

This paper cites Qwen2.5-Omni Technical Report.

AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models Qwen2.5-Omni Technical Report

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-06-30T13:44:41.232632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T13:35:01.226818Z digest=sha256:675d197630121fec832c0dc9cd2902611730695ff6cf0e60d36c197ccab836a6

Observation 92e1a51d-d954-4a38-96c4-9afc9a0a6028 · outbound

This paper cites Qwen3-Omni Technical Report.

AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models Qwen3-Omni Technical Report

Reference 34

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T13:44:41.234807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T13:35:01.226818Z digest=sha256:88e31703b5cb59c48e93b5a032c5e2db9402262fce98ee31d3b1e89544499ca8

Observation 639c7393-d3f3-4561-8a00-0cdaebfa3cce · outbound

This paper cites Qwen3 Technical Report.

AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models Qwen3 Technical Report

Reference 35

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T13:44:41.235165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T13:35:01.226818Z digest=sha256:ee7be2576f5fc788bee14833975bd20a8b3203cefa89d82e63a040c989acfef2

Observation f171955d-ed2e-423f-be91-f16c0f96b876 · outbound

This paper cites MISMATCHED.

AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models MISMATCHED

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T01:35:54.194052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T13:35:01.226818Z digest=sha256:ac030fda22f8909106a3be66c1eab514371f3eaedddd1ee088379a76110b9205

Observation 6ce298de-ace0-48f5-948b-841b9cfeba1b · outbound

This paper cites woman" with.

AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models woman" with

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T01:35:54.180547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T13:35:01.226818Z digest=sha256:beb66a5c1d0ecccdd0de4af9cdbed476835316a746fa7789e4b49b95da9aa6a6

Pith citing papers

Observation cac83026-c1cb-4992-9284-e277ce7c22a2 · inbound

MultiRef-Compass: Towards Comprehensive Evaluation of Multi-Reference-to-Audio-Video Generation cites this paper.

MultiRef-Compass: Towards Comprehensive Evaluation of Multi-Reference-to-Audio-Video Generation AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T03:12:37.865575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:12:37.865575Z digest=sha256:799cdddea52f377c91fa7b64a3a0c95fc332bea9042b899a61d62521e27633ee

Observation 17b66798-4277-4afc-8754-d4da0fed1034 · inbound

OmniVR: Joint Video-Audio Conditional Generation for Restoring Degraded Historical Films cites this paper.

OmniVR: Joint Video-Audio Conditional Generation for Restoring Degraded Historical Films AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models

Reference 42

Resolution
metadata mismatch
local_arxiv, observed 2026-08-08T00:17:54.065292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-08T00:17:53.910165Z digest=sha256:585866e93d4155ef479974972955e335dabb778497dbcf29ee2324256491ec13