Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-30T13:35:01.226818Z
Paper Citation Record · LEDGER
As of 22 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 2 inbound Pith citation observations for arXiv:2605.24652.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-30T13:35:01.226818Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-08T00:17:53.910165Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-08T00:17:54.057786Z
37 of 37 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 2e3e5160-e4fe-4abc-89a1-35dd252e9d24 · outbound
AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models T2AV-Compass: Towards Unified Evaluation for Text-to-Audio-Video Generation
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 8996143b-dd65-46d4-8987-5f3bd63df3b3 · outbound
AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models JointAVBench: A Benchmark for Joint Audio-Visual Reasoning Evaluation
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 9c9f1201-4455-474a-99b1-53b841c23dbc · outbound
AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models In: 2025 IEEE 37th International Conference on Tools with Artificial Intelligence (ICTAI)
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ecb2003d-3b8c-4c9f-bf19-2a13a20370de · outbound
AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 59c1b992-f5dc-48b5-821b-75e6962f9b72 · outbound
AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models Qwen2-Audio Technical Report
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation f1b68f7e-bbba-447b-ad0c-550d348a1fab · outbound
AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models In: Work- shop on Multi-view Lip-reading, ACCV (2016)
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 292d3d76-eead-4c21-884f-0d81ba4e4d7e · outbound
AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models IEEE Open Journal of Signal Pro- cessing pp
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 2abe84ea-869c-4b58-8472-f95b90ed5640 · outbound
AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models In: ICASSP 2023-2023 IEEE Interna- tional Conference on Acoustics, Speech and Signal Processing (ICASSP)
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation d24ba60d-a8cc-4542-8fff-df0f0b0b24fd · outbound
AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models In: CVPR (2023)
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 7241cfab-f5bf-4f66-9a12-9227d12dcf2f · outbound
AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models Unresolved cited work
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation e4624a68-a529-4f14-9881-c38c50d3f1c5 · outbound
AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models Dreamid-omni: Unified framework for controllable human-centric audio-video generation
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation f5fbe09d-327d-4aac-b3e9-0b40f1dd756f · outbound
AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models In: Proceedings of the IEEE/CVF International Conference on Computer Vision
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 3386859a-2bfe-4c4c-ab81-38d77215093d · outbound
AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models MMDisCo: Multi-Modal Discriminator-Guided Cooperative Diffusion for Joint Audio and Video Generation
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation be1052fb-6d9a-4307-9cc8-bc1c562a24b3 · outbound
AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models VABench: A Comprehensive Benchmark for Audio-Video Generation
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation cc3bf316-8386-48a7-94cc-31e39530d1e3 · outbound
AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (2024)
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ea6fbebe-35f3-49b4-a253-ef814f737230 · outbound
AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models Taming Visually Guided Sound Generation
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b9fa4640-186b-4f4d-bf0e-6044e4b3a10b · outbound
AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models LatentSync: Taming Audio-Conditioned Latent Diffusion Models for Lip Sync with SyncNet Supervision
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 6900eba7-bfd3-4bba-bd7e-28e4a01cfcee · outbound
AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models In: Proceedings of the Computer Vision and Pattern Recognition Conference
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 2113bfcd-1a7f-4f40-b01a-28e473799c9a · outbound
AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models Proceed- ings of the International Conference on Machine Learning pp
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation eeac333f-e2bf-4d12-90f3-41074316eb55 · outbound
AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ebdac6ca-0c7a-4d93-9fec-e08f87560b26 · outbound
AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models In: Interspeech (2021),https://api.semanticscholar.org/CorpusID:233296150
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b4276ef5-ac85-495f-beaf-e8240d092c68 · outbound
AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models Unresolved cited work
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation e8f4db8c-3b8a-47b4-89b0-3c1004835392 · outbound
AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models Unresolved cited work
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 3426f835-e79d-47a0-8e5c-1480ae042109 · outbound
AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models In: International conference on machine learning
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 223e80ae-1f4e-4695-b610-02facb2c57f8 · outbound
AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models LAION-5B: An open large-scale dataset for training next generation image-text models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 82d8deff-88c0-487b-9608-c46a012b312e · outbound
AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models Seedance 1.5 pro: A Native Audio-Visual Joint Generation Foundation Model
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 51e221fe-c2db-4c46-acb7-4321c1340019 · outbound
AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models HunyuanVideo-Foley: Multimodal Diffusion with Representation Alignment for High-Fidelity Foley Audio Generation
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 6c1d55db-e97a-4f06-b3ca-42f193b15c9e · outbound
AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 200f3fc9-e83f-4985-b3ea-918ae710aed9 · outbound
AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models Wan: Open and Advanced Large-Scale Video Generative Models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation c6d103ef-583b-41b8-88a6-e5e711cd2edb · outbound
AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models arXiv preprint arXiv:2601.04151 (2026)
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation a34e84b6-6b7c-415e-a663-a463f083e867 · outbound
AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation aa13d13b-664e-4639-9772-7c04aa43b297 · outbound
AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models In: Proceedings of the IEEE/CVF international conference on computer vision
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 513fd16c-25df-4e8a-b300-b6af867d39a7 · outbound
AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models Qwen2.5-Omni Technical Report
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 92e1a51d-d954-4a38-96c4-9afc9a0a6028 · outbound
AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models Qwen3-Omni Technical Report
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 639c7393-d3f3-4561-8a00-0cdaebfa3cce · outbound
AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models Qwen3 Technical Report
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation f171955d-ed2e-423f-be91-f16c0f96b876 · outbound
AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models MISMATCHED
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 6ce298de-ace0-48f5-948b-841b9cfeba1b · outbound
AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models woman" with
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation cac83026-c1cb-4992-9284-e277ce7c22a2 · inbound
MultiRef-Compass: Towards Comprehensive Evaluation of Multi-Reference-to-Audio-Video Generation AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17b66798-4277-4afc-8754-d4da0fed1034 · inbound
OmniVR: Joint Video-Audio Conditional Generation for Restoring Degraded Historical Films AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.