Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-02T14:42:23.038497Z
Paper Citation Record · LEDGER
As of 20 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 2 inbound Pith citation observations for arXiv:2607.00726.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-02T14:42:23.038497Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-08T00:17:53.934343Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-02T14:47:03.157753Z
31 of 31 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation b96bc9b3-2b1b-41c0-95d6-6ba4c5e5200e · outbound
AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 00d0b861-d2ea-4287-8bf9-fb913f143414 · outbound
AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization The framework defines audio–visual synchroniza- tion performance along two key dimensions: temporal consis- tency perception and semantic consistency perception
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 3a69b4b3-1634-4920-a985-fd45696be4fb · outbound
AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization Setup All experiments are conducted on two NVIDIA H20 GPUs, with each job allocated 4 vCPUs (Intel Xeon Platinum 8469C)
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 8d44d570-26fb-4d81-961e-ba08dbab8dfb · outbound
AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization First, the semantic editing tasks rely on gener- ative methods such as DDSP and OpenV oice V2
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 7c8dc27b-dd10-44c8-b960-bf10d554a38f · outbound
AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation abffc545-e574-4a8e-935d-edc9b3a35bab · outbound
AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization These tools did not contribute to the creation of any sci- entific content, data, or conclusions
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 0dfc2863-2403-4a50-a467-25835e2f291a · outbound
AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization Look, listen and learn,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 8c2ae7eb-f2a5-4ef8-8dfe-1f81eb112b12 · outbound
AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization Audio-visual scene analysis with self-supervised multisensory features,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation d1aca361-54fb-4a35-a8c4-86c0a801c501 · outbound
AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization Audio-visual event localization in unconstrained videos,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation e5050b42-d424-47ff-b4d7-0163ff4e08ad · outbound
AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization The sound of pixels,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation ab648870-6c1e-4c68-8d60-9dd12ff4e6ba · outbound
AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization Diff-foley: Synchronized video-to-audio synthesis with latent diffusion models,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 38afd828-36e0-42ce-a6bd-3870e8bbedc7 · outbound
AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization MMAudio: Taming multimodal joint training for high-quality video-to-audio synthesis,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation f19e6ef5-3530-4536-839b-c0668901d07b · outbound
AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization FreeAudio: Training-free timing planning for controllable long-form text-to- audio generation,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation ab8a855c-2e6f-48bb-8648-533d8ab41082 · outbound
AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization ControlAudio: Tackling Text-Guided, Timing-Indicated and Intelligible Audio Generation via Progressive Diffusion Modeling
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 70cc7f8b-7691-43e7-ba35-78253d719dae · outbound
AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization V ATT: Transformers for multimodal self- supervised learning from raw video, audio and text,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation c613800e-aaa5-4da7-96b5-25a94636db3c · outbound
AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization Qwen3-Omni Technical Report
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 5f2eef17-dc85-47ed-91de-8f40910f7232 · outbound
AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization Kling-foley: Multimodal diffusion transformer for high-quality video-to-audio generation,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation f59c7858-f8f8-4982-879b-4af53c945390 · outbound
AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization V2a-mapper: A lightweight solution for vision-to-audio generation by connecting foundation models
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation ddf727fa-1b29-447d-9dc5-311031373b2d · outbound
AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization Audio-Visual Synchronisation in the wild
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation b7077ea2-4d93-409e-90c3-4a86707f59b2 · outbound
AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization CLAP: Learning Audio Concepts From Natural Language Supervision
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 2b846a7b-5816-4628-a9ed-e0627b39a5b4 · outbound
AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization Imagebind: One embedding space to bind them all,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 3056c054-64e9-4ff7-878b-84f989c9470f · outbound
AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization Contrastive audio-visual masked autoencoder,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation d1ef6c23-6694-4236-a69c-c20940ac0a91 · outbound
AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization Sparse in space and time: Audio-visual synchronisation with trainable selectors,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation d3404882-9cb8-489e-a8c1-0f880128c825 · outbound
AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization Synchformer: Efficient synchronization from sparse cues,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation cf41efdb-72d0-45db-8434-ad33a2a3e5f6 · outbound
AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization DDSP: Differen- tiable digital signal processing,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation ef41d4c5-0a1f-491c-8b4c-1776e0b78ce9 · outbound
AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization Audio set: An ontology and human-labeled dataset for audio events
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 4fb5c988-c551-4201-b01b-d2aa47c65092 · outbound
AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization Vggsound: A large-scale audio-visual dataset,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 433f1be3-63b6-4695-a454-924ac0f43320 · outbound
AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization OpenVoice: Versatile Instant Voice Cloning
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 43dd8e30-9ac5-41b3-a90d-ace7de62fd71 · outbound
AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization OpenV oiceV2 model card,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 96a0de66-fe88-45c2-8657-9322d0017715 · outbound
AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization Cav-mae sync: Improving contrastive audio-visual mask au- toencoders via fine-grained alignment,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 66493f19-15a6-4695-bd8b-f4c0d5cd2318 · outbound
AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization Gemini 3 Flash: frontier intelligence built for speed,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation b96bc9b3-2b1b-41c0-95d6-6ba4c5e5200e · inbound
AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 27213d46-d234-409c-abef-b4577fd34cae · inbound
OmniVR: Joint Video-Audio Conditional Generation for Restoring Degraded Historical Films AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.