Pith. sign in

Paper Citation Record · LEDGER

AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization

As of 20 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 2 inbound Pith citation observations for arXiv:2607.00726.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.00726 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-02T14:42:23.038497Z

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T00:17:53.934343Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-02T14:47:03.157753Z

Reference resolution

31 of 31 outbound references displayed

  • verified exact7
  • verified fuzzy21
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b96bc9b3-2b1b-41c0-95d6-6ba4c5e5200e · outbound

This paper cites AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization.

AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T14:47:03.159103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-02T14:42:23.038497Z digest=sha256:eadb904380c128e97d8398c7257063a026b5377d6ca81ab9a3a47a9cd4884b58

Observation 00d0b861-d2ea-4287-8bf9-fb913f143414 · outbound

This paper cites The framework defines audio–visual synchroniza- tion performance along two key dimensions: temporal consis- tency perception and semantic consistency perception.

AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization The framework defines audio–visual synchroniza- tion performance along two key dimensions: temporal consis- tency perception and semantic consistency perception

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-02T14:47:03.167535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-02T14:42:23.038497Z digest=sha256:e781a67e423eafdd758c6ab41abdee097f6213cfbd23422bba3c7ffc4cbceb03

Observation 3a69b4b3-1634-4920-a985-fd45696be4fb · outbound

This paper cites Setup All experiments are conducted on two NVIDIA H20 GPUs, with each job allocated 4 vCPUs (Intel Xeon Platinum 8469C).

AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization Setup All experiments are conducted on two NVIDIA H20 GPUs, with each job allocated 4 vCPUs (Intel Xeon Platinum 8469C)

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-02T14:47:03.173478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-02T14:42:23.038497Z digest=sha256:e770a142e855df0ba3d05b8a2cdbd98f7f527b078c99c27520bbaf1590aaedba

Observation 8d44d570-26fb-4d81-961e-ba08dbab8dfb · outbound

This paper cites First, the semantic editing tasks rely on gener- ative methods such as DDSP and OpenV oice V2.

AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization First, the semantic editing tasks rely on gener- ative methods such as DDSP and OpenV oice V2

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T00:11:39.956528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-02T14:42:23.038497Z digest=sha256:209d733d3c21f9b3ec8183f51925683ceb0abedaca4c919adb44c7487770bb4b

Observation 7c8dc27b-dd10-44c8-b960-bf10d554a38f · outbound

This paper cites an unresolved cited work.

AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-07-06T00:11:39.960685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-02T14:42:23.038497Z digest=sha256:17caed3d2398b179db86b6aed108a1d74f01c1c37e4948b9fef8fbedd2ba6dd2

Observation abffc545-e574-4a8e-935d-edc9b3a35bab · outbound

This paper cites These tools did not contribute to the creation of any sci- entific content, data, or conclusions.

AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization These tools did not contribute to the creation of any sci- entific content, data, or conclusions

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T00:11:39.920655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-02T14:42:23.038497Z digest=sha256:8af694d3d92f58ede1ef4adaa30de4aa1d45f5bf86e7dc7404cd723fa8aed3e4

Observation 0dfc2863-2403-4a50-a467-25835e2f291a · outbound

This paper cites Look, listen and learn,.

AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization Look, listen and learn,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T00:11:39.946485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-02T14:42:23.038497Z digest=sha256:2b8094430e54bb5df42193c35088e298657e1beea3d9af53a8ca09a3773aac29

Observation 8c2ae7eb-f2a5-4ef8-8dfe-1f81eb112b12 · outbound

This paper cites Audio-visual scene analysis with self-supervised multisensory features,.

AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization Audio-visual scene analysis with self-supervised multisensory features,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T00:11:39.948681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-02T14:42:23.038497Z digest=sha256:4cc4cc45807d729c3fa1d60267d7932d00c41decc31181004e3fb080670d1f21

Observation d1aca361-54fb-4a35-a8c4-86c0a801c501 · outbound

This paper cites Audio-visual event localization in unconstrained videos,.

AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization Audio-visual event localization in unconstrained videos,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T00:11:39.950693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-02T14:42:23.038497Z digest=sha256:2788cd24fe62bfb908628f1c8954a9c667848a9ece90bb6c00d7567b2e5a9e03

Observation e5050b42-d424-47ff-b4d7-0163ff4e08ad · outbound

This paper cites The sound of pixels,.

AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization The sound of pixels,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T00:11:39.944727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-02T14:42:23.038497Z digest=sha256:0343602bb98f15f552109b2914dc4f5f34ab9bb2377196234324eafbd423327f

Observation ab648870-6c1e-4c68-8d60-9dd12ff4e6ba · outbound

This paper cites Diff-foley: Synchronized video-to-audio synthesis with latent diffusion models,.

AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization Diff-foley: Synchronized video-to-audio synthesis with latent diffusion models,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T00:11:39.942884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-02T14:42:23.038497Z digest=sha256:7a9473ca140fe732bc49e56702d0d61948341486f6af961c9e16d709f8056b50

Observation 38afd828-36e0-42ce-a6bd-3870e8bbedc7 · outbound

This paper cites MMAudio: Taming multimodal joint training for high-quality video-to-audio synthesis,.

AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization MMAudio: Taming multimodal joint training for high-quality video-to-audio synthesis,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T00:11:39.938805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-02T14:42:23.038497Z digest=sha256:6724bda8cd8e50e75dabeb8b070905365261ab330e7d68011cccf629cc74b1ab

Observation f19e6ef5-3530-4536-839b-c0668901d07b · outbound

This paper cites FreeAudio: Training-free timing planning for controllable long-form text-to- audio generation,.

AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization FreeAudio: Training-free timing planning for controllable long-form text-to- audio generation,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T00:11:39.940853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-02T14:42:23.038497Z digest=sha256:4e98eedf78d1c56ef2e751195b90d0823c9b837c94b77aa1e8b9f41400fc0373

Observation ab8a855c-2e6f-48bb-8648-533d8ab41082 · outbound

This paper cites ControlAudio: Tackling Text-Guided, Timing-Indicated and Intelligible Audio Generation via Progressive Diffusion Modeling.

AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization ControlAudio: Tackling Text-Guided, Timing-Indicated and Intelligible Audio Generation via Progressive Diffusion Modeling

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-02T14:47:03.176945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-02T14:42:23.038497Z digest=sha256:fc10c833e5b57444fe145501faeea8f8968bc1b690289c6bf6bd8125701eaa79

Observation 70cc7f8b-7691-43e7-ba35-78253d719dae · outbound

This paper cites V ATT: Transformers for multimodal self- supervised learning from raw video, audio and text,.

AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization V ATT: Transformers for multimodal self- supervised learning from raw video, audio and text,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T00:11:39.952652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-02T14:42:23.038497Z digest=sha256:ee5ea7d3facdadef81ede38bc7eb3ec6dac7136998202f4260357a858ac97892

Observation c613800e-aaa5-4da7-96b5-25a94636db3c · outbound

This paper cites Qwen3-Omni Technical Report.

AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization Qwen3-Omni Technical Report

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-02T14:47:03.174375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-02T14:42:23.038497Z digest=sha256:9e2aed09eafe7dee8667e4a0a3feb0d56efb6cf55a9e6fdd43769e32f776b425

Observation 5f2eef17-dc85-47ed-91de-8f40910f7232 · outbound

This paper cites Kling-foley: Multimodal diffusion transformer for high-quality video-to-audio generation,.

AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization Kling-foley: Multimodal diffusion transformer for high-quality video-to-audio generation,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T00:11:39.958922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-02T14:42:23.038497Z digest=sha256:5d625b25b5ac9d2bf4f49d577e85f61aaa5cca52c4caa6525595e2df692ed00f

Observation f59c7858-f8f8-4982-879b-4af53c945390 · outbound

This paper cites V2a-mapper: A lightweight solution for vision-to-audio generation by connecting foundation models.

AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization V2a-mapper: A lightweight solution for vision-to-audio generation by connecting foundation models

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T14:47:03.170542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-02T14:42:23.038497Z digest=sha256:cb8af0a5ceddfdf070c113976275a35b2146ba4972f0f1c6fe75a3ddaf35653b

Observation ddf727fa-1b29-447d-9dc5-311031373b2d · outbound

This paper cites Audio-Visual Synchronisation in the wild.

AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization Audio-Visual Synchronisation in the wild

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-02T14:47:03.165821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-02T14:42:23.038497Z digest=sha256:154969b8c9014820e1137ba384658179ecc2db337f3cb7be5ee727bd0d0c3307

Observation b7077ea2-4d93-409e-90c3-4a86707f59b2 · outbound

This paper cites CLAP: Learning Audio Concepts From Natural Language Supervision.

AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization CLAP: Learning Audio Concepts From Natural Language Supervision

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-02T14:47:03.176468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-02T14:42:23.038497Z digest=sha256:4302ab39f486626207c7e5e6b092a9d9b91775bf870b671aa871d141b3743149

Observation 2b846a7b-5816-4628-a9ed-e0627b39a5b4 · outbound

This paper cites Imagebind: One embedding space to bind them all,.

AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization Imagebind: One embedding space to bind them all,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T00:11:39.963662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-02T14:42:23.038497Z digest=sha256:a1767d3b0f67d0359d8441987af72fd146c63833de99250804c11fae800b258b

Observation 3056c054-64e9-4ff7-878b-84f989c9470f · outbound

This paper cites Contrastive audio-visual masked autoencoder,.

AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization Contrastive audio-visual masked autoencoder,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T00:11:39.922708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-02T14:42:23.038497Z digest=sha256:1cbaaec0b2dee22c73c0a6c42bb3e5324d659f94463ba2535faaee0015bd14bb

Observation d1ef6c23-6694-4236-a69c-c20940ac0a91 · outbound

This paper cites Sparse in space and time: Audio-visual synchronisation with trainable selectors,.

AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization Sparse in space and time: Audio-visual synchronisation with trainable selectors,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T00:11:39.931007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-02T14:42:23.038497Z digest=sha256:34f47c2239f68a567d645a450dfd591ca0e1b3c5a18088a97afa9859a50b5ca8

Observation d3404882-9cb8-489e-a8c1-0f880128c825 · outbound

This paper cites Synchformer: Efficient synchronization from sparse cues,.

AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization Synchformer: Efficient synchronization from sparse cues,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T00:11:39.932945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-02T14:42:23.038497Z digest=sha256:f19b0dadee6ac941b88cec4f94b88c0c7e60d5409dfe650b2c67dd8ff9302059

Observation cf41efdb-72d0-45db-8434-ad33a2a3e5f6 · outbound

This paper cites DDSP: Differen- tiable digital signal processing,.

AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization DDSP: Differen- tiable digital signal processing,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T00:11:39.936929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-02T14:42:23.038497Z digest=sha256:6b0ff80f5d00bb9c10e09e8195dba815277a8b58c15b90ad243f6eea19c15c6a

Observation ef41d4c5-0a1f-491c-8b4c-1776e0b78ce9 · outbound

This paper cites Audio set: An ontology and human-labeled dataset for audio events.

AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization Audio set: An ontology and human-labeled dataset for audio events

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T00:11:39.926725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-02T14:42:23.038497Z digest=sha256:42e4bd22bfbb551f16f5500353a62f4bfc847f1f9db21d8f500a9a23ba9f67b6

Observation 4fb5c988-c551-4201-b01b-d2aa47c65092 · outbound

This paper cites Vggsound: A large-scale audio-visual dataset,.

AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization Vggsound: A large-scale audio-visual dataset,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T00:11:39.929149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-02T14:42:23.038497Z digest=sha256:75b7b4657b52f8415b3882dbd8a45aa9c5b23fd6b349241c294477e6cfd8559b

Observation 433f1be3-63b6-4695-a454-924ac0f43320 · outbound

This paper cites OpenVoice: Versatile Instant Voice Cloning.

AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization OpenVoice: Versatile Instant Voice Cloning

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-02T14:47:03.159614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-02T14:42:23.038497Z digest=sha256:b215168e68011c6dd4a5e09b23a5fa9d193c1ab4e5febd1f6950ceca5a2b3905

Observation 43dd8e30-9ac5-41b3-a90d-ace7de62fd71 · outbound

This paper cites OpenV oiceV2 model card,.

AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization OpenV oiceV2 model card,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T00:11:39.954531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-02T14:42:23.038497Z digest=sha256:ebd648e479965dc6b27db665afed91124834bed795a9a5fca8f1d9aabb7a5c90

Observation 96a0de66-fe88-45c2-8657-9322d0017715 · outbound

This paper cites Cav-mae sync: Improving contrastive audio-visual mask au- toencoders via fine-grained alignment,.

AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization Cav-mae sync: Improving contrastive audio-visual mask au- toencoders via fine-grained alignment,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T00:11:39.924698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-02T14:42:23.038497Z digest=sha256:5c59ffbea9912ab917c5615420e441069300150456a12831158f7586a5979afe

Observation 66493f19-15a6-4695-bd8b-f4c0d5cd2318 · outbound

This paper cites Gemini 3 Flash: frontier intelligence built for speed,.

AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization Gemini 3 Flash: frontier intelligence built for speed,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T00:11:39.934978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-02T14:42:23.038497Z digest=sha256:c9d2e27a93e2922ee19e0217b6f249242799e34e83f62cddc16ad0b225c589dd

Pith citing papers

Observation b96bc9b3-2b1b-41c0-95d6-6ba4c5e5200e · inbound

AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization cites this paper.

AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T14:47:03.159103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-02T14:42:23.038497Z digest=sha256:eadb904380c128e97d8398c7257063a026b5377d6ca81ab9a3a47a9cd4884b58

Observation 27213d46-d234-409c-abef-b4577fd34cae · inbound

OmniVR: Joint Video-Audio Conditional Generation for Restoring Degraded Historical Films cites this paper.

OmniVR: Joint Video-Audio Conditional Generation for Restoring Degraded Historical Films AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-08T00:17:53.934343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:17:53.934343Z digest=sha256:2de7875bea7394fd454c584bff0a3c836e422d9d40705e3923d59c0c3b5d69a9