Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-26T23:17:45.299833Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 1 inbound Pith citation observation for arXiv:2606.23712.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-26T23:17:45.299833Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-26T23:17:45.299833Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-03T22:49:01.384665Z
35 of 35 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 5deb219c-beb9-49e1-b4f4-5d425e4dabe3 · outbound
Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement Recent deep learn- ing approaches have significantly improved performance [1–3], with generative modeling frameworks emerging as a powerful direction
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89dbd48a-c87c-4bca-acaa-4f190bf0b482 · outbound
Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a9516dbd-c7dc-4cc8-9b84-34a969b732b5 · outbound
Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement 1, describing how the audio-visual con- trastive loss is computed and the motivation behind it
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29ed98d9-149b-47af-ad22-c702fad92144 · outbound
Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement As baselines, we con- sider A V-DiffUSEEN, with cross-attention fusion [8, 13], the audio-only version, AO-DiffUSEEN [13], and the supervised- generative FlowA VSE model [7]
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78d6f946-dbcc-498e-a989-35221d479999 · outbound
Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement Matched condition: TCD-DEMAND Under matched conditions (Table 1), the proposed model im- proves all metrics compared to A V-DiffUSEEN
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 524e131a-c194-485f-a7c3-9ed5ce3046b2 · outbound
Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement During the pretraining of the visual-conditioned speech diffusion model, we augment the denoising score matching ob- jective with a contrastive audio-visual alignment loss
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8dccfb8-f91f-46f7-82d7-19fd41abd647 · outbound
Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement Unresolved cited work
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation abe0aa72-a4df-4085-a5ff-2157cc113e88 · outbound
Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement Generative AI tools were only used to edit and polish some portions of the manuscript
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc98de02-c35e-4519-9557-50233c8b473a · outbound
Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement SEGAN: Speech enhancement generative adversarial network,
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42c0a82b-ce2c-4996-9b44-0c0f8463d353 · outbound
Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement Conv-TasNet: Surpassing ideal time– frequency magnitude masking for speech separation,
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7836dbb8-7d74-40fb-ae77-af0bb744b29d · outbound
Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement DCCRN: Deep complex convolution recurrent network for phase-aware speech enhancement,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e54e6c0-1e8e-4fe8-9274-037b1169f730 · outbound
Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement Speech enhancement and dereverberation with diffusion-based gen- erative models,
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ccded98c-1c52-4c7c-a183-7c800ab26646 · outbound
Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement The conversation: Deep audio-visual speech enhancement,
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e168e88-24da-43a7-9283-ded77279913f · outbound
Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement Audio-visual speech enhancement using multimodal deep con- volutional neural networks,
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fa0d510-f391-4a3f-97a8-4915a7be2514 · outbound
Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement FlowA VSE: Efficient audio-visual speech enhancement with conditional flow matching,
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c7e79e7-2730-4bf5-a57d-4f0ba8828a25 · outbound
Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement Diffusion-based unsupervised audio-visual speech enhancement,
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00df5926-de7b-48df-95f8-0b002ed8d94a · outbound
Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement An overview of deep-learning-based audio-visual speech enhancement and separation,
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b129cf7-9d46-4a10-83dc-ebecdac0bbec · outbound
Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement Learning transfer- able visual models from natural language supervision,
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ed6430b-a876-42ef-a7b3-fc3397cc4d8c · outbound
Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement CLAP: Learn- ing audio concepts from natural language supervision,
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 032e5418-8cae-463e-87a7-1922dc297312 · outbound
Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement SA V-SE: Scene-aware audio-visual speech enhancement with selective state space model,
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9107a07-13aa-4db5-aaa1-dd0767c09fdb · outbound
Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement Diffusion-based Frameworks for Unsupervised Speech Enhancement
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9ec4f8fc-b3ab-4292-9e12-fb2a8417dbf2 · outbound
Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement Score-based generative modeling through stochastic differ- ential equations,
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ad42cd0-8bd9-49ef-ac26-3299a6263287 · outbound
Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement Nonnegative matrix factor- ization with the itakura-saito divergence: With application to music analysis,
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4637457c-8421-4014-a872-d69d9cc1cda5 · outbound
Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement Tweedie’s formula and selection bias,
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eea11571-501f-4479-acdc-02b3503f7786 · outbound
Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement Deep residual learning for im- age recognition,
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 524a6f4f-0fc6-4ba4-a4ca-47e4a843d284 · outbound
Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7d69518e-b9cd-4010-9d25-aa3d18057d95 · outbound
Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement Representation Learning with Contrastive Predictive Coding
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fcbf7481-a563-4ee6-a647-5048a5bc8b76 · outbound
Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement TCD-TIMIT: An audio-visual corpus of con- tinuous speech,
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 320a8621-46b7-4b0d-8d90-173f4644e06e · outbound
Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement The diverse environments multi- channel acoustic noise database (DEMAND): A database of multi- channel environmental noise recordings,
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63597d9c-061d-4644-930c-d95d8a30ea95 · outbound
Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement LRS3-TED: a large-scale dataset for visual speech recognition
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7016b475-36d9-4462-bdd4-24d74b2ec68b · outbound
Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement NTCD-TIMIT: A new database and base- line for noise-robust audio-visual speech recognition
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 709ebb50-95b1-4e58-b620-921464db7ada · outbound
Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement SDR–half- baked or well done?
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bcd495df-920c-485f-85e3-ddb03a86f8fc · outbound
Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement Performance measure- ment in blind audio source separation,
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 272ef1c0-95a9-4e6f-a1f0-8d00467d0975 · outbound
Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement Perceptual evaluation of speech quality (PESQ)-a new method for speech quality assessment of telephone networks and codecs,
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0983f300-b362-4d03-a785-f905f04a78b2 · outbound
Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement An algo- rithm for intelligibility prediction of time–frequency weighted noisy speech,
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89dbd48a-c87c-4bca-acaa-4f190bf0b482 · inbound
Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.