Pith. sign in

Paper Citation Record · LEDGER

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement

As of 7 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 1 inbound Pith citation observation for arXiv:2606.23712.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.23712 v1

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-26T23:17:45.299833Z

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-26T23:17:45.299833Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-03T22:49:01.384665Z

Reference resolution

35 of 35 outbound references displayed

  • verified exact4
  • verified fuzzy0
  • unresolved29
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5deb219c-beb9-49e1-b4f4-5d425e4dabe3 · outbound

This paper cites Recent deep learn- ing approaches have significantly improved performance [1–3], with generative modeling frameworks emerging as a powerful direction.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement Recent deep learn- ing approaches have significantly improved performance [1–3], with generative modeling frameworks emerging as a powerful direction

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:4d125c59411005ab760b5dbc37fa4211e36fa9d403b40ac196b83088f7031923

Observation 89dbd48a-c87c-4bca-acaa-4f190bf0b482 · outbound

This paper cites Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T22:49:01.387089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:87c4853963e199c81fc96940bcd3e8df0efbcf3c554d1e802a6e98520529b571

Observation a9516dbd-c7dc-4cc8-9b84-34a969b732b5 · outbound

This paper cites 1, describing how the audio-visual con- trastive loss is computed and the motivation behind it.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement 1, describing how the audio-visual con- trastive loss is computed and the motivation behind it

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:7131629ab46d8d5e5ce418484f66a3e9841b1f2e8c5bba6d56f19af388181512

Observation 29ed98d9-149b-47af-ad22-c702fad92144 · outbound

This paper cites As baselines, we con- sider A V-DiffUSEEN, with cross-attention fusion [8, 13], the audio-only version, AO-DiffUSEEN [13], and the supervised- generative FlowA VSE model [7].

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement As baselines, we con- sider A V-DiffUSEEN, with cross-attention fusion [8, 13], the audio-only version, AO-DiffUSEEN [13], and the supervised- generative FlowA VSE model [7]

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:ca8518939458ac03e15caf9655bfadeb8a641218ffa940cb8e91f84b0c84582e

Observation 78d6f946-dbcc-498e-a989-35221d479999 · outbound

This paper cites Matched condition: TCD-DEMAND Under matched conditions (Table 1), the proposed model im- proves all metrics compared to A V-DiffUSEEN.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement Matched condition: TCD-DEMAND Under matched conditions (Table 1), the proposed model im- proves all metrics compared to A V-DiffUSEEN

Reference 5

Resolution
malformed identifier
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:bd7136a29782621996c05ba2e5942ec7455f56d4b36b881ab49cba9304c22450

Observation 524e131a-c194-485f-a7c3-9ed5ce3046b2 · outbound

This paper cites During the pretraining of the visual-conditioned speech diffusion model, we augment the denoising score matching ob- jective with a contrastive audio-visual alignment loss.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement During the pretraining of the visual-conditioned speech diffusion model, we augment the denoising score matching ob- jective with a contrastive audio-visual alignment loss

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:3eeb26736c0d52b1efcaffc16db85473ca0cb8bc385ad6c66524f50eee48cf40

Observation a8dccfb8-f91f-46f7-82d7-19fd41abd647 · outbound

This paper cites an unresolved cited work.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:6c7b99cd956752ef7d8841574294367c250c30a661f29fd63219bbd4f043eab1

Observation abe0aa72-a4df-4085-a5ff-2157cc113e88 · outbound

This paper cites Generative AI tools were only used to edit and polish some portions of the manuscript.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement Generative AI tools were only used to edit and polish some portions of the manuscript

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:7f9487fccfba62675dd8263fee943b2ffb767236216f3c9b2f9fc6601388ca45

Observation bc98de02-c35e-4519-9557-50233c8b473a · outbound

This paper cites SEGAN: Speech enhancement generative adversarial network,.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement SEGAN: Speech enhancement generative adversarial network,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:80e1293cf50a43da3ea57c3832d0954ea0576e3fb4b480f7e7a3c4f1f9c3b83c

Observation 42c0a82b-ce2c-4996-9b44-0c0f8463d353 · outbound

This paper cites Conv-TasNet: Surpassing ideal time– frequency magnitude masking for speech separation,.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement Conv-TasNet: Surpassing ideal time– frequency magnitude masking for speech separation,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:6cd3cfc267188422c16aa88928bc99e7b549f890965cb29ebab1ebd5733c2d65

Observation 7836dbb8-7d74-40fb-ae77-af0bb744b29d · outbound

This paper cites DCCRN: Deep complex convolution recurrent network for phase-aware speech enhancement,.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement DCCRN: Deep complex convolution recurrent network for phase-aware speech enhancement,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:982d08eebbe36d6a50d761cbc4f9dc4592f6902223f1affda3e025b7f8427490

Observation 7e54e6c0-1e8e-4fe8-9274-037b1169f730 · outbound

This paper cites Speech enhancement and dereverberation with diffusion-based gen- erative models,.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement Speech enhancement and dereverberation with diffusion-based gen- erative models,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:ac3f4a23728422c575369ea295fa439caabc4b02a45fa31979aeda7150cef081

Observation ccded98c-1c52-4c7c-a183-7c800ab26646 · outbound

This paper cites The conversation: Deep audio-visual speech enhancement,.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement The conversation: Deep audio-visual speech enhancement,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:5574511e33434c78bc9bc80664b0720bcede66da9836b1f579d8d356aef5d631

Observation 6e168e88-24da-43a7-9283-ded77279913f · outbound

This paper cites Audio-visual speech enhancement using multimodal deep con- volutional neural networks,.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement Audio-visual speech enhancement using multimodal deep con- volutional neural networks,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:2857e4eebebb3d3b38550bd1eecb02baf4888fc2693689fd438ea5d52f1b258e

Observation 2fa0d510-f391-4a3f-97a8-4915a7be2514 · outbound

This paper cites FlowA VSE: Efficient audio-visual speech enhancement with conditional flow matching,.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement FlowA VSE: Efficient audio-visual speech enhancement with conditional flow matching,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:53e84aafccb96319ccfab414f585338bea853d7da22d385566a19d7a49649a92

Observation 0c7e79e7-2730-4bf5-a57d-4f0ba8828a25 · outbound

This paper cites Diffusion-based unsupervised audio-visual speech enhancement,.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement Diffusion-based unsupervised audio-visual speech enhancement,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:d165d0f48e96b4fcd8d107e4dce5b6d75486d7b1fe207a66570ffb00cbadadab

Observation 00df5926-de7b-48df-95f8-0b002ed8d94a · outbound

This paper cites An overview of deep-learning-based audio-visual speech enhancement and separation,.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement An overview of deep-learning-based audio-visual speech enhancement and separation,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:303030881aa8b8da575a056f64aa2bc603e64589c603cb2e188e6ea73f83735a

Observation 3b129cf7-9d46-4a10-83dc-ebecdac0bbec · outbound

This paper cites Learning transfer- able visual models from natural language supervision,.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement Learning transfer- able visual models from natural language supervision,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:36fd81d30c730522eecddf535494a4472a7eb331da4a4b6cd45787fcc94c9a49

Observation 3ed6430b-a876-42ef-a7b3-fc3397cc4d8c · outbound

This paper cites CLAP: Learn- ing audio concepts from natural language supervision,.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement CLAP: Learn- ing audio concepts from natural language supervision,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:e06f6cd08ea7625e36b10e7a9946a7e926835ac93665cec7127114bf0b4774b5

Observation 032e5418-8cae-463e-87a7-1922dc297312 · outbound

This paper cites SA V-SE: Scene-aware audio-visual speech enhancement with selective state space model,.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement SA V-SE: Scene-aware audio-visual speech enhancement with selective state space model,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:434941894ed514d5f900b7fb61ad26d6042c5c03eb9cbd249a708245b96e1bd5

Observation f9107a07-13aa-4db5-aaa1-dd0767c09fdb · outbound

This paper cites Diffusion-based Frameworks for Unsupervised Speech Enhancement.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement Diffusion-based Frameworks for Unsupervised Speech Enhancement

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-07-03T22:49:01.372695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:7702eb7ce4c14c107be7da5986437edabf1d74615b821decf9fc6b23612f9701

Observation 9ec4f8fc-b3ab-4292-9e12-fb2a8417dbf2 · outbound

This paper cites Score-based generative modeling through stochastic differ- ential equations,.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement Score-based generative modeling through stochastic differ- ential equations,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:18b04882661d9f19230a77bf8789f7fb829904ea880580dbe286d718d0d4cd7b

Observation 4ad42cd0-8bd9-49ef-ac26-3299a6263287 · outbound

This paper cites Nonnegative matrix factor- ization with the itakura-saito divergence: With application to music analysis,.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement Nonnegative matrix factor- ization with the itakura-saito divergence: With application to music analysis,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:40de59fb3211e2351032b7b98d444a9fe56049b7817b67088493ebe54d086be2

Observation 4637457c-8421-4014-a872-d69d9cc1cda5 · outbound

This paper cites Tweedie’s formula and selection bias,.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement Tweedie’s formula and selection bias,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:db7a3531c081caaa96ad631e80f5397a1ac306df771b9103944284874260d018

Observation eea11571-501f-4479-acdc-02b3503f7786 · outbound

This paper cites Deep residual learning for im- age recognition,.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement Deep residual learning for im- age recognition,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:657dac0bd6db7cd89d1e0552134d06be65d9dd4fd6a9c9a9bd8f9005d45be393

Observation 524a6f4f-0fc6-4ba4-a4ca-47e4a843d284 · outbound

This paper cites Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-03T22:49:01.368464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:46c37b32f1c248e0540c53b442ab54f8d7bbc0e77f40df7fb5f25d2be616f9b1

Observation 7d69518e-b9cd-4010-9d25-aa3d18057d95 · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement Representation Learning with Contrastive Predictive Coding

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-07-03T22:49:01.377870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:fd3933534a8104e39e5869e49f22f7fb2c0e31432b3c484d56332307152f9788

Observation fcbf7481-a563-4ee6-a647-5048a5bc8b76 · outbound

This paper cites TCD-TIMIT: An audio-visual corpus of con- tinuous speech,.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement TCD-TIMIT: An audio-visual corpus of con- tinuous speech,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:767823c10a272da2bc6d85f55ac18786fbd41a20d6edc87da2b6262ce86d5447

Observation 320a8621-46b7-4b0d-8d90-173f4644e06e · outbound

This paper cites The diverse environments multi- channel acoustic noise database (DEMAND): A database of multi- channel environmental noise recordings,.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement The diverse environments multi- channel acoustic noise database (DEMAND): A database of multi- channel environmental noise recordings,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:dfb7d5e968c82cf569c59d34ad11b6c3f80a05ad7c6013934e8a1ccb41c359ac

Observation 63597d9c-061d-4644-930c-d95d8a30ea95 · outbound

This paper cites LRS3-TED: a large-scale dataset for visual speech recognition.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement LRS3-TED: a large-scale dataset for visual speech recognition

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-07-03T22:49:01.382024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:f6c45eca63019005f341e906247a1b2bae7674e74fa93306bd63f1100051fd34

Observation 7016b475-36d9-4462-bdd4-24d74b2ec68b · outbound

This paper cites NTCD-TIMIT: A new database and base- line for noise-robust audio-visual speech recognition.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement NTCD-TIMIT: A new database and base- line for noise-robust audio-visual speech recognition

Reference 31

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:b9d6e7fe7fdb426305eda4f3fdb505fee8897a00466424971458f0a82e0b80db

Observation 709ebb50-95b1-4e58-b620-921464db7ada · outbound

This paper cites SDR–half- baked or well done?.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement SDR–half- baked or well done?

Reference 32

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:c0de6a42bc26260b4f2f20d7d8da7f890cc9287347dc4fc58bafa6af622857b1

Observation bcd495df-920c-485f-85e3-ddb03a86f8fc · outbound

This paper cites Performance measure- ment in blind audio source separation,.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement Performance measure- ment in blind audio source separation,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:180f92700a66d7f4163dadb98533d8a804c49776fc0c68acf53dde1d81989b10

Observation 272ef1c0-95a9-4e6f-a1f0-8d00467d0975 · outbound

This paper cites Perceptual evaluation of speech quality (PESQ)-a new method for speech quality assessment of telephone networks and codecs,.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement Perceptual evaluation of speech quality (PESQ)-a new method for speech quality assessment of telephone networks and codecs,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:b646984b3339f24c9a0a934ff62d8987875838603c4ab22770125371d2e20454

Observation 0983f300-b362-4d03-a785-f905f04a78b2 · outbound

This paper cites An algo- rithm for intelligibility prediction of time–frequency weighted noisy speech,.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement An algo- rithm for intelligibility prediction of time–frequency weighted noisy speech,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:bc57f889469e8e9c55d5d064114d70d24c75421a04ae5736f4c4f8c80ef34b5c

Pith citing papers

Observation 89dbd48a-c87c-4bca-acaa-4f190bf0b482 · inbound

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement cites this paper.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T22:49:01.387089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:87c4853963e199c81fc96940bcd3e8df0efbcf3c554d1e802a6e98520529b571