Pith. sign in

Paper Citation Record · LEDGER

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement

As of 7 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 1 inbound Pith citation observation for arXiv:2606.23712.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.23712 v1

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-26T23:17:45.299833Z

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-26T23:17:45.299833Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-03T22:49:01.384665Z

Reference resolution

35 of 35 outbound references displayed

  • verified exact4
  • verified fuzzy0
  • unresolved29
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5deb219c-beb9-49e1-b4f4-5d425e4dabe3 · outbound

This paper cites Recent deep learn- ing approaches have significantly improved performance [1–3], with generative modeling frameworks emerging as a powerful direction.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement Recent deep learn- ing approaches have significantly improved performance [1–3], with generative modeling frameworks emerging as a powerful direction

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:604fd136d785598cfd82a62a58d525a407e6c083d9983a4200657930b15f3ac5

Observation 89dbd48a-c87c-4bca-acaa-4f190bf0b482 · outbound

This paper cites Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T22:49:01.387089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:b35099bf49056bd0a25cf5df64092d955343a892513ec4190c7dc5501417c5cf

Observation a9516dbd-c7dc-4cc8-9b84-34a969b732b5 · outbound

This paper cites 1, describing how the audio-visual con- trastive loss is computed and the motivation behind it.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement 1, describing how the audio-visual con- trastive loss is computed and the motivation behind it

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:70330015ec3b4d8991c547d6a5e447ac862d84f3c403cf99f4a8b3f000f4e5b9

Observation 29ed98d9-149b-47af-ad22-c702fad92144 · outbound

This paper cites As baselines, we con- sider A V-DiffUSEEN, with cross-attention fusion [8, 13], the audio-only version, AO-DiffUSEEN [13], and the supervised- generative FlowA VSE model [7].

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement As baselines, we con- sider A V-DiffUSEEN, with cross-attention fusion [8, 13], the audio-only version, AO-DiffUSEEN [13], and the supervised- generative FlowA VSE model [7]

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:374af770c15f8b5dcd44e8adc576fdee8eb848fc51b2b7bcdac1b11dca78c79a

Observation 78d6f946-dbcc-498e-a989-35221d479999 · outbound

This paper cites Matched condition: TCD-DEMAND Under matched conditions (Table 1), the proposed model im- proves all metrics compared to A V-DiffUSEEN.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement Matched condition: TCD-DEMAND Under matched conditions (Table 1), the proposed model im- proves all metrics compared to A V-DiffUSEEN

Reference 5

Resolution
malformed identifier
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:4ed56e53e1db267de9e69b33bdc4feef09effcd2fb5c6b969a84ec4cf603eeda

Observation 524e131a-c194-485f-a7c3-9ed5ce3046b2 · outbound

This paper cites During the pretraining of the visual-conditioned speech diffusion model, we augment the denoising score matching ob- jective with a contrastive audio-visual alignment loss.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement During the pretraining of the visual-conditioned speech diffusion model, we augment the denoising score matching ob- jective with a contrastive audio-visual alignment loss

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:d2e4b84a15f8dd0b451f40b88a431111de1171fe672a6692d2beb067145e6594

Observation a8dccfb8-f91f-46f7-82d7-19fd41abd647 · outbound

This paper cites an unresolved cited work.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:3c2815b0f9e345609fabf13646d69163a51ed7cc8bf62e3d9a95b871a2d994e0

Observation abe0aa72-a4df-4085-a5ff-2157cc113e88 · outbound

This paper cites Generative AI tools were only used to edit and polish some portions of the manuscript.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement Generative AI tools were only used to edit and polish some portions of the manuscript

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:f582be93aa4ba9456edde9ad5f61089bba87071ca01d65a922e360353e31b4c7

Observation bc98de02-c35e-4519-9557-50233c8b473a · outbound

This paper cites SEGAN: Speech enhancement generative adversarial network,.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement SEGAN: Speech enhancement generative adversarial network,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:448bc96014980eae8b1cea3dc8fae784943ba79ecb3acfbbec6cdcac3c93d2e6

Observation 42c0a82b-ce2c-4996-9b44-0c0f8463d353 · outbound

This paper cites Conv-TasNet: Surpassing ideal time– frequency magnitude masking for speech separation,.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement Conv-TasNet: Surpassing ideal time– frequency magnitude masking for speech separation,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:df947689728a7c2bdded0ace5a15058796354a049804e4209142716e42faebc9

Observation 7836dbb8-7d74-40fb-ae77-af0bb744b29d · outbound

This paper cites DCCRN: Deep complex convolution recurrent network for phase-aware speech enhancement,.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement DCCRN: Deep complex convolution recurrent network for phase-aware speech enhancement,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:b3ad7cae4d1a986824bdbbc3e577f13b14afaf99ff4c4f7fd3651824f193d657

Observation 7e54e6c0-1e8e-4fe8-9274-037b1169f730 · outbound

This paper cites Speech enhancement and dereverberation with diffusion-based gen- erative models,.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement Speech enhancement and dereverberation with diffusion-based gen- erative models,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:63a843e86d57cff84b50141b8777828b6ddcdab90c43b67f373d1387f56a408a

Observation ccded98c-1c52-4c7c-a183-7c800ab26646 · outbound

This paper cites The conversation: Deep audio-visual speech enhancement,.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement The conversation: Deep audio-visual speech enhancement,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:0f87d8eaead9300ff7589e006ac9dc165c9971419997e1be014efa36f9ef628e

Observation 6e168e88-24da-43a7-9283-ded77279913f · outbound

This paper cites Audio-visual speech enhancement using multimodal deep con- volutional neural networks,.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement Audio-visual speech enhancement using multimodal deep con- volutional neural networks,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:90adfb36af9e734d8800408c43422a633b0fd06136bac5a465de898caec9a9ee

Observation 2fa0d510-f391-4a3f-97a8-4915a7be2514 · outbound

This paper cites FlowA VSE: Efficient audio-visual speech enhancement with conditional flow matching,.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement FlowA VSE: Efficient audio-visual speech enhancement with conditional flow matching,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:e86710dedafa4a85e72820c0993f3a71c624eeac5f331384b51b2aac4f5ffd0c

Observation 0c7e79e7-2730-4bf5-a57d-4f0ba8828a25 · outbound

This paper cites Diffusion-based unsupervised audio-visual speech enhancement,.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement Diffusion-based unsupervised audio-visual speech enhancement,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:b13045319bd583a66080e4f613a8dd8a1f67ba260ff616ee799965bc40bdd50d

Observation 00df5926-de7b-48df-95f8-0b002ed8d94a · outbound

This paper cites An overview of deep-learning-based audio-visual speech enhancement and separation,.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement An overview of deep-learning-based audio-visual speech enhancement and separation,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:f440d87cbce4b003cff8328edc8ee32e277bd7a9dbf61b03781d797f9a0543d5

Observation 3b129cf7-9d46-4a10-83dc-ebecdac0bbec · outbound

This paper cites Learning transfer- able visual models from natural language supervision,.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement Learning transfer- able visual models from natural language supervision,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:142bc0b6768af0ab2d02555b0ede7fcf1aa2ed0f97854462c9f53c2fc3ce13c6

Observation 3ed6430b-a876-42ef-a7b3-fc3397cc4d8c · outbound

This paper cites CLAP: Learn- ing audio concepts from natural language supervision,.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement CLAP: Learn- ing audio concepts from natural language supervision,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:28c16e284ad3476bccf63dda1909ae68d3dc1ba3520ed87f64e4eb0135e0b0cf

Observation 032e5418-8cae-463e-87a7-1922dc297312 · outbound

This paper cites SA V-SE: Scene-aware audio-visual speech enhancement with selective state space model,.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement SA V-SE: Scene-aware audio-visual speech enhancement with selective state space model,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:de7d4c1d3303a5f9a37619a55dd31fa95f40425612e860e99b2dcf9c3da83889

Observation f9107a07-13aa-4db5-aaa1-dd0767c09fdb · outbound

This paper cites Diffusion-based Frameworks for Unsupervised Speech Enhancement.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement Diffusion-based Frameworks for Unsupervised Speech Enhancement

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-07-03T22:49:01.372695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:fc538fc91fbad14757bdc496444d20009c69b4657c3aa5c2d1d50b286eeab333

Observation 9ec4f8fc-b3ab-4292-9e12-fb2a8417dbf2 · outbound

This paper cites Score-based generative modeling through stochastic differ- ential equations,.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement Score-based generative modeling through stochastic differ- ential equations,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:6b1e2d2b7164a23da81f7c913b9ae656a3f839f32653d17c08fb623ea0d957d9

Observation 4ad42cd0-8bd9-49ef-ac26-3299a6263287 · outbound

This paper cites Nonnegative matrix factor- ization with the itakura-saito divergence: With application to music analysis,.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement Nonnegative matrix factor- ization with the itakura-saito divergence: With application to music analysis,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:14c5988783f24f4b00c6693acff53d8335a27c3c236e9d38c9d7481a1531b30b

Observation 4637457c-8421-4014-a872-d69d9cc1cda5 · outbound

This paper cites Tweedie’s formula and selection bias,.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement Tweedie’s formula and selection bias,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:4ea07e07402e68b0cc52e81766a568bc6624070685f6b3e85f56c393ab29bd27

Observation eea11571-501f-4479-acdc-02b3503f7786 · outbound

This paper cites Deep residual learning for im- age recognition,.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement Deep residual learning for im- age recognition,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:7b88adddf07b4c474c6b3141791e8b0778d340deb92038ed6b746d8cefcbb4e0

Observation 524a6f4f-0fc6-4ba4-a4ca-47e4a843d284 · outbound

This paper cites Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-03T22:49:01.368464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:b66e6cf70b935b408cd2617bb3c6979f59d347babe175fc6ff87a86528b4e533

Observation 7d69518e-b9cd-4010-9d25-aa3d18057d95 · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement Representation Learning with Contrastive Predictive Coding

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-07-03T22:49:01.377870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:432420804b4871dd76b8b1da90f16bdfeed0e019ccb923899fc3c430a0a4ad0e

Observation fcbf7481-a563-4ee6-a647-5048a5bc8b76 · outbound

This paper cites TCD-TIMIT: An audio-visual corpus of con- tinuous speech,.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement TCD-TIMIT: An audio-visual corpus of con- tinuous speech,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:37dfa2aa984a0d4fb03589d9ff66fc3f11b8d44e8ae74068ca1f66bf1d4279e9

Observation 320a8621-46b7-4b0d-8d90-173f4644e06e · outbound

This paper cites The diverse environments multi- channel acoustic noise database (DEMAND): A database of multi- channel environmental noise recordings,.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement The diverse environments multi- channel acoustic noise database (DEMAND): A database of multi- channel environmental noise recordings,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:7c62d8b8d2768fafe733af7ed8784b5282ec5db1bdfe7d179210a05bcd5bc3c3

Observation 63597d9c-061d-4644-930c-d95d8a30ea95 · outbound

This paper cites LRS3-TED: a large-scale dataset for visual speech recognition.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement LRS3-TED: a large-scale dataset for visual speech recognition

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-07-03T22:49:01.382024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:4c136698c724cef14392266df357c7cc56703e3e341952765af6b2771c5f9057

Observation 7016b475-36d9-4462-bdd4-24d74b2ec68b · outbound

This paper cites NTCD-TIMIT: A new database and base- line for noise-robust audio-visual speech recognition.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement NTCD-TIMIT: A new database and base- line for noise-robust audio-visual speech recognition

Reference 31

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:8f1a2897a14c5409526b75ebb38615f6f7f3d1ee2a383cc4043e4e98eb2cb981

Observation 709ebb50-95b1-4e58-b620-921464db7ada · outbound

This paper cites SDR–half- baked or well done?.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement SDR–half- baked or well done?

Reference 32

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:a76e87559c6e3d17db341c3844d4ce7641f178231b35a6f08e3b599d710b4f60

Observation bcd495df-920c-485f-85e3-ddb03a86f8fc · outbound

This paper cites Performance measure- ment in blind audio source separation,.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement Performance measure- ment in blind audio source separation,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:bb5cc16619f651feac47e0a365436a8cd5303b2afc28d913b704eee4165e15f5

Observation 272ef1c0-95a9-4e6f-a1f0-8d00467d0975 · outbound

This paper cites Perceptual evaluation of speech quality (PESQ)-a new method for speech quality assessment of telephone networks and codecs,.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement Perceptual evaluation of speech quality (PESQ)-a new method for speech quality assessment of telephone networks and codecs,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:41674ad959c5b95133aaf21100ed2497b2a1877960dbb73e73d05cf194388b54

Observation 0983f300-b362-4d03-a785-f905f04a78b2 · outbound

This paper cites An algo- rithm for intelligibility prediction of time–frequency weighted noisy speech,.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement An algo- rithm for intelligibility prediction of time–frequency weighted noisy speech,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:b3042246c2ab28953c9d681f28935a4769e75e11b13115d6d8e8a63598265501

Pith citing papers

Observation 89dbd48a-c87c-4bca-acaa-4f190bf0b482 · inbound

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement cites this paper.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T22:49:01.387089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:b35099bf49056bd0a25cf5df64092d955343a892513ec4190c7dc5501417c5cf