Pith. sign in

Paper Citation Record · LEDGER

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation

As of 14 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 1 inbound Pith citation observation for arXiv:2608.09288.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.09288 v1

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T20:11:17.649998Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T20:11:17.344737Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-11T20:11:18.008189Z

Reference resolution

41 of 41 outbound references displayed

  • verified exact0
  • verified fuzzy30
  • unresolved8
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fc445400-b253-469c-9d90-86a84a66ae2c · outbound

This paper cites an unresolved cited work.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-11T20:11:18.851075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:11:17.333888Z digest=sha256:25c74d5526cdee672d9f700aa6c4795a8be6d418b8102f1e31ed01e7000fbc6d

Observation 25cdecb6-faa7-4f86-8dc5-9fec91c89c82 · outbound

This paper cites DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation

Reference 2

Resolution
malformed identifier
local_arxiv, observed 2026-08-11T20:11:18.015625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:11:17.344737Z digest=sha256:ab8b5560da1991d6538946f78c778c48f530a938cf5083193103f5062a48f623

Observation a91461eb-53fb-4c65-a36f-5604f41ea4e0 · outbound

This paper cites As illustrated in Figure 1, the framework comprises three components.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation As illustrated in Figure 1, the framework comprises three components

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:18.829440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:11:17.354145Z digest=sha256:78edda61b2defaa6dae5762ab400c75de59a70e3682f6a7926585395ab4782dc

Observation 66e2904c-72a7-44cf-942e-09160a56e2a8 · outbound

This paper cites Each GPU processes a batch of 2 three-second segments.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation Each GPU processes a batch of 2 three-second segments

Reference 4

Resolution
malformed identifier
raw_fallback, observed 2026-08-11T20:11:17.979671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:11:17.363813Z digest=sha256:b0a4a15bfef5f28fe0e93c3b470c3c37193e9eab91d9e9c52b28ece68426d63e

Observation 75eea7ea-8709-4388-8bb1-1e3ec166f98e · outbound

This paper cites an unresolved cited work.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-11T20:11:18.809126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:11:17.375982Z digest=sha256:f9298040e0e89d76b285ba0c937840c6fbff514062333b04efe4f3658db61d9c

Observation 83323de6-c4ea-426e-952e-d990f5845e29 · outbound

This paper cites Conv-TasNet: Surpassing ideal time- frequency magnitude masking for speech separation,.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation Conv-TasNet: Surpassing ideal time- frequency magnitude masking for speech separation,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:18.787881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:11:17.384365Z digest=sha256:f73373ef314f27e5d781fe59d368b2d584a6a87a39098ccae31a17c4d2a39f6f

Observation 115e74d8-e6ce-45f5-a227-59b676efcbc3 · outbound

This paper cites Permutation invari- ant training of deep models for speaker-independent multi-talker speech separation,.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation Permutation invari- ant training of deep models for speaker-independent multi-talker speech separation,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T20:11:17.392252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:11:17.392252Z digest=sha256:4d61677c3f7df5dc23de8765aaeacacbc10e18ff175283c7a3d1ca596ef8a2d5

Observation e8a1fcb8-2e8a-470c-b1f9-6c4cf83f3fd5 · outbound

This paper cites Deep clus- tering: Discriminative embeddings for segmentation and separa- tion,.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation Deep clus- tering: Discriminative embeddings for segmentation and separa- tion,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T20:11:17.401926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:11:17.401926Z digest=sha256:fd0b5223694be7c09f68b00520b8ea80ece0497f5becf43af12e79b16a35bb35

Observation 57262555-fa9a-4c87-9042-8815d639a69f · outbound

This paper cites Ps4: Proxy-supervised joint training for real target speaker extraction,.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation Ps4: Proxy-supervised joint training for real target speaker extraction,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:18.733292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:11:17.408624Z digest=sha256:1da7c1bb609efe5a437d57e599e04157b20c76a50707e322253a204a8659997f

Observation 6bb8351a-2688-4b3f-a87f-a37375ae6694 · outbound

This paper cites Attention is all you need in speech separation,.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation Attention is all you need in speech separation,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T20:11:17.468283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:11:17.468283Z digest=sha256:098c48ee15562429167bc2ba23a5dbb6488708ea48526eb2aa6b0a5d084348f1

Observation 5e71a295-8a89-476d-bb03-6856ba402afb · outbound

This paper cites Looking to listen at the cock- tail party: A speaker-independent audio-visual model for speech separation,.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation Looking to listen at the cock- tail party: A speaker-independent audio-visual model for speech separation,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:18.712514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:11:17.426200Z digest=sha256:d8a914fa120d0a16486e26db1395bc3984a4eadaffcb27d9361d8abfa9491c97

Observation 96b59652-e80d-4c7c-8076-0f3dc98bea56 · outbound

This paper cites The conversation: Deep audio-visual speech enhancement,.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation The conversation: Deep audio-visual speech enhancement,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:18.693661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:11:17.436060Z digest=sha256:3766ef9de374e6b8097f58168140570e3d32e0d06b3402b9f2c9ea2ab208405f

Observation 992753b2-b15a-4bfb-b1e8-9c9bbfe2381a · outbound

This paper cites The fifth CHiME speech separation and recogni- tion challenge: Dataset, task and baselines,.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation The fifth CHiME speech separation and recogni- tion challenge: Dataset, task and baselines,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:18.673732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:11:17.447103Z digest=sha256:267072d72136c02d56ca3f75b2b9a9a2e20e72895fb82bfe45db40ff1a59ee81

Observation 3f8a515a-c3a2-4c16-84e0-966d748376b7 · outbound

This paper cites Time domain audio visual speech separation,.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation Time domain audio visual speech separation,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:18.653153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:11:17.455231Z digest=sha256:2cc558f4f559e2e85644aceecafe3370dc3b029d3d125c0f572e5926c43bfdd4

Observation d71aea48-a37b-45d7-8ae9-f0a815ffa552 · outbound

This paper cites Dual-path RNN: Efficient long sequence modeling for time-domain single-channel speech sepa- ration,.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation Dual-path RNN: Efficient long sequence modeling for time-domain single-channel speech sepa- ration,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:18.631428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:11:17.461553Z digest=sha256:ed579705ee3c0d7ea49b14b1e181a8886bacb14717caf42a0b68978184003116

Observation 253453a5-10ff-4d60-a783-d704dc62d06e · outbound

This paper cites Tiger: Time-frequency interleaved gain extraction and reconstruction for efficient speech separation,.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation Tiger: Time-frequency interleaved gain extraction and reconstruction for efficient speech separation,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:18.443078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:11:17.522082Z digest=sha256:eba92e1b148419a08a237fe668427ae490254c221e186ed273014e48587dd51e

Observation d2608980-c4e3-496e-b4f6-25dfff961332 · outbound

This paper cites The AliMeeting corpus: A multi-modal meeting corpus,.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation The AliMeeting corpus: A multi-modal meeting corpus,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:18.595219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:11:17.473940Z digest=sha256:a731c929fb39cc873256f9d8b4b6c90a4a5d7d420c6e0f4eee045a6c17358d24

Observation fd809e29-c043-4c39-9c87-f7b69a1723b6 · outbound

This paper cites MISP: A multi-modal interactive speech process- ing system,.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation MISP: A multi-modal interactive speech process- ing system,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:18.572367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:11:17.483059Z digest=sha256:f3e3c68120f02e35aafce42a9c14282024f916a853f953c5759ca60e6d7c0acf

Observation 3550d6bb-89c4-4abc-b4b9-12293e97788f · outbound

This paper cites AISHELL-4: An open source dataset for speech separation, diarization and recognition,.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation AISHELL-4: An open source dataset for speech separation, diarization and recognition,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:18.549663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:11:17.494065Z digest=sha256:d8ef8abfeba003d9b73be1d41c5fe88e613f266c91fee799c78885eda19a142b

Observation 7d840653-aa0e-466d-96f8-e7feaa5005db · outbound

This paper cites Pyroomacoustics: A Python package for audio room simulation and array processing algorithms,.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation Pyroomacoustics: A Python package for audio room simulation and array processing algorithms,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:18.517409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:11:17.501409Z digest=sha256:49fe33d5345e15c23a6984f0230d11cac27477b82ba589f1bb8e80f77453d984

Observation 493c7620-d86e-44ec-b7c9-abc3c55dfb05 · outbound

This paper cites A tutorial on hidden markov models and selected applications in speech recognition,.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation A tutorial on hidden markov models and selected applications in speech recognition,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:18.485875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:11:17.508585Z digest=sha256:88a923593eebbe935ecfa5123884815e993a67e6de88311c5c8d59d6b7359c8f

Observation 17111b39-8046-4a0f-80e2-243022073d20 · outbound

This paper cites WeSpeaker: A research and production oriented speaker embedding learning toolkit,.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation WeSpeaker: A research and production oriented speaker embedding learning toolkit,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:18.324783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:11:17.570686Z digest=sha256:01cdeb9cad8cdcb16828f4309f25e16fddbbe379303189b853580f104ff6dc5c

Observation 0091574d-79e4-4869-bf45-747464705955 · outbound

This paper cites FunASR: A fundamental end-to-end speech recog- nition toolkit,.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation FunASR: A fundamental end-to-end speech recog- nition toolkit,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:18.422808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:11:17.530326Z digest=sha256:f7c9b417fe3fb4d82307e266699caf570173f7c9ffd805a7e97d331cdf99c762

Observation 10891722-f7dd-4c05-974e-b8140a655b55 · outbound

This paper cites Room impulse response generator,.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation Room impulse response generator,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:18.394748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:11:17.537504Z digest=sha256:6cd62da8e3f66bbba963370c5efa4cf423d6c85715e9bef13c5e87c9695dd1f5

Observation 4c61cefb-b562-47f4-adbc-271e34d9e0f3 · outbound

This paper cites MUSAN: A Music, Speech, and Noise Corpus.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation MUSAN: A Music, Speech, and Noise Corpus

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T20:11:17.545140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:11:17.545140Z digest=sha256:f382ac5cbf3c47cee8552210929d9a096d93d75805e056902a568053641d924d

Observation d49fe0c5-3e18-47d9-bfa1-bbf262df550c · outbound

This paper cites Overcoming catastrophic forgetting in neural networks,.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation Overcoming catastrophic forgetting in neural networks,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T20:11:17.553114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:11:17.553114Z digest=sha256:1568f3f4a637f52796b07b04b534c669726917c24db92d03ebfb66c4f5846aa9

Observation cd12754a-dbc8-4030-b37a-449c3f8bbe28 · outbound

This paper cites SDR— half-baked or well done?.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation SDR— half-baked or well done?

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:18.350823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:11:17.561434Z digest=sha256:f3d423f552a3d8e63cf5ce13e4aa2954e50e5e1ad6551acdc9f5c44f2a9ff60d

Observation 80506bb9-9d1d-4262-9005-0f181893d814 · outbound

This paper cites Out of time: automated lip sync in the wild,.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation Out of time: automated lip sync in the wild,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:18.182349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:11:17.607522Z digest=sha256:33f612658ebf5f63633f26daf172ae9986ea505bd3b2b44aa33535d8a03f0935

Observation 4621ba01-de89-496c-b84a-3f695fe93e4c · outbound

This paper cites An al- gorithm for intelligibility prediction of time-frequency weighted noisy speech,.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation An al- gorithm for intelligibility prediction of time-frequency weighted noisy speech,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:18.292973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:11:17.576768Z digest=sha256:6a6130bffc07a8110ccce35ac893ff1f9d41496342a81232f072752705c9ca51

Observation 37610fde-13b3-4199-b35a-cf321ca2121e · outbound

This paper cites Per- ceptual evaluation of speech quality (PESQ)—a new method for speech quality assessment of telephone networks and codecs,.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation Per- ceptual evaluation of speech quality (PESQ)—a new method for speech quality assessment of telephone networks and codecs,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:18.259065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:11:17.583386Z digest=sha256:6b32f656f0f5472f982c38989909a6425af122e5a2ec3de65e958986ec0a7563

Observation c499a89d-fa14-4693-a02d-f02d1cfabeca · outbound

This paper cites UTMOS: UTokyo-SaruLab system for V oiceMOS challenge 2022,.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation UTMOS: UTokyo-SaruLab system for V oiceMOS challenge 2022,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:18.233862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:11:17.589283Z digest=sha256:426734acde89577d76aece0e784f4c48307153089d3a5a43bc0862dcb8b12176

Observation 38e85c96-5376-4b67-a3d7-376e9a4f4a18 · outbound

This paper cites CN-Celeb: A challenging Chinese speaker recognition dataset,.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation CN-Celeb: A challenging Chinese speaker recognition dataset,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:18.207925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:11:17.594834Z digest=sha256:f8b2aa1616805bd6bf9dc26a4393b5acc91738d84bf9d544edfb23a52fb3903c

Observation bdaae359-ff6e-41a6-a41f-d9f99bc321c9 · outbound

This paper cites LatentSync: Taming Audio-Conditioned Latent Diffusion Models for Lip Sync with SyncNet Supervision.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation LatentSync: Taming Audio-Conditioned Latent Diffusion Models for Lip Sync with SyncNet Supervision

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T20:11:17.601296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:11:17.601296Z digest=sha256:d84116796176095a9240815dc82f838a78ab9a1e0e8247ba7f3d48c9cd1d6cf4

Observation 98347166-25ef-4caa-be48-4d0cea703896 · outbound

This paper cites DNSMOS: A non-intrusive perceptual objective speech quality metric to evaluate noise sup- pressors,.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation DNSMOS: A non-intrusive perceptual objective speech quality metric to evaluate noise sup- pressors,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:18.056216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:11:17.645482Z digest=sha256:b096d2ddd269465b852c36dfce2f8beb437083a960146bd9d1faacdbbf514a5d

Observation bf281f55-d30e-403c-91c5-025d59aa647a · outbound

This paper cites Perfect match: Improved cross-modal embeddings for audio-visual synchronisa- tion,.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation Perfect match: Improved cross-modal embeddings for audio-visual synchronisa- tion,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:18.161583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:11:17.613825Z digest=sha256:9d6053710a0fee1750ca6f99cefbfcb91308983fd13b9281f10a76eee17e0955

Observation 09a28137-d12b-41e7-b2de-3002d75bf121 · outbound

This paper cites How far are we from solving the 2D & 3D face alignment problem? (and a dataset of 230,000 3D facial landmarks),.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation How far are we from solving the 2D & 3D face alignment problem? (and a dataset of 230,000 3D facial landmarks),

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:18.141478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:11:17.622982Z digest=sha256:dae1bd2e21923557743cd44e5d5dd2651e05027071623236a46c4a3ff0a3be56

Observation 756fc988-95c4-4891-bd19-25e58cd4a8ae · outbound

This paper cites SEGAN: Speech en- hancement generative adversarial network,.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation SEGAN: Speech en- hancement generative adversarial network,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:18.121792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:11:17.629274Z digest=sha256:c7a541e386ffd39cf1b302e5cc20e0de83d920e383fbae024224dfa4316b5fa9

Observation 2e55f7fb-f54a-4281-959d-9e2b2fe8e520 · outbound

This paper cites MossFormer2: Combining transformer and RNN-free recurrent network for enhanced time-domain monaural speech separation,.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation MossFormer2: Combining transformer and RNN-free recurrent network for enhanced time-domain monaural speech separation,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:18.101171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:11:17.636161Z digest=sha256:226063fb5966a7085a5a6a497590882f6d0b22ac48b66d0917c1c082a8e7f531

Observation 37ccb782-15f1-44b6-aa2c-b3cf76884711 · outbound

This paper cites Recommendation ITU- R BS.1770-4: Algorithms to measure audio programme loudness and true-peak audio level,.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation Recommendation ITU- R BS.1770-4: Algorithms to measure audio programme loudness and true-peak audio level,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:18.078598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:11:17.641112Z digest=sha256:f368dd562fe80e77338567450b08d28862e2d37cf4075d41b04bc99d65fc3e61

Observation 58be2944-db1e-452e-a1ce-174982bea642 · outbound

This paper cites Adam: A method for stochastic opti- mization,.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation Adam: A method for stochastic opti- mization,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:18.037042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:11:17.649998Z digest=sha256:a47720af618a147e4f4631226b3a5ced9c97dcc10e81d7f5b1986a9d9fdf0c79

Observation 3a2c279c-6dfb-463f-8a7b-c2b2b003c788 · outbound

This paper cites PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction

Reference 2026

Resolution
metadata mismatch
local_arxiv, observed 2026-08-11T20:11:17.777692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:11:17.417789Z digest=sha256:e7ecd1bc11adae997e1c99ec9262156377ebf0b42d4882eff7deb6f73781b43f

Pith citing papers

Observation 25cdecb6-faa7-4f86-8dc5-9fec91c89c82 · inbound

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation cites this paper.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation

Reference 2

Resolution
malformed identifier
local_arxiv, observed 2026-08-11T20:11:18.015625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:11:17.344737Z digest=sha256:ab8b5560da1991d6538946f78c778c48f530a938cf5083193103f5062a48f623