Pith. sign in

Paper Citation Record · LEDGER

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment

As of 8 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 7 inbound Pith citation observations for arXiv:2506.19398.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.19398 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:11:24.534728Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:11:24.362189Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T13:08:08.743449Z

Reference resolution

45 of 45 outbound references displayed

  • verified exact1
  • verified fuzzy38
  • unresolved5
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 69262a98-01cc-4375-8be0-ce2617879210 · outbound

This paper cites While crucial for these applications, ac- curately processing speech is challenged by the often degraded quality of real-world audio.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment While crucial for these applications, ac- curately processing speech is challenged by the often degraded quality of real-world audio

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:25.131306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:11:24.356498Z digest=sha256:df4a6d45420b180064a952789129688f9eef3bfa7f4a186851a770e7269f470e

Observation 714fd3b6-e2d1-45c8-b151-2d020951eb2c · outbound

This paper cites ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T23:11:24.362189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:11:24.362189Z digest=sha256:bb007aa4970f73650a2ca284fa3f38f2a7d0aab6da2a225c254162ff170204d0

Observation ffbe1fee-e561-42d0-b535-2900b9f5f465 · outbound

This paper cites Training strategies 3.1.1.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment Training strategies 3.1.1

Reference 3

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T23:11:25.112661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:11:24.366582Z digest=sha256:b0e53f3a240c1960cec629e1749996834dc408a2c795598c288b5d11302051ba

Observation 72b8ed4b-36ef-4577-8a86-85b840ff6436 · outbound

This paper cites Beyond the presented evaluations, ClearerV oice- Studio is available for live demos on HuggingFace and Mod- elScope, enabling users to experiment with real-world record- ings.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment Beyond the presented evaluations, ClearerV oice- Studio is available for live demos on HuggingFace and Mod- elScope, enabling users to experiment with real-world record- ings

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:25.095263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:11:24.371021Z digest=sha256:f981bd51d3c32de04e56e6a6e531b6c32e59381a0b8cd12b957c14cf10bde74c

Observation 641f26e5-0e0d-4eb6-aae4-80b3166b30ec · outbound

This paper cites Deep learning for audio signal processing,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment Deep learning for audio signal processing,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:25.078020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:11:24.376338Z digest=sha256:d70ea76c9c20535b229a425d34b89ffa9ad68dba4076185189f9bb45f79b9b03

Observation 84304d9c-cd5d-423d-9c93-b1b942667db1 · outbound

This paper cites Mamba in Speech: Towards an Alternative to Self-Attention.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment Mamba in Speech: Towards an Alternative to Self-Attention

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T23:11:24.380674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:11:24.380674Z digest=sha256:2f0e978d6705cf8a2d099a04d49ff88b802a4b18cd5c791b5205ef09dbbb7e54

Observation 6082d4d3-b631-4c69-ad1b-0d83a76af01a · outbound

This paper cites DeepMMSE: A deep learning approach to mmse-based noise power spectral density estimation,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment DeepMMSE: A deep learning approach to mmse-based noise power spectral density estimation,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:25.063117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:11:24.385499Z digest=sha256:3f160332d47a1436f3a429290b958cbb51ef91f7dea642f3e35a9822e0be1928

Observation 2a35cea7-d23a-4042-93f2-bdc8a40e17dd · outbound

This paper cites SpeechBrain: A general-purpose speech toolkit,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment SpeechBrain: A general-purpose speech toolkit,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:25.047106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:11:24.389448Z digest=sha256:c510a6a0b02bc5e981ffffe26c173b650209ecf10f5e4c429cf908566e2a61d8

Observation 651b4301-b4e2-413f-8fbd-95a5ae5587c8 · outbound

This paper cites AudioSR: Versatile audio super-resolution at scale,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment AudioSR: Versatile audio super-resolution at scale,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.978245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:11:24.416110Z digest=sha256:20c7e558dd9b123ea25bb6b40c0910c85f1440d9a778dd0c35ffe28ba5d45fa7

Observation 0a8da161-6948-4863-9ea9-090d76f982a9 · outbound

This paper cites ESPnet: End-to-end speech processing toolkit,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment ESPnet: End-to-end speech processing toolkit,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:25.032797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:11:24.399073Z digest=sha256:678e12c25846b723e9cc056e5d1a6dff9df3e8923fc6465f6b248e44ecb94ac3

Observation d4baf6c7-b9da-4399-9bab-d6526be4a8df · outbound

This paper cites Summary on the multimodal information-based speech processing 2023 challenge,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment Summary on the multimodal information-based speech processing 2023 challenge,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:25.020116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:11:24.403507Z digest=sha256:95decba0b9b53161cf80999e64feaae103ebd85d46889ede333537e5e527b065

Observation 51bb2732-9407-4c47-a76c-6aebd1a7ea79 · outbound

This paper cites As- teroid: the PyTorch-based audio source separation toolkit for re- searchers,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment As- teroid: the PyTorch-based audio source separation toolkit for re- searchers,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:25.007293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:11:24.407313Z digest=sha256:87aa5b14a6e0cac9f91b5b71401b08f5d6e92736eca8e5603d918bf4d295c938

Observation 95a7cd92-8c58-4079-bed2-62ae14bd37ca · outbound

This paper cites DeepFilterNet: Perceptually motivated real-time speech en- hancement,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment DeepFilterNet: Perceptually motivated real-time speech en- hancement,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.993239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:11:24.411596Z digest=sha256:719b0799766f809fdf0b016bf735fb104b88667c4fdd2958c6141bd911ab851d

Observation cc184939-26b3-4313-8d7f-57dec752ac13 · outbound

This paper cites Hifi-SR: A unified generative transformer-convolutional adversarial net- work for high-fidelity speech super-resolution,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment Hifi-SR: A unified generative transformer-convolutional adversarial net- work for high-fidelity speech super-resolution,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.915062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:11:24.435224Z digest=sha256:2f5be844e71454d72770d227a53c1bb029b23317bec1ee45d8a943fad7421186

Observation b444fbaa-fe93-4bfe-8501-0343628453ef · outbound

This paper cites FlowA VSE: Ef- ficient audio-visual speech enhancement with conditional flow matching,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment FlowA VSE: Ef- ficient audio-visual speech enhancement with conditional flow matching,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.964648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:11:24.420128Z digest=sha256:488e27d186aebb1237dc22c116fe59680506b254a47d936cdab0bb463cbea92e

Observation 60e16eda-03ee-4f93-af38-77890af77535 · outbound

This paper cites FRCRN: Boosting feature representation using frequency recurrence for monaural speech enhancement,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment FRCRN: Boosting feature representation using frequency recurrence for monaural speech enhancement,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.950815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:11:24.424200Z digest=sha256:38afaa2179963d2963a84392f8c496134a31052db41d0e7274763ee228781f63

Observation f9a92ae9-a8d1-46e3-9c7c-c7226a044498 · outbound

This paper cites MossFormer2: Combin- ing transformer and rnn-free recurrent network for enhanced time- domain monaural speech separation,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment MossFormer2: Combin- ing transformer and rnn-free recurrent network for enhanced time- domain monaural speech separation,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.937368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:11:24.427496Z digest=sha256:4be4ace0ab0a520f09e41796ca98adf5d4345ab5ad4c2b1745281109d8ec2258

Observation 3f76367b-dfd3-42af-8818-5c8d59f4bbf4 · outbound

This paper cites MossFormer: Pushing the performance limit of monaural speech separation using gated single-head trans- former with convolution-augmented joint self-attentions,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment MossFormer: Pushing the performance limit of monaural speech separation using gated single-head trans- former with convolution-augmented joint self-attentions,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.926695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:11:24.431574Z digest=sha256:f08e311f9d3fc318e0ff9b62cd684bafe12a5ee9d196d89a38363b55a3b2b98d

Observation 5658879f-5b6d-49b8-a6d9-67f991f513c3 · outbound

This paper cites NeuroHeed: Neuro-Steered Speaker Extraction using EEG Signals.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment NeuroHeed: Neuro-Steered Speaker Extraction using EEG Signals

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-06T23:11:24.586002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:11:24.452262Z digest=sha256:7a5aea9576cbc07899e246b183214440d7b734d69aa2c73e69344f67f2dc6a9e

Observation de92d52c-4bf6-4860-8822-ba96854e3fe5 · outbound

This paper cites HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.902975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:11:24.438524Z digest=sha256:f129081a4f3cfbe6825d08cc4560865358821a39c81dafa25ae9cb15d83c278d

Observation aabc8ee1-2c97-4e64-aa86-a8da807dc826 · outbound

This paper cites Scenario-aware audio-visual TF- Gridnet for target speech extraction,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment Scenario-aware audio-visual TF- Gridnet for target speech extraction,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.891068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:11:24.441735Z digest=sha256:0dd555ea6afe41da83d4870ca625a2cda76a4e62b2e1657cb978372085a25ede

Observation d8c239ae-1c8d-4c32-859d-5764628568ce · outbound

This paper cites Speaker extraction with co-speech gestures cue,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment Speaker extraction with co-speech gestures cue,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.878761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:11:24.445326Z digest=sha256:d58c77971c9c23d210552ab6d502744ae6e001005aa51dff958d96bc08a9bb32

Observation bac71848-182d-4296-86e5-a0f824999631 · outbound

This paper cites SpEx+: A complete time domain speaker extraction network,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment SpEx+: A complete time domain speaker extraction network,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.866309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:11:24.448614Z digest=sha256:9e9f2981d28817b78cd5500772366f2ad71c5fbf61525fbf71de2499dc852258

Observation ec631b81-503b-41f1-8a50-b5ab29dfa231 · outbound

This paper cites DCCRN+: Channel-wise subband dccrn with snr estimation for speech enhancement,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment DCCRN+: Channel-wise subband dccrn with snr estimation for speech enhancement,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.800145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:11:24.472350Z digest=sha256:93683a7903c45010ba7841982d8cc427c8087dc75d36420ebaef323e742bcb27

Observation 56a28c20-a421-4a64-9fef-8351e699a4d6 · outbound

This paper cites The interspeech 2020 deep noise suppression challenge: Datasets, subjective testing framework, and challenge results,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment The interspeech 2020 deep noise suppression challenge: Datasets, subjective testing framework, and challenge results,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.852653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:11:24.456214Z digest=sha256:0ed17e57b711eac72e60981be44f2354bd2264a970d5224bf79c22e9fc3c5703

Observation 3f12de00-69f9-4ec6-9155-361df8278134 · outbound

This paper cites CSTR VCTK Corpus: English multi-speaker corpus for cstr voice cloning toolkit (version 0.92),.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment CSTR VCTK Corpus: English multi-speaker corpus for cstr voice cloning toolkit (version 0.92),

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.840151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:11:24.459983Z digest=sha256:93a7dea19c9142805623c3a657923a1c36465a4b5370a041bcbc4c93e0e46bd2

Observation 659d3a06-5962-4ebf-b2e5-03c71bfa0038 · outbound

This paper cites Audio set: An ontology and human-labeled dataset for audio events,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment Audio set: An ontology and human-labeled dataset for audio events,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.825770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:11:24.463692Z digest=sha256:593f0c9462629116a9bdd071e803f5047fe3df8091475cc7e91f659f3643d17a

Observation 95fdaf25-0bc1-40ec-a534-9a7e5ab7d920 · outbound

This paper cites DEMAND: a collection of multi-channel recordings of acoustic noise in diverse environments,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment DEMAND: a collection of multi-channel recordings of acoustic noise in diverse environments,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.812415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:11:24.468069Z digest=sha256:b1bb5fd898040e1e282d5746ce1c16da5c259df729da81079063e4db0de14b94

Observation 84ac29bb-be00-47d6-9960-e826474b4b6e · outbound

This paper cites Phase- sensitive and recognition-boosted speech separation using deep recurrent neural networks,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment Phase- sensitive and recognition-boosted speech separation using deep recurrent neural networks,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.743338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:11:24.492570Z digest=sha256:1cfc9884e7b0ed201b0c8374966b932df0ce4f5849315653e5c1ccca4645c192

Observation ae34a824-1169-40b3-8f35-5e3e3c074c7b · outbound

This paper cites A mask free neural network for monaural speech enhancement,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment A mask free neural network for monaural speech enhancement,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.788296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:11:24.476740Z digest=sha256:faea35db3dd16133935cd7cac0ddf856c8089b1739006521fc06570751411503

Observation 33a9f56d-411a-468a-9246-0760cb623b87 · outbound

This paper cites TridentSE: Guiding speech enhancement with 32 global tokens,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment TridentSE: Guiding speech enhancement with 32 global tokens,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.777289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:11:24.480847Z digest=sha256:7939f8925c0c6e06283d3debe02b042f39bbb31341a66277a38cb006953040d7

Observation 5ab46bd6-5543-4693-94c5-762f92a30bea · outbound

This paper cites LibriTTS: A corpus derived from librispeech for text- to-speech,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment LibriTTS: A corpus derived from librispeech for text- to-speech,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.766236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:11:24.484737Z digest=sha256:c09e72eccb46bcb246ee3cea0571a993424dbc99428a02c898593bad26481170

Observation 37ab05c7-7c56-42e3-8b87-94489e6e4ff0 · outbound

This paper cites Explor- ing strategies for training deep neural networks,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment Explor- ing strategies for training deep neural networks,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.754403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:11:24.488634Z digest=sha256:6b511b003dd79f7264433d6be6c619f09acf8c90c61cd45656ded11208c02d1b

Observation 1a55352d-409e-441e-8e99-b07c8b85bb82 · outbound

This paper cites An efficient encoder-decoder archi- tecture with top-down attention for speech separation,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment An efficient encoder-decoder archi- tecture with top-down attention for speech separation,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.683161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:11:24.511237Z digest=sha256:961316c6f14b15a4dd8a9d0accafdc990a1b39c2b73eea6379d567aecffd7071

Observation a3042b7b-1c2e-4e49-9fcb-38ad66742e36 · outbound

This paper cites CMGAN: Conformer-based metric gan for speech enhancement,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment CMGAN: Conformer-based metric gan for speech enhancement,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.730747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:11:24.496146Z digest=sha256:ba0ecfdee8d99361987610b6438e4e4695afe056e0a1b1d6ec2997207c8f7f92

Observation d3c53e10-c4f2-4b40-953d-c1a1122527eb · outbound

This paper cites Permutation invari- ant training of deep models for speaker-independent multi-talker speech separation,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment Permutation invari- ant training of deep models for speaker-independent multi-talker speech separation,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.719254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:11:24.499581Z digest=sha256:4da09c0d714532ccf4d62fd3e13ab568f632ac6c8df213d3106c3be389653ea9

Observation 9f24950c-a1cd-4283-9cd1-8ab001d8e2c4 · outbound

This paper cites Dual-Path RNN: Efficient long sequence modeling for time-domain single-channel speech sepa- ration,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment Dual-Path RNN: Efficient long sequence modeling for time-domain single-channel speech sepa- ration,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.706871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:11:24.503833Z digest=sha256:f9d473b9b7a579689bba83ca606657ad5bb5ab35f1eb1612f975bc7c7049dabb

Observation d376de3c-6851-490f-a8d7-be8a5d857958 · outbound

This paper cites Attention is all you need in speech separation,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment Attention is all you need in speech separation,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.695833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:11:24.507563Z digest=sha256:f6fef7b0cfaacd021a22d1fd5398b5cbd5558fe5bb7ad5aa6ae68013af96fcd3

Observation 22a3c68e-0b4a-41ee-a4f2-14383c8ead30 · outbound

This paper cites Selective listening by synchronizing speech with lips,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment Selective listening by synchronizing speech with lips,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.640308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:11:24.530987Z digest=sha256:6b0ff6c8d14ac9f771aa378dd463fbe5b193b49e9cc00547765ec8809030ad46

Observation 07b9b447-c93c-4a44-9e37-c8d9f9885845 · outbound

This paper cites TF-GridNet: Making time-frequency domain models great again for monaural speaker separation,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment TF-GridNet: Making time-frequency domain models great again for monaural speaker separation,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T23:11:24.515143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:11:24.515143Z digest=sha256:a84f6115ee923582fa5abae03d315f16050e5efe2ae7128509beb97432ce734f

Observation 100526fc-2a2d-40f2-a4e1-ed08c42b5971 · outbound

This paper cites SPMamba: State-space model is all you need in speech separation.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment SPMamba: State-space model is all you need in speech separation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T23:11:24.519069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:11:24.519069Z digest=sha256:9214d07b1cd86ecf48366595dc93b22b5dfdaf43d2882a500fa22919b269271d

Observation a9606353-6c7f-47a4-a8e3-39ca7e2241b3 · outbound

This paper cites Time domain audio visual speech separation,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment Time domain audio visual speech separation,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.663849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:11:24.523257Z digest=sha256:a020071cb3672c3e004bc748647bec1042b96e9fa3edcb4a6d86d213edf39e2b

Observation f5854460-642f-4afb-9029-0d21a5d977fe · outbound

This paper cites MuSE: Multi-modal target speaker extraction with visual cues,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment MuSE: Multi-modal target speaker extraction with visual cues,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.651853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:11:24.526741Z digest=sha256:d90d0109a9c4a7f29dd9ada27529bbe26f6627ba651af252e5124fd91945f1ad

Observation ea871c47-8135-4ac8-a3a6-cdb06e9d4e2a · outbound

This paper cites USEV: Universal speaker extraction with visual cue,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment USEV: Universal speaker extraction with visual cue,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.629442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:11:24.534728Z digest=sha256:3d86266ff325e9606c23fdfe2cf4acc01809f7739f06c93de5a3d6c82ac80f5f

Observation 3ae2ef09-51e3-4364-9d36-43eabf7fb5d6 · outbound

This paper cites SpeechBrain: A General-Purpose Speech Toolkit.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment SpeechBrain: A General-Purpose Speech Toolkit

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T23:11:24.394051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:11:24.394051Z digest=sha256:3422e87a55f8f81c0bffa008ff7b13496860d9591e9d54248735fc6144556484

Pith citing papers

Observation 714fd3b6-e2d1-45c8-b151-2d020951eb2c · inbound

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment cites this paper.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T23:11:24.362189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:11:24.362189Z digest=sha256:bb007aa4970f73650a2ca284fa3f38f2a7d0aab6da2a225c254162ff170204d0

Observation 02546d07-92d4-4255-a078-45e13b8b6517 · inbound

CueNet: Robust Audio-Visual Speaker Extraction through Cross-Modal Cue Mining and Interaction cites this paper.

CueNet: Robust Audio-Visual Speaker Extraction through Cross-Modal Cue Mining and Interaction ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T19:39:50.604200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:39:50.604200Z digest=sha256:e1fc1f347d108aa1d8cee599dd7658c102797efaf1f3415bdd40f53b7becfba8

Observation 54563704-4c79-430a-bd33-33beede6abb6 · inbound

Beyond Monologue: Interactive Talking-Listening Avatar Generation with Conversational Audio Context-Aware Kernels cites this paper.

Beyond Monologue: Interactive Talking-Listening Avatar Generation with Conversational Audio Context-Aware Kernels ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:56:01.909103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:16:53.603971Z digest=sha256:f662a7bcd2fa2cb0832339c5154cc5ede6b0d4836084169b1f9543dd32968239

Observation deafe58a-e056-440e-8781-a9f84efcd48a · inbound

Hierarchical Codec Diffusion for Video-to-Speech Generation cites this paper.

Hierarchical Codec Diffusion for Video-to-Speech Generation ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment

Reference 66

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T08:12:26.283645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T08:12:01.260833Z digest=sha256:d5ef607cdc6bafdaf724b5ec1847683c5d321cf878f9a6a01545442902ead766

Observation c8e55c73-53a0-4cd6-8931-5862900542ea · inbound

Can Large Audio Language Models Ignore Multilingual Distractors? An Evaluation of Their Selective Auditory Attention Capabilities cites this paper.

Can Large Audio Language Models Ignore Multilingual Distractors? An Evaluation of Their Selective Auditory Attention Capabilities ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-19T23:17:57.459132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-19T23:17:08.124240Z digest=sha256:5acdb964d2653d960d0d649c1f50f38f97ed3ba7e9339f7fe6d1554d18a23d87

Observation 6c4486fa-1b94-4290-bd04-2b0d6369226b · inbound

Feature-Aligned Speech Watermarking for Robustness to Reconstruction Distortions cites this paper.

Feature-Aligned Speech Watermarking for Robustness to Reconstruction Distortions ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T13:08:08.745045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T08:27:00.881610Z digest=sha256:bfbc4bb22bb8979fd01968833a4efbd19ccb924d07f494d158e3e2a7585c6b33

Observation f08e5579-001a-4a99-96e0-c7ceb8cf0d7d · inbound

Where Speech Enhancement Hurts Recognition: An Inference Time Polar Projection Diagnosis cites this paper.

Where Speech Enhancement Hurts Recognition: An Inference Time Polar Projection Diagnosis ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-14T06:37:05.257817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T06:37:05.257817Z digest=sha256:644db6b1394c978d79f48346e78a94e2622535a9098362c2babaad28c33bff21