Pith. sign in

Paper Citation Record · LEDGER

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment

As of 19 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 7 inbound Pith citation observations for arXiv:2506.19398.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.19398 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:11:24.534728Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:11:24.362189Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T13:08:08.743449Z

Reference resolution

45 of 45 outbound references displayed

  • verified exact1
  • verified fuzzy38
  • unresolved5
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 69262a98-01cc-4375-8be0-ce2617879210 · outbound

This paper cites While crucial for these applications, ac- curately processing speech is challenged by the often degraded quality of real-world audio.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment While crucial for these applications, ac- curately processing speech is challenged by the often degraded quality of real-world audio

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:25.131306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:11:24.356498Z digest=sha256:e3340f65472c559c7d83b50e69e4c4bb9a42692b107a495f63430ae9154b5fce

Observation 714fd3b6-e2d1-45c8-b151-2d020951eb2c · outbound

This paper cites ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T23:11:24.362189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:11:24.362189Z digest=sha256:87a67b46b1edb98d5c67737dc0b39bd57442df818437b26ad67ec71afae39f4d

Observation ffbe1fee-e561-42d0-b535-2900b9f5f465 · outbound

This paper cites Training strategies 3.1.1.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment Training strategies 3.1.1

Reference 3

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T23:11:25.112661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:11:24.366582Z digest=sha256:ec793c00a0e6b545257cc053ae6bdb735b832d2e13bd808dd8559c95ba56803a

Observation 72b8ed4b-36ef-4577-8a86-85b840ff6436 · outbound

This paper cites Beyond the presented evaluations, ClearerV oice- Studio is available for live demos on HuggingFace and Mod- elScope, enabling users to experiment with real-world record- ings.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment Beyond the presented evaluations, ClearerV oice- Studio is available for live demos on HuggingFace and Mod- elScope, enabling users to experiment with real-world record- ings

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:25.095263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:11:24.371021Z digest=sha256:53b495204ea7d9f22da8735c045df86a539981c908925acadb9787aff9254ac2

Observation 641f26e5-0e0d-4eb6-aae4-80b3166b30ec · outbound

This paper cites Deep learning for audio signal processing,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment Deep learning for audio signal processing,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:25.078020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:11:24.376338Z digest=sha256:1317a4779911c6d124a63a509cbf23c40d9d7683f90d8ef6e92c05b8b7b28f85

Observation 84304d9c-cd5d-423d-9c93-b1b942667db1 · outbound

This paper cites Mamba in Speech: Towards an Alternative to Self-Attention.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment Mamba in Speech: Towards an Alternative to Self-Attention

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T23:11:24.380674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:11:24.380674Z digest=sha256:73ae97190d8d3155beeeda361bad3ecd4f1ab38093720c69fa35ff97c049970e

Observation 6082d4d3-b631-4c69-ad1b-0d83a76af01a · outbound

This paper cites DeepMMSE: A deep learning approach to mmse-based noise power spectral density estimation,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment DeepMMSE: A deep learning approach to mmse-based noise power spectral density estimation,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:25.063117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:11:24.385499Z digest=sha256:ab358a4e6a5aa51a683662af8eaf2b10fc8c1740f86eb5ba3a90b985c1054871

Observation 2a35cea7-d23a-4042-93f2-bdc8a40e17dd · outbound

This paper cites SpeechBrain: A general-purpose speech toolkit,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment SpeechBrain: A general-purpose speech toolkit,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:25.047106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:11:24.389448Z digest=sha256:b97dfd41fbce485067754b8539edd25811ed4d43f53e3c14a5d2b1b201e04272

Observation 651b4301-b4e2-413f-8fbd-95a5ae5587c8 · outbound

This paper cites AudioSR: Versatile audio super-resolution at scale,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment AudioSR: Versatile audio super-resolution at scale,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.978245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:11:24.416110Z digest=sha256:b5e5c7de01a417391d99058cd67523b39131da76d08deb9785b18d39d67a3d54

Observation 0a8da161-6948-4863-9ea9-090d76f982a9 · outbound

This paper cites ESPnet: End-to-end speech processing toolkit,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment ESPnet: End-to-end speech processing toolkit,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:25.032797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:11:24.399073Z digest=sha256:a0de0dfbd37085403b1617e9ec444c54f055783b0f623b40e662c1dacbd2bf60

Observation d4baf6c7-b9da-4399-9bab-d6526be4a8df · outbound

This paper cites Summary on the multimodal information-based speech processing 2023 challenge,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment Summary on the multimodal information-based speech processing 2023 challenge,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:25.020116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:11:24.403507Z digest=sha256:037383870f2a4a31665c82bda69c12cff016873a3769345bb2b05fc1c3f37e60

Observation 51bb2732-9407-4c47-a76c-6aebd1a7ea79 · outbound

This paper cites As- teroid: the PyTorch-based audio source separation toolkit for re- searchers,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment As- teroid: the PyTorch-based audio source separation toolkit for re- searchers,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:25.007293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:11:24.407313Z digest=sha256:9ffa979824d1fe7558fb8794142fa8b06e72944f09297599f55d599bc9bb7ae5

Observation 95a7cd92-8c58-4079-bed2-62ae14bd37ca · outbound

This paper cites DeepFilterNet: Perceptually motivated real-time speech en- hancement,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment DeepFilterNet: Perceptually motivated real-time speech en- hancement,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.993239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:11:24.411596Z digest=sha256:80b0588522db1d97093b2c4da8cd94abb626fcc6ec7ee13c097d0c08cbe77aa3

Observation cc184939-26b3-4313-8d7f-57dec752ac13 · outbound

This paper cites Hifi-SR: A unified generative transformer-convolutional adversarial net- work for high-fidelity speech super-resolution,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment Hifi-SR: A unified generative transformer-convolutional adversarial net- work for high-fidelity speech super-resolution,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.915062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:11:24.435224Z digest=sha256:e49839ade0225a74336dc0bd76d18dd0d1c2993e533f68009a4ef47f0ccd45ef

Observation b444fbaa-fe93-4bfe-8501-0343628453ef · outbound

This paper cites FlowA VSE: Ef- ficient audio-visual speech enhancement with conditional flow matching,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment FlowA VSE: Ef- ficient audio-visual speech enhancement with conditional flow matching,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.964648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:11:24.420128Z digest=sha256:d4eed72e5d39b3706fcfd688e7f3340a66b967dea440dea888fdb8e681753375

Observation 60e16eda-03ee-4f93-af38-77890af77535 · outbound

This paper cites FRCRN: Boosting feature representation using frequency recurrence for monaural speech enhancement,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment FRCRN: Boosting feature representation using frequency recurrence for monaural speech enhancement,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.950815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:11:24.424200Z digest=sha256:9a820cfb75fc25b7a3c51ed2a062f0e4bac68e505e54bcad095acc6f5540e639

Observation f9a92ae9-a8d1-46e3-9c7c-c7226a044498 · outbound

This paper cites MossFormer2: Combin- ing transformer and rnn-free recurrent network for enhanced time- domain monaural speech separation,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment MossFormer2: Combin- ing transformer and rnn-free recurrent network for enhanced time- domain monaural speech separation,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.937368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:11:24.427496Z digest=sha256:d7f6bc92adbdd27df3762dde244595f62e3ed61f4e7e869a5337c028dbd15ef4

Observation 3f76367b-dfd3-42af-8818-5c8d59f4bbf4 · outbound

This paper cites MossFormer: Pushing the performance limit of monaural speech separation using gated single-head trans- former with convolution-augmented joint self-attentions,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment MossFormer: Pushing the performance limit of monaural speech separation using gated single-head trans- former with convolution-augmented joint self-attentions,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.926695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:11:24.431574Z digest=sha256:4e2222911f875642542c1b1d0a99363b0ccc58831d9dec94ba566177cd13c210

Observation 5658879f-5b6d-49b8-a6d9-67f991f513c3 · outbound

This paper cites NeuroHeed: Neuro-Steered Speaker Extraction using EEG Signals.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment NeuroHeed: Neuro-Steered Speaker Extraction using EEG Signals

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-06T23:11:24.586002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:11:24.452262Z digest=sha256:72cce8c6ba26d5e525bcb607ad1e0650606c79e520970b7cdb34395ddd3e83df

Observation de92d52c-4bf6-4860-8822-ba96854e3fe5 · outbound

This paper cites HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.902975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:11:24.438524Z digest=sha256:87ed1eab2905410ec7d251bd28c6a53b44e847ec5f4c0249a9e25c8543b7d96e

Observation aabc8ee1-2c97-4e64-aa86-a8da807dc826 · outbound

This paper cites Scenario-aware audio-visual TF- Gridnet for target speech extraction,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment Scenario-aware audio-visual TF- Gridnet for target speech extraction,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.891068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:11:24.441735Z digest=sha256:6f2c633a903923fa42f51762ed65596aca880ed83d1a95ae55ca226f56278bd8

Observation d8c239ae-1c8d-4c32-859d-5764628568ce · outbound

This paper cites Speaker extraction with co-speech gestures cue,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment Speaker extraction with co-speech gestures cue,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.878761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:11:24.445326Z digest=sha256:862dbf3ea142e02075927b2a7d25e52579a7877374b344ac69fdab21f0ebea7b

Observation bac71848-182d-4296-86e5-a0f824999631 · outbound

This paper cites SpEx+: A complete time domain speaker extraction network,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment SpEx+: A complete time domain speaker extraction network,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.866309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:11:24.448614Z digest=sha256:8f50bc8adced5998adb8794dcd1f636d1fbabe975ccce82f7ece628ca80afecc

Observation ec631b81-503b-41f1-8a50-b5ab29dfa231 · outbound

This paper cites DCCRN+: Channel-wise subband dccrn with snr estimation for speech enhancement,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment DCCRN+: Channel-wise subband dccrn with snr estimation for speech enhancement,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.800145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:11:24.472350Z digest=sha256:deebd5c5d4cf2b5fd333d3803b96b13d13d57843b0e3baf7e01555e24e288e9d

Observation 56a28c20-a421-4a64-9fef-8351e699a4d6 · outbound

This paper cites The interspeech 2020 deep noise suppression challenge: Datasets, subjective testing framework, and challenge results,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment The interspeech 2020 deep noise suppression challenge: Datasets, subjective testing framework, and challenge results,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.852653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:11:24.456214Z digest=sha256:c8308f4ebd8bd115fb74f5410ff09c111599539f8844771ffec4d063f8e992d0

Observation 3f12de00-69f9-4ec6-9155-361df8278134 · outbound

This paper cites CSTR VCTK Corpus: English multi-speaker corpus for cstr voice cloning toolkit (version 0.92),.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment CSTR VCTK Corpus: English multi-speaker corpus for cstr voice cloning toolkit (version 0.92),

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.840151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:11:24.459983Z digest=sha256:bdd0e677687ac072c5982c150099396bde1ed8f0b9dd734963abd86436e07fb1

Observation 659d3a06-5962-4ebf-b2e5-03c71bfa0038 · outbound

This paper cites Audio set: An ontology and human-labeled dataset for audio events,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment Audio set: An ontology and human-labeled dataset for audio events,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.825770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:11:24.463692Z digest=sha256:6bce4b0bb32380789b0a8aa6f32428c2a45c154b5edf872ce5de3962ec49fe29

Observation 95fdaf25-0bc1-40ec-a534-9a7e5ab7d920 · outbound

This paper cites DEMAND: a collection of multi-channel recordings of acoustic noise in diverse environments,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment DEMAND: a collection of multi-channel recordings of acoustic noise in diverse environments,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.812415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:11:24.468069Z digest=sha256:fa6836b901269249cafb82ddd5629829c0dd712463c0858dd0d6814afceccb1e

Observation 84ac29bb-be00-47d6-9960-e826474b4b6e · outbound

This paper cites Phase- sensitive and recognition-boosted speech separation using deep recurrent neural networks,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment Phase- sensitive and recognition-boosted speech separation using deep recurrent neural networks,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.743338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:11:24.492570Z digest=sha256:5157f8aea8d73585cf70e145f4903f73e7021d735059f233580e3aa6abf887eb

Observation ae34a824-1169-40b3-8f35-5e3e3c074c7b · outbound

This paper cites A mask free neural network for monaural speech enhancement,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment A mask free neural network for monaural speech enhancement,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.788296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:11:24.476740Z digest=sha256:637734ef5052702f17ffbff4585d71d9f3d9c622658250b83036cd5adac22db0

Observation 33a9f56d-411a-468a-9246-0760cb623b87 · outbound

This paper cites TridentSE: Guiding speech enhancement with 32 global tokens,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment TridentSE: Guiding speech enhancement with 32 global tokens,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.777289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:11:24.480847Z digest=sha256:3569d44f652f961af29b829ede6727fe844b8f2be79f734ee2f7863466b6020d

Observation 5ab46bd6-5543-4693-94c5-762f92a30bea · outbound

This paper cites LibriTTS: A corpus derived from librispeech for text- to-speech,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment LibriTTS: A corpus derived from librispeech for text- to-speech,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.766236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:11:24.484737Z digest=sha256:e6a3d4e17c7350f86952565641933b46a9ed307514d1f924b21daa071646c22a

Observation 37ab05c7-7c56-42e3-8b87-94489e6e4ff0 · outbound

This paper cites Explor- ing strategies for training deep neural networks,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment Explor- ing strategies for training deep neural networks,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.754403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:11:24.488634Z digest=sha256:8a5b612502b97a7055a67cb5bb8d1ebe47d5ef198b65379edae02b063b7b149b

Observation 1a55352d-409e-441e-8e99-b07c8b85bb82 · outbound

This paper cites An efficient encoder-decoder archi- tecture with top-down attention for speech separation,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment An efficient encoder-decoder archi- tecture with top-down attention for speech separation,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.683161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:11:24.511237Z digest=sha256:3b50c5214df794bdb51b436abab479c288ad95f8a5be0d8ec79a3b476a07af07

Observation a3042b7b-1c2e-4e49-9fcb-38ad66742e36 · outbound

This paper cites CMGAN: Conformer-based metric gan for speech enhancement,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment CMGAN: Conformer-based metric gan for speech enhancement,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.730747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:11:24.496146Z digest=sha256:f455e0dbe0d70be1793bd28b36f37fc59670e8ddeef365e24b807f756ef202f8

Observation d3c53e10-c4f2-4b40-953d-c1a1122527eb · outbound

This paper cites Permutation invari- ant training of deep models for speaker-independent multi-talker speech separation,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment Permutation invari- ant training of deep models for speaker-independent multi-talker speech separation,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.719254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:11:24.499581Z digest=sha256:df28d00ec79eb55734d347429ddc77c62443bce29487706ef6e64f11fa1d7c89

Observation 9f24950c-a1cd-4283-9cd1-8ab001d8e2c4 · outbound

This paper cites Dual-Path RNN: Efficient long sequence modeling for time-domain single-channel speech sepa- ration,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment Dual-Path RNN: Efficient long sequence modeling for time-domain single-channel speech sepa- ration,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.706871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:11:24.503833Z digest=sha256:318ce41f3e30ae596802f517eaef4cdb5aa1382e8db2bbd9ab633d3a3951f295

Observation d376de3c-6851-490f-a8d7-be8a5d857958 · outbound

This paper cites Attention is all you need in speech separation,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment Attention is all you need in speech separation,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.695833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:11:24.507563Z digest=sha256:34b9e5327cab76028083c762ceaadcc8baeb3447d132dce0f6cada14d4f27856

Observation 22a3c68e-0b4a-41ee-a4f2-14383c8ead30 · outbound

This paper cites Selective listening by synchronizing speech with lips,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment Selective listening by synchronizing speech with lips,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.640308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:11:24.530987Z digest=sha256:0f86afb1513dd0e6625eeb57ff5906e159733db4c5671380f0d303bd47ba82fe

Observation 07b9b447-c93c-4a44-9e37-c8d9f9885845 · outbound

This paper cites TF-GridNet: Making time-frequency domain models great again for monaural speaker separation,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment TF-GridNet: Making time-frequency domain models great again for monaural speaker separation,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T23:11:24.515143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:11:24.515143Z digest=sha256:6653843bf8fc4ee44bc4b7d2054a3e95afaede01f0c7f72a943c27a7571a8207

Observation 100526fc-2a2d-40f2-a4e1-ed08c42b5971 · outbound

This paper cites SPMamba: State-space model is all you need in speech separation.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment SPMamba: State-space model is all you need in speech separation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T23:11:24.519069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:11:24.519069Z digest=sha256:4d061cf74cf67b9b4d3ff27a4174a6d49acd5b01f837a5a1d7a72d4f92be02be

Observation a9606353-6c7f-47a4-a8e3-39ca7e2241b3 · outbound

This paper cites Time domain audio visual speech separation,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment Time domain audio visual speech separation,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.663849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:11:24.523257Z digest=sha256:f8626fbb027a65c06cdf5d055cce42053e6f0d0dd36973839565f5e4cd91cc88

Observation f5854460-642f-4afb-9029-0d21a5d977fe · outbound

This paper cites MuSE: Multi-modal target speaker extraction with visual cues,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment MuSE: Multi-modal target speaker extraction with visual cues,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.651853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:11:24.526741Z digest=sha256:846e94b3e3827ec00fce669b4ad57f4ea533c56bd8fc544324ee4316ba3a45fa

Observation ea871c47-8135-4ac8-a3a6-cdb06e9d4e2a · outbound

This paper cites USEV: Universal speaker extraction with visual cue,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment USEV: Universal speaker extraction with visual cue,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.629442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:11:24.534728Z digest=sha256:f544c09eef471b36a5c1dcc85c3cc4b3b9ee194c48329318c637c09510a5d74a

Observation 3ae2ef09-51e3-4364-9d36-43eabf7fb5d6 · outbound

This paper cites SpeechBrain: A General-Purpose Speech Toolkit.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment SpeechBrain: A General-Purpose Speech Toolkit

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T23:11:24.394051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:11:24.394051Z digest=sha256:e64362c5aa4a4339b9768410c57043044dc975c77212ec44e0fcc6a09412ef38

Pith citing papers

Observation 714fd3b6-e2d1-45c8-b151-2d020951eb2c · inbound

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment cites this paper.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T23:11:24.362189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:11:24.362189Z digest=sha256:87a67b46b1edb98d5c67737dc0b39bd57442df818437b26ad67ec71afae39f4d

Observation 02546d07-92d4-4255-a078-45e13b8b6517 · inbound

CueNet: Robust Audio-Visual Speaker Extraction through Cross-Modal Cue Mining and Interaction cites this paper.

CueNet: Robust Audio-Visual Speaker Extraction through Cross-Modal Cue Mining and Interaction ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T19:39:50.604200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:39:50.604200Z digest=sha256:3c0a0c5cf99600324616efb767f1d3613b2f85bcaf401eb7f24aec693618602b

Observation 54563704-4c79-430a-bd33-33beede6abb6 · inbound

Beyond Monologue: Interactive Talking-Listening Avatar Generation with Conversational Audio Context-Aware Kernels cites this paper.

Beyond Monologue: Interactive Talking-Listening Avatar Generation with Conversational Audio Context-Aware Kernels ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:56:01.909103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T15:16:53.603971Z digest=sha256:de3eafdcc3dfd03b65deb259155a303d9c25ea3f6b6ccc31d0e6f886a8d52551

Observation deafe58a-e056-440e-8781-a9f84efcd48a · inbound

Hierarchical Codec Diffusion for Video-to-Speech Generation cites this paper.

Hierarchical Codec Diffusion for Video-to-Speech Generation ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment

Reference 66

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T08:12:26.283645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T08:12:01.260833Z digest=sha256:66bb35e9a4a1714c96eac565636b9edf68735b5affa4dbfdac0f98eb25a0ecf1

Observation c8e55c73-53a0-4cd6-8931-5862900542ea · inbound

Can Large Audio Language Models Ignore Multilingual Distractors? An Evaluation of Their Selective Auditory Attention Capabilities cites this paper.

Can Large Audio Language Models Ignore Multilingual Distractors? An Evaluation of Their Selective Auditory Attention Capabilities ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-19T23:17:57.459132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-19T23:17:08.124240Z digest=sha256:27f0b0a7356cb7976e697a519f8caa6bc48eb1898666a7c3270d290a56bd5172

Observation 6c4486fa-1b94-4290-bd04-2b0d6369226b · inbound

Feature-Aligned Speech Watermarking for Robustness to Reconstruction Distortions cites this paper.

Feature-Aligned Speech Watermarking for Robustness to Reconstruction Distortions ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T13:08:08.745045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T08:27:00.881610Z digest=sha256:e80d3b423ea37f26c044f0dfe785426e649d9c044acae96766d1cf4877da6b38

Observation f08e5579-001a-4a99-96e0-c7ceb8cf0d7d · inbound

Where Speech Enhancement Hurts Recognition: An Inference Time Polar Projection Diagnosis cites this paper.

Where Speech Enhancement Hurts Recognition: An Inference Time Polar Projection Diagnosis ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-14T06:37:05.257817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T06:37:05.257817Z digest=sha256:b5ea5a339bc6d93bd3a6676c055be1f41f2520a839029744091913e167567a59