Pith. sign in

Paper Citation Record · LEDGER

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline

As of 19 August 2026, this Paper Citation Record lists 80 of 80 outbound references and 3 inbound Pith citation observations for arXiv:2505.19314.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.19314 v3

Coverage vector

measured 80 of 80 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:22:07.475291Z

measured 83 of 83 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T14:20:49.132229Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

80 of 80 outbound references displayed

  • verified exact7
  • verified fuzzy55
  • unresolved17
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d86250cd-c963-4ac1-9256-ddbe38fad0ee · outbound

This paper cites The cocktail-party problem revisited: early pro- cessing and selection of multi-talker speech,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline The cocktail-party problem revisited: early pro- cessing and selection of multi-talker speech,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:16.316621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:22:07.154451Z digest=sha256:19d649ce817cf14dcc72869fec71b3879dc3d99bc49e22f15bec4e470f41b338

Observation d24a1ce8-4e6c-4d93-9999-a8c2e4cf756a · outbound

This paper cites Neural target speech extraction: An overview,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Neural target speech extraction: An overview,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:16.167867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:22:07.158760Z digest=sha256:1357bcdc72e28571662b61c5944f8b4d4f2c5f2a65457e7c7d79a8fd3d9325b6

Observation a5f4c12f-df00-4fbe-9813-b06d63e71704 · outbound

This paper cites Neural spatial filter: Target speaker speech separation assisted with directional information,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Neural spatial filter: Target speaker speech separation assisted with directional information,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:16.021303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:22:07.163056Z digest=sha256:242d1f27bcba5a6dbd347d40d1992e13d7cee4ba4b76c5b487145f50e439b4d8

Observation 8b003eec-47d2-42bc-b634-cecfc738d80d · outbound

This paper cites Far-field location guided target speech extraction using end-to-end speech recognition objectives,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Far-field location guided target speech extraction using end-to-end speech recognition objectives,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:15.913298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:22:07.168216Z digest=sha256:bfa224b478125069207ad00dded9c3bc3026213981b79aa7e778236fbbe747b4

Observation 1eedf57c-0046-4639-a424-985e93779fbf · outbound

This paper cites Looking to listen at the cocktail party: a speaker-independent audio-visual model for speech separation,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Looking to listen at the cocktail party: a speaker-independent audio-visual model for speech separation,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:15.435237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:22:07.171996Z digest=sha256:026215f7375e82dd4e901176d66bd12c19d639ae7585b42d132187e508f311d5

Observation 27460b18-8024-4c28-a89d-4479dbe0a0ef · outbound

This paper cites Conceptbeam: Concept driven target speech extraction,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Conceptbeam: Concept driven target speech extraction,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:15.283829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:22:07.176146Z digest=sha256:62389760a54ad22e10b9aa5a4ecb93f8445acab8f33da00e367bf299f5a1135a

Observation 693bb2b9-cf52-4e47-8d02-945d341d136c · outbound

This paper cites V oicefilter: Targeted voice separation by speaker-conditioned spectrogram masking,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline V oicefilter: Targeted voice separation by speaker-conditioned spectrogram masking,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:15.164815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:22:07.180844Z digest=sha256:9069b96e02b4f2ac580393d6efa52d6104ac73b2e907ae036b0e5205b4921f8a

Observation 930acadf-25cf-4460-b1bf-70371d4bf6ca · outbound

This paper cites Speakerbeam: Speaker aware neural network for target speaker extraction in speech mixtures,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Speakerbeam: Speaker aware neural network for target speaker extraction in speech mixtures,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:15.002077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:22:07.184966Z digest=sha256:797ec66e73208b256052b16c0889b7c362a0315a9f5bb9c0fdd25b03f4ef383f

Observation 33604da3-1634-4163-ac55-ae8543afb007 · outbound

This paper cites Wesep: A scalable and flexible toolkit towards generalizable target speaker extraction,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Wesep: A scalable and flexible toolkit towards generalizable target speaker extraction,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:14.875258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:22:07.188682Z digest=sha256:4bfad8e0546ddf35a31089fc501a6fddf5b212861c862c3d9b81adecaed6b9f2

Observation 460ed315-7fe9-41ff-b469-ddcbc96234dd · outbound

This paper cites Target confusion in end-to-end speaker extraction: Analysis and approaches,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Target confusion in end-to-end speaker extraction: Analysis and approaches,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:14.758242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:22:07.192850Z digest=sha256:12c9fac245ffe9c2850aa70c178f6228e92793b12f062b28b3c14e5a2ef9ba2a

Observation f22eb8bc-4f60-4f6c-803d-194c16e497aa · outbound

This paper cites Dpccn: Densely-connected pyramid complex convolutional network for robust speech separation 11 and extraction,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Dpccn: Densely-connected pyramid complex convolutional network for robust speech separation 11 and extraction,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:14.633307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:22:07.196434Z digest=sha256:d73fce15ca2d73b5afc1631cd8534052f0016d471a68fcea506669c00f162397

Observation a5f74896-a48a-44e2-a4d5-4a405cdc5bdc · outbound

This paper cites Improving target sound extraction with timestamp information,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Improving target sound extraction with timestamp information,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:14.516695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:22:07.200583Z digest=sha256:989820cb9cd953e98f423d49e2d38c0ee10884412fb88f616ddd08ff08d334db

Observation 610f6f0c-164e-427f-ad8d-55bce01922a0 · outbound

This paper cites WeSep: A Scalable and Flexible Toolkit Towards Generalizable Target Speaker Extraction.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline WeSep: A Scalable and Flexible Toolkit Towards Generalizable Target Speaker Extraction

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:07.204106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:07.204106Z digest=sha256:22fde1a83bff5a200748bf223e9744e23fc61572a5d8bfd853a86c3e1be21b8a

Observation 611ca561-de03-4463-8427-7df5c5ba0e9c · outbound

This paper cites Spex: Multi-scale time domain speaker extraction network,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Spex: Multi-scale time domain speaker extraction network,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:14.357625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:22:07.208667Z digest=sha256:0074fbd10015118f94c0321ddabe0b76ab99bcf0b1218a27c4c793da70891980

Observation 3858436a-b8b9-43ec-9a1f-78e582e3d462 · outbound

This paper cites Spex+: A complete time domain speaker extraction network,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Spex+: A complete time domain speaker extraction network,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:14.233683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:22:07.212662Z digest=sha256:17dbde88fab58e705defde2ed0f26a015390453e0a8acd24b5f07b2df773321a

Observation 85db493b-472e-4c86-895d-e4730dc4985c · outbound

This paper cites X-SEPFORMER: end-to- end speaker extraction network with explicit optimization on speaker confusion,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline X-SEPFORMER: end-to- end speaker extraction network with explicit optimization on speaker confusion,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:14.115190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:22:07.216422Z digest=sha256:945c193242b140d46fb1ba3e7255ca51befdd27149bb94017eaca174572da6e3

Observation 8da9cead-1b4d-4c0c-877a-787dddb08cab · outbound

This paper cites X-tf-gridnet: A time-frequency domain target speaker extraction network with adaptive speaker embedding fusion,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline X-tf-gridnet: A time-frequency domain target speaker extraction network with adaptive speaker embedding fusion,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:13.979985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:22:07.221377Z digest=sha256:1259d4e0d455f13278c0274aef0fa5992675c8ef00c2a408dc8cad7be2c013ac

Observation 1de9266a-258a-4493-b63b-e2782b33b22b · outbound

This paper cites USEF-TSE: Universal Speaker Embedding Free Target Speaker Extraction.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline USEF-TSE: Universal Speaker Embedding Free Target Speaker Extraction

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:22:07.834953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:22:07.225317Z digest=sha256:2b337bc31e25c4d4c3734dd399f7ba312cb692c6be3a2ebe649c4fbc5853ae58

Observation ed61ae57-6b30-475d-b20a-a71d244102f2 · outbound

This paper cites Target speech extraction with conditional diffusion model,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Target speech extraction with conditional diffusion model,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:13.866038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:22:07.230095Z digest=sha256:b1c33fb8db4f8ef7e5b914616c6d520e8a16d6984f67e8e5eade3d1f0f7fbcca

Observation 3f2ee616-3578-4865-9a29-f62fd92bdbb1 · outbound

This paper cites Noise-robust Speech Separation with Fast Generative Correction.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Noise-robust Speech Separation with Fast Generative Correction

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:22:07.817493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:22:07.234297Z digest=sha256:50c0e8942de97a06521d332947c1ea2a0c3dd872f6c2af9b7b33332525fde1e0

Observation 9f531854-3c29-43d1-bf74-867007a691f8 · outbound

This paper cites Speech enhancement and dereverberation with diffusion-based generative models,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Speech enhancement and dereverberation with diffusion-based generative models,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:13.683036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:22:07.239647Z digest=sha256:62d864a3d4f25efc0ab751b89550d07b851e5f1e0c43081b27c6f669338e58c6

Observation 1c4af11a-89ad-4505-9ea5-a24441951168 · outbound

This paper cites Diffusion-based generative speech source separation,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Diffusion-based generative speech source separation,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:07.243716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:07.243716Z digest=sha256:9e871dccc334e65abf2df755c3188ca9839bab8a9f3339b440ee26a13e0d7ed9

Observation 63742595-d3a9-4838-bd75-3a4f27419613 · outbound

This paper cites Generative pre-training for speech with flow matching,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Generative pre-training for speech with flow matching,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:13.475291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:22:07.248153Z digest=sha256:14fd863dedff18b358000f23c236260de3226b1cbcb5b4c09a5b51a6d047c565

Observation ffb056b2-5e5a-4357-a714-17287e0219e7 · outbound

This paper cites Metis: A Foundation Speech Generation Model with Masked Generative Pre-training.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Metis: A Foundation Speech Generation Model with Masked Generative Pre-training

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:07.252044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:07.252044Z digest=sha256:7f8e21bbb6bdb5b07d9c991e68d80a15777cc9fb506cf68e5f46cf89110fc273

Observation af0f0304-7d6e-4afa-91ce-5577d88b2cbd · outbound

This paper cites SoloAudio: Target Sound Extraction with Language-oriented Audio Diffusion Transformer.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline SoloAudio: Target Sound Extraction with Language-oriented Audio Diffusion Transformer

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:07.256774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:07.256774Z digest=sha256:ea3ea0caed2b51f711cd6ec9c9910ea54e158f3e75e1fc77a6559a63238176e4

Observation bf62dc7e-f961-46a4-be01-5f8568be862f · outbound

This paper cites Separate And Diffuse: Using a Pretrained Diffusion Model for Improving Source Separation.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Separate And Diffuse: Using a Pretrained Diffusion Model for Improving Source Separation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:07.261084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:07.261084Z digest=sha256:847f17f483a09d6282d9d41a8499c46e4e629b90ea8fe26264538384662013e2

Observation eadbdd15-2a50-446f-88e2-1527b08990b5 · outbound

This paper cites Attention is all you need,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Attention is all you need,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:07.265276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:07.265276Z digest=sha256:9aae17d9e365e469299ac5b398cd4adebee2dd09447e983971bfc9dcd88d833a

Observation 17e636f9-cabc-4df8-b41f-2cbca8dbe101 · outbound

This paper cites Large language model based generative error correction: A challenge and baselines for speech recognition, speaker tagging, and emotion recognition,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Large language model based generative error correction: A challenge and baselines for speech recognition, speaker tagging, and emotion recognition,

Reference 28

Resolution
verified exact
raw_fallback, observed 2026-08-07T14:22:07.780630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:22:07.269521Z digest=sha256:3d198ba9d1a797fff57c1d1927b8171eba4ec9f0fc09944f45f4d93a22009ddb

Observation 58f778a2-4cd6-4848-b207-4291a63e5542 · outbound

This paper cites SepALM: Audio Language Models Are Error Correctors for Robust Speech Separation.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline SepALM: Audio Language Models Are Error Correctors for Robust Speech Separation

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:22:07.710817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:22:07.274041Z digest=sha256:d4a9fa48bbd11707ac7fd42f8fd2509613b347135cde557feece4dd39206e388

Observation ff992483-89bd-4c28-8710-1079063a1ac4 · outbound

This paper cites LibriMix: An Open-Source Dataset for Generalizable Speech Separation.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline LibriMix: An Open-Source Dataset for Generalizable Speech Separation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:07.278378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:07.278378Z digest=sha256:383eb5a294786b2c778edd3699fccf84c4cc5c481e0b467ab799704e8dd0a68e

Observation 335f9c72-14f9-4d0b-9d3d-7ba3e47b1cba · outbound

This paper cites Target speech extraction with conditional diffusion model,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Target speech extraction with conditional diffusion model,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:13.348423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:22:07.282398Z digest=sha256:287f80eb52209fb8490bed56d8e47b798d8c7c7b4b13234e301bcac4b8bfd891

Observation e99c4cc7-f709-4a9a-a5f6-b9cb4bf63e49 · outbound

This paper cites Dpm-tse: A diffusion probabilistic model for target sound extraction,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Dpm-tse: A diffusion probabilistic model for target sound extraction,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:13.098187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:22:07.286482Z digest=sha256:bd6d60759aad80915c303187fd1e35a7b594368476e62df635f0698ecd7f9353

Observation 73ea1afc-061b-4f60-a9f9-019697bca4fb · outbound

This paper cites Diffusion- based generative speech source separation,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Diffusion- based generative speech source separation,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:12.750074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:22:07.290826Z digest=sha256:eb0746cfee37d6530c2353dd13a57f42d0ce6eb861d57016867d88b71e8e061f

Observation 246e030c-ca12-4386-8f20-fdd375bd3a4e · outbound

This paper cites Generative Speech Foundation Model Pretraining for High-Quality Speech Extraction and Restoration.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Generative Speech Foundation Model Pretraining for High-Quality Speech Extraction and Restoration

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:22:07.683925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:22:07.295184Z digest=sha256:ceca0a446668ccb9eb888af5e22a6c0c35474fecf14372ca2f4afe7305be8f07

Observation aa515a27-f2d8-4c0c-99c9-7d5b713dd5b7 · outbound

This paper cites Generation- based target speech extraction with speech discretization and vocoder,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Generation- based target speech extraction with speech discretization and vocoder,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:12.458042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:22:07.299420Z digest=sha256:03e92415e4aaa787a8ed0dd12d2b9b1e0ae2726e232e6758a837339ec74c4df7

Observation 25dbd4bf-1922-457e-abd1-dd3165775919 · outbound

This paper cites Enhancing Intelligibility for Generative Target Speech Extraction via Joint Optimization with Target Speaker ASR.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Enhancing Intelligibility for Generative Target Speech Extraction via Joint Optimization with Target Speaker ASR

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:22:07.667941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:22:07.303374Z digest=sha256:89483f388d2da8d3b6956d3b1955d336e898b1acd5575595cb17af8135855912

Observation 72098b94-b6c8-4c80-a0a5-18864715c414 · outbound

This paper cites Diffusion-based signal refiner for speech separation,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Diffusion-based signal refiner for speech separation,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:07.307054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:07.307054Z digest=sha256:14030b2f1e0ab98f8ad93900cd147d9530a479b0994943f2d1a96ece7b84433a

Observation 4842be6c-412c-46b7-83ea-3ef0264395a8 · outbound

This paper cites Storm: A diffusion-based stochastic regeneration model for speech enhancement and dereverberation,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Storm: A diffusion-based stochastic regeneration model for speech enhancement and dereverberation,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:12.042351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:22:07.311336Z digest=sha256:3dd8a548f3096f6a51155b6b996faf59abbaab7597adc24852f6e3dfcc482f90

Observation 22f71e89-d785-415d-b410-649fa43d864b · outbound

This paper cites Ddtse: Discriminative diffusion model for target speech extraction,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Ddtse: Discriminative diffusion model for target speech extraction,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:11.826875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:22:07.315289Z digest=sha256:ad60fd3313c8042ef5ffb2b775aa268e9946bdb0c118e1bd601ea1a2453622d6

Observation 4c8713c5-9933-4c50-ba86-8f2066bbe735 · outbound

This paper cites Speaker-aware neural network based beamformer for speaker extraction in speech mixtures,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Speaker-aware neural network based beamformer for speaker extraction in speech mixtures,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:11.286750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:22:07.319275Z digest=sha256:07c567d84389a729c23c31fd879356c2041332546da79c50867c9fe4848e6cb7

Observation 52a51c3a-7be6-4adb-80a6-4269e55219eb · outbound

This paper cites X- vectors: Robust DNN embeddings for speaker recognition,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline X- vectors: Robust DNN embeddings for speaker recognition,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:10.825851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:22:07.323186Z digest=sha256:69abaca02ca765c7749e3e7b1cd76d912d19ae77f16625a0b3efc3e782d270be

Observation 4f347263-8123-49eb-a228-bd5c259de03c · outbound

This paper cites Probing self-supervised learning models with target speech extraction,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Probing self-supervised learning models with target speech extraction,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:10.383416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:22:07.326526Z digest=sha256:40cce33d0c14eec3c878e792ab042485b2e8337ee6588e72882ccd166dfd67be

Observation 70347c62-4802-4526-b4b1-884eaeda3a8c · outbound

This paper cites Target speech extraction with pre-trained self-supervised learning models,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Target speech extraction with pre-trained self-supervised learning models,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:10.184130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:22:07.329756Z digest=sha256:b446b0965fdec42ba983c284f4ca9c6b94c75484cb2017bbbcce8fbf9b1dcf6a

Observation a6131e0a-85aa-4405-8e71-6099872bfdb6 · outbound

This paper cites Smma-net: An audio clue-based target speaker extraction network with spectrogram matching and mutual attention,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Smma-net: An audio clue-based target speaker extraction network with spectrogram matching and mutual attention,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:10.039313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:22:07.333165Z digest=sha256:2ea44127af776d43ae8c37ac75a33110a6cfdb766d63eb2466eb6050d5cc4f1c

Observation dc6868b9-d78f-42f0-848f-bb5f43732e07 · outbound

This paper cites Target speaker extraction by directly exploiting contextual information in the time-frequency domain,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Target speaker extraction by directly exploiting contextual information in the time-frequency domain,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:09.902393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:22:07.336782Z digest=sha256:54e5ee2a84d5340342d94d976a00b7306c1bd8ef5087e302bd40eeead1582150

Observation f0f6c734-18fe-4623-9839-4327c25098b4 · outbound

This paper cites Target speaker extraction with ultra-short reference speech by VE-VE framework,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Target speaker extraction with ultra-short reference speech by VE-VE framework,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:09.751317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:22:07.340177Z digest=sha256:e50f4b65344ee2104334a133851649578b36af405476013f8a6de770611237cf

Observation cfcbd0d9-aa6b-4853-b2a3-7dd89ec2c885 · outbound

This paper cites Sef-net: Speaker embedding free target speaker extraction network,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Sef-net: Speaker embedding free target speaker extraction network,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:09.585458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:22:07.344503Z digest=sha256:2c69828195229f2d05588f414881d2aa779262d22b7f7dd16dbf35c6e163c257

Observation 6dc8bc8b-9d24-4d81-9c76-b40ddc55d071 · outbound

This paper cites Common diffusion noise schedules and sample steps are flawed,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Common diffusion noise schedules and sample steps are flawed,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:09.460733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:22:07.347592Z digest=sha256:43296a170d67c13297c9466a40becd4299cf8c70043d7b429aaa26656f04d5fb

Observation 23f2da4c-8b3e-435f-ad39-abb56ef48d70 · outbound

This paper cites Progressive distillation for fast sampling of diffusion models,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Progressive distillation for fast sampling of diffusion models,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:09.360479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:22:07.350993Z digest=sha256:12a9a26c840818f635f074addf16e0085271bc2ff454641bda388afdde14327a

Observation f157c4e6-1f2e-4093-a974-7438ab5c6e9d · outbound

This paper cites High- fidelity audio compression with improved RVQGAN,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline High- fidelity audio compression with improved RVQGAN,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:09.211137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:22:07.354792Z digest=sha256:e9a91f86a9a428ae0c7560969eff1286e2392807ad4c3a63f40997a8ab4aa34f

Observation 010a6aa9-89f3-47cb-9fbb-8b2e41a8ab8f · outbound

This paper cites Stable Audio Open.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Stable Audio Open

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:07.358306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:07.358306Z digest=sha256:af712a05b4dd9636303f95e4b80f89d24f99e4f45e0712713ba1f34c9173b1d4

Observation 96e4134f-1562-4fc0-96fe-b16912434989 · outbound

This paper cites EzAudio: Enhancing Text-to-Audio Generation with Efficient Diffusion Transformer.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline EzAudio: Enhancing Text-to-Audio Generation with Efficient Diffusion Transformer

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:07.362796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:07.362796Z digest=sha256:a1b0c6e53dcd32949d12c3ae50679a5051727659feb59f603439f4dd94173cd0

Observation 15bc5cc7-3c30-4004-a002-2d6a7760ec0b · outbound

This paper cites Tf- gridnet: Integrating full- and sub-band modeling for speech separation,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Tf- gridnet: Integrating full- and sub-band modeling for speech separation,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:09.018402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:22:07.366834Z digest=sha256:8a3fdf882bc14b6159f48442f385cc50ded8fe69f4a00dd6fc0173d5ad956432

Observation e6a27d0f-83b0-47a4-b831-3b0e46ed4003 · outbound

This paper cites SPMamba: State-space model is all you need in speech separation.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline SPMamba: State-space model is all you need in speech separation

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:07.370792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:07.370792Z digest=sha256:40c925fbfa1e2cdf4238e63eec4d9f81e786f87b3f8cb45bb82f7ea729d566e9

Observation 11054119-9270-47de-9613-91804548ef4b · outbound

This paper cites Complex ratio masking for monaural speech separation,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Complex ratio masking for monaural speech separation,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:08.871580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:22:07.374906Z digest=sha256:c829b923e5bcf562917bf84608e8f0db43295e09183f1547d14463e727e5f74a

Observation aadc30e0-17e9-4e6e-ba89-151dfc7f0c01 · outbound

This paper cites auraloss: Audio focused loss functions in pytorch,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline auraloss: Audio focused loss functions in pytorch,

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:07.379255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:07.379255Z digest=sha256:8def2050a0e056fa0cdeee6fe81e7a22df19c80b84c5acc39447ed3663d552c9

Observation 54a7bc8f-5781-4c69-b7e1-ad5fe4ce80f4 · outbound

This paper cites High fidelity neural audio compression,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline High fidelity neural audio compression,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:08.674926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:22:07.383128Z digest=sha256:0ad7f79d790f292e64908c644831f9a257882e2b5e6c21a897c5dc4bc7d62fcb

Observation de2e9680-5b1a-41a2-9c2f-441ef16b5157 · outbound

This paper cites Scalable diffusion models with transformers,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Scalable diffusion models with transformers,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:08.454405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:22:07.387236Z digest=sha256:190c6cdbab5a722e92521f63a1825441ce1d0bdb4fa6fcf9c5d2d59bf3d831ba

Observation 31cfd854-1be9-4e3c-804e-60f026a25267 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline An image is worth 16x16 words: Transformers for image recognition at scale,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:08.264101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:22:07.390920Z digest=sha256:7fdf2b4b8362d57e57d96a05f3f6ec5efebec1818c9810b074af7bfd369fbc70

Observation a0bfd558-bae4-425f-91e7-49c5aed8eebf · outbound

This paper cites Sigmoid-weighted linear units for neural network function approximation in reinforcement learning,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Sigmoid-weighted linear units for neural network function approximation in reinforcement learning,

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:07.395030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:07.395030Z digest=sha256:6b3c6474d11b48fb38d6b287c18f3e48dec677d9ef2ce7a73124314db70d118c

Observation 2f748628-177d-4a38-b1bd-f2f82efc59c0 · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Roformer: Enhanced transformer with rotary position embedding,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:08.143124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:22:07.399020Z digest=sha256:2af7580f2ba37db9bdbe8494665aa888485a68647752339b422526f6dd43408d

Observation 0edd31f9-8f3b-4bb4-9518-44c1cbd8072d · outbound

This paper cites Single- channel multi-speaker separation using deep clustering,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Single- channel multi-speaker separation using deep clustering,

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:08.040444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:22:07.402914Z digest=sha256:d4e94a5fd03c09dfd29eb1c4f3202a942cf827e0821a5557cf8aa15aa707619a

Observation 75bdcf37-727d-4132-8a8d-0786387e5952 · outbound

This paper cites Conv-tasnet: Surpassing ideal time-frequency magnitude masking for speech separation,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Conv-tasnet: Surpassing ideal time-frequency magnitude masking for speech separation,

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:08.028706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:22:07.406508Z digest=sha256:e6b6e7a4b75f62ec15b15677e54e6e80bf5486cd4e31c63f07112da17adc4f12

Observation 4918245f-567b-4278-be56-986fac214ce5 · outbound

This paper cites Wham!: Extending speech separation to noisy environments,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Wham!: Extending speech separation to noisy environments,

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:08.017035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:22:07.410086Z digest=sha256:4fa1d370e922c0fee7163be2ed1f34391ed8064e520557467e0ed20ecc178991

Observation 79627c97-0cab-4265-9437-c08bde14e8bb · outbound

This paper cites Librispeech: An ASR corpus based on public domain audio books,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Librispeech: An ASR corpus based on public domain audio books,

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:08.005422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:22:07.413724Z digest=sha256:81a66bf958ee48992b6df48c8932ca63d4bb36b6c94df2503bf20d225a0d2e0a

Observation 73d7aa18-fb73-4828-97c6-8dd42e67a898 · outbound

This paper cites Improving speaker discrimination of target speech extraction with time-domain speakerbeam,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Improving speaker discrimination of target speech extraction with time-domain speakerbeam,

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:07.993118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:22:07.417535Z digest=sha256:1f2d2163b19e9df89dcf0e1790b497a8c6a3516beeded71f8891ad1df7266662

Observation 7deb9100-49d1-4a44-81eb-7d1dac4d1144 · outbound

This paper cites MUSAN: A Music, Speech, and Noise Corpus.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline MUSAN: A Music, Speech, and Noise Corpus

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:07.421318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:07.421318Z digest=sha256:4db3021fa69214cb9392805c9e3f164709ebb07517621fb88a077e99d6dda60e

Observation d5dd377b-9cfe-4767-b06e-9e8dc831bfbf · outbound

This paper cites Multichannel audio database in various acoustic environments,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Multichannel audio database in various acoustic environments,

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:07.979644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:22:07.425128Z digest=sha256:b4da7609adab910f5e35f29775ec9491283bb2d3774c14745e59f0cdfcf793f6

Observation 6be34bc1-4554-44a6-bbce-60ee64db0d1c · outbound

This paper cites The fifth ’chime’ speech separation and recognition challenge: Dataset, task and baselines,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline The fifth ’chime’ speech separation and recognition challenge: Dataset, task and baselines,

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:07.966005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:22:07.429542Z digest=sha256:48599b6f240190997cd3fa9e90ed9005e5c07d7afa759caae4d9e01d7233ebe5

Observation b89ca30e-01c2-46bf-b953-d0904c8f36e3 · outbound

This paper cites SonicSim: A customizable simulation platform for speech processing in moving sound source scenarios.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline SonicSim: A customizable simulation platform for speech processing in moving sound source scenarios

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:07.433411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:07.433411Z digest=sha256:02ff2e5af0e3c4e5e5040bd64143b6c04d07d3da2aaae0ede66f9fd750743055

Observation 35abf0fb-854c-4bf8-ae44-0209e167cd0d · outbound

This paper cites Perceptual evaluation of speech quality (pesq)-a new method for speech quality assessment of telephone networks and codecs,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Perceptual evaluation of speech quality (pesq)-a new method for speech quality assessment of telephone networks and codecs,

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:07.951031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:22:07.437013Z digest=sha256:797bc03913cac6b762a3a89901ffe6e03d19a70e3967e5ba2ac871a8a3fb4b94

Observation e24fe973-e9ee-44e8-88a7-9561bd4b5938 · outbound

This paper cites An algorithm for predicting the intelligibility of speech masked by modulated noise maskers,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline An algorithm for predicting the intelligibility of speech masked by modulated noise maskers,

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:07.931359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:22:07.440926Z digest=sha256:ff27636434fab05709648d857033a053bfa38152a53e697a219919def2d76dfe

Observation 9ca04559-6a14-4eb9-95b6-5a04d202f30a · outbound

This paper cites Dnsmos P.835: A non- intrusive perceptual objective speech quality metric to evaluate noise suppressors,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Dnsmos P.835: A non- intrusive perceptual objective speech quality metric to evaluate noise suppressors,

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:07.911881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:22:07.445060Z digest=sha256:ff0ba976905a703d9cad718d8f6d74f33b66f26d2a9ce4865efa1a16cf5496a0

Observation a04353fe-156d-442a-ba94-d5b2718ed1be · outbound

This paper cites Robust speech recognition via large-scale weak supervision,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Robust speech recognition via large-scale weak supervision,

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:07.888087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:22:07.453503Z digest=sha256:cea129ec6d5b71915267fdd81c625749c96a01e502a3cca6fef5ea5320091b07

Observation 7f382b3e-dce3-42ac-8325-5ba18d26efb8 · outbound

This paper cites Wavlm: Large-scale self-supervised pre- training for full stack speech processing,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Wavlm: Large-scale self-supervised pre- training for full stack speech processing,

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:07.457443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:07.457443Z digest=sha256:51486c57a12c9f768ab80821fe16a046a40263977c40ad239d9dbc992ada5b6f

Observation e5c1acbc-6245-400f-8509-f8751ee5b120 · outbound

This paper cites Improving Generalization of Speech Separation in Real-World Scenarios: Strategies in Simulation, Optimization, and Evaluation.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Improving Generalization of Speech Separation in Real-World Scenarios: Strategies in Simulation, Optimization, and Evaluation

Reference 76

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:22:07.533270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:22:07.460985Z digest=sha256:2122ee899e751e0ad00e2b531d1a21f11cc37dc41c7a986847e7b471f099e64a

Observation 31812576-3bd2-4f26-836a-81dda307c00c · outbound

This paper cites Classifier-Free Diffusion Guidance.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Classifier-Free Diffusion Guidance

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:07.466008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:07.466008Z digest=sha256:cdc75896a7312c74cae186ef92e1eb1bf67dd9ef19cb455ccae52be45ad923e3

Observation 2da558ad-64b7-47c4-b910-1665ea66ae00 · outbound

This paper cites Photorealistic text-to-image diffusion models with deep language understanding,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Photorealistic text-to-image diffusion models with deep language understanding,

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:07.869297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:22:07.470504Z digest=sha256:46a022d252d18b4e9fa6f6799f7dbe62356903b3eec7414d5cb95b37dd36e325

Observation 6f599b9a-c3e8-4b82-9d3d-e8ebea2852e9 · outbound

This paper cites Available: https://openreview.net/forum?id=08Yk-n5l2Al.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Available: https://openreview.net/forum?id=08Yk-n5l2Al

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:07.857925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:22:07.475291Z digest=sha256:1047328b104d935374520ce3fca5a7aa058de70371f1d265d4a9196d44438d06

Observation f9ec7a85-4eca-4bf2-b7a6-cc964b864693 · outbound

This paper cites an unresolved cited work.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Unresolved cited work

Reference 2022

Resolution
parse uncertain
no resolver link, observed 2026-08-07T14:22:07.449434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:07.449434Z digest=sha256:b19675f18e909745f975997ac809299954c40b20b6a632dc22d3c41502553aaa

Pith citing papers

Observation cc329c18-7386-4acd-8846-95a8f1672c13 · inbound

GenTSE: Enhancing Target Speaker Extraction via a Coarse-to-Fine Generative Language Model cites this paper.

GenTSE: Enhancing Target Speaker Extraction via a Coarse-to-Fine Generative Language Model SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T14:20:49.132229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:20:49.132229Z digest=sha256:a2a78761bcd4ef83816e6f1118a35d549618ff2cd219a91ba5eb06dd39f8750c

Observation 0293e581-3ff6-4913-8d5f-4960ebfa64ed · inbound

Enroll-on-Wakeup: A First Comparative Study of Target Speech Extraction for Seamless Interaction in Real Noisy Human-Machine Dialogue Scenarios cites this paper.

Enroll-on-Wakeup: A First Comparative Study of Target Speech Extraction for Seamless Interaction in Real Noisy Human-Machine Dialogue Scenarios SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T22:51:20.849476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:51:20.849476Z digest=sha256:78170b4365fc09b8701640d06baf1f2802b7f46c2e36e29402c2086a4c66fd67

Observation 29ebe771-cc57-4139-8478-42c13570334e · inbound

Beyond Acoustic Prefixes: Persistent Grounding in Serialized Acoustic Memory for LLM-Based Multi-Talker Speech Recognition cites this paper.

Beyond Acoustic Prefixes: Persistent Grounding in Serialized Acoustic Memory for LLM-Based Multi-Talker Speech Recognition SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-13T17:09:15.281290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T17:09:15.281290Z digest=sha256:7823cd5abe5f57d5470b7dd6fd6b964ae92c45564a6afe947e69c4e52f1d8198