Pith. sign in

Paper Citation Record · LEDGER

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline

As of 8 August 2026, this Paper Citation Record lists 80 of 80 outbound references and 3 inbound Pith citation observations for arXiv:2505.19314.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.19314 v3

Coverage vector

measured 80 of 80 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:22:07.475291Z

measured 83 of 83 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T14:20:49.132229Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

80 of 80 outbound references displayed

  • verified exact7
  • verified fuzzy55
  • unresolved17
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d86250cd-c963-4ac1-9256-ddbe38fad0ee · outbound

This paper cites The cocktail-party problem revisited: early pro- cessing and selection of multi-talker speech,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline The cocktail-party problem revisited: early pro- cessing and selection of multi-talker speech,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:16.316621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:22:07.154451Z digest=sha256:cd11d2f400ab9eaa0037fbc1ce5319d96e22ce76ed5892f97b1f8671270aed98

Observation d24a1ce8-4e6c-4d93-9999-a8c2e4cf756a · outbound

This paper cites Neural target speech extraction: An overview,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Neural target speech extraction: An overview,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:16.167867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:22:07.158760Z digest=sha256:b40fec389698e8880cab3f8176c19a129c77444355aac7362923ec53ef0f6c89

Observation a5f4c12f-df00-4fbe-9813-b06d63e71704 · outbound

This paper cites Neural spatial filter: Target speaker speech separation assisted with directional information,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Neural spatial filter: Target speaker speech separation assisted with directional information,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:16.021303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:22:07.163056Z digest=sha256:ef174e8829bff1032970094b6f3d9dc2d4389cb569f9942537ca7b242ff99fd5

Observation 8b003eec-47d2-42bc-b634-cecfc738d80d · outbound

This paper cites Far-field location guided target speech extraction using end-to-end speech recognition objectives,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Far-field location guided target speech extraction using end-to-end speech recognition objectives,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:15.913298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:22:07.168216Z digest=sha256:dda94a09f1ccb69d66f8cbb5f37d43f0304ea64ea025935937e5cae2fb1a5a08

Observation 1eedf57c-0046-4639-a424-985e93779fbf · outbound

This paper cites Looking to listen at the cocktail party: a speaker-independent audio-visual model for speech separation,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Looking to listen at the cocktail party: a speaker-independent audio-visual model for speech separation,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:15.435237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:22:07.171996Z digest=sha256:bd2482bf54f1eeb61a3dea279dc76f5c7649ae85154a60908d0c6232a8112789

Observation 27460b18-8024-4c28-a89d-4479dbe0a0ef · outbound

This paper cites Conceptbeam: Concept driven target speech extraction,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Conceptbeam: Concept driven target speech extraction,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:15.283829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:22:07.176146Z digest=sha256:212d7c1500fd92c0dfb16f688ec05ec666d29872bf37e741290bd8d9f15853d3

Observation 693bb2b9-cf52-4e47-8d02-945d341d136c · outbound

This paper cites V oicefilter: Targeted voice separation by speaker-conditioned spectrogram masking,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline V oicefilter: Targeted voice separation by speaker-conditioned spectrogram masking,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:15.164815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:22:07.180844Z digest=sha256:14740606a53516dffbf28844722a3f9548bf53b9fd53ff0b6a454b04c7a478d6

Observation 930acadf-25cf-4460-b1bf-70371d4bf6ca · outbound

This paper cites Speakerbeam: Speaker aware neural network for target speaker extraction in speech mixtures,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Speakerbeam: Speaker aware neural network for target speaker extraction in speech mixtures,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:15.002077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:22:07.184966Z digest=sha256:3756ef3177d4d4f66f712e7b8b8f468ef817f53bac2c36aec4825d292fae7866

Observation 33604da3-1634-4163-ac55-ae8543afb007 · outbound

This paper cites Wesep: A scalable and flexible toolkit towards generalizable target speaker extraction,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Wesep: A scalable and flexible toolkit towards generalizable target speaker extraction,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:14.875258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:22:07.188682Z digest=sha256:e2571df4c3f7d996e47123c3c277f007e96a5e5b0343e8c6c844c79140e406d7

Observation 460ed315-7fe9-41ff-b469-ddcbc96234dd · outbound

This paper cites Target confusion in end-to-end speaker extraction: Analysis and approaches,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Target confusion in end-to-end speaker extraction: Analysis and approaches,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:14.758242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:22:07.192850Z digest=sha256:da652a1696c486a88510e78165d4a90efffe96d50f88d2303720d35b2614df15

Observation f22eb8bc-4f60-4f6c-803d-194c16e497aa · outbound

This paper cites Dpccn: Densely-connected pyramid complex convolutional network for robust speech separation 11 and extraction,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Dpccn: Densely-connected pyramid complex convolutional network for robust speech separation 11 and extraction,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:14.633307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:22:07.196434Z digest=sha256:cd2ec6d2c3c5eeb5bc240523221c5ee0b76e0d133d8a6f0e5d615922ce6cb46d

Observation a5f74896-a48a-44e2-a4d5-4a405cdc5bdc · outbound

This paper cites Improving target sound extraction with timestamp information,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Improving target sound extraction with timestamp information,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:14.516695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:22:07.200583Z digest=sha256:c44b7dbc1f76a71c62a0f8e931d15d5b76256beff10ba14cfc8def57c8a18e1b

Observation 610f6f0c-164e-427f-ad8d-55bce01922a0 · outbound

This paper cites WeSep: A Scalable and Flexible Toolkit Towards Generalizable Target Speaker Extraction.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline WeSep: A Scalable and Flexible Toolkit Towards Generalizable Target Speaker Extraction

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:07.204106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:07.204106Z digest=sha256:22e7d22ba13c488d3a046d04a59bee0c907ae4078333c9b7e8b1c6b157f34b9b

Observation 611ca561-de03-4463-8427-7df5c5ba0e9c · outbound

This paper cites Spex: Multi-scale time domain speaker extraction network,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Spex: Multi-scale time domain speaker extraction network,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:14.357625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:22:07.208667Z digest=sha256:bfd25cd8c98ff010be3e7d09d4ccf01951090c19c5afb6298304d577b865c071

Observation 3858436a-b8b9-43ec-9a1f-78e582e3d462 · outbound

This paper cites Spex+: A complete time domain speaker extraction network,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Spex+: A complete time domain speaker extraction network,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:14.233683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:22:07.212662Z digest=sha256:479e5812a23ebbe0dbfa2fa6a498e10e68120593ac8f4186a7132353ce3bd56a

Observation 85db493b-472e-4c86-895d-e4730dc4985c · outbound

This paper cites X-SEPFORMER: end-to- end speaker extraction network with explicit optimization on speaker confusion,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline X-SEPFORMER: end-to- end speaker extraction network with explicit optimization on speaker confusion,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:14.115190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:22:07.216422Z digest=sha256:c9c257c01448702569e5a03c2b671dfa75d6729abd5b0649096c3ac4ef8fd2ef

Observation 8da9cead-1b4d-4c0c-877a-787dddb08cab · outbound

This paper cites X-tf-gridnet: A time-frequency domain target speaker extraction network with adaptive speaker embedding fusion,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline X-tf-gridnet: A time-frequency domain target speaker extraction network with adaptive speaker embedding fusion,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:13.979985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:22:07.221377Z digest=sha256:f0609c86be202b30dcbbd91dc715a5e662a7320ec923313d3286c25977b5137f

Observation 1de9266a-258a-4493-b63b-e2782b33b22b · outbound

This paper cites USEF-TSE: Universal Speaker Embedding Free Target Speaker Extraction.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline USEF-TSE: Universal Speaker Embedding Free Target Speaker Extraction

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:22:07.834953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:22:07.225317Z digest=sha256:b18049c88de05f35a50701018fbc83af7b83290c42a3d60e31a10f1b4cfdccbe

Observation ed61ae57-6b30-475d-b20a-a71d244102f2 · outbound

This paper cites Target speech extraction with conditional diffusion model,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Target speech extraction with conditional diffusion model,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:13.866038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:22:07.230095Z digest=sha256:df974eb9e88bd2ac6a2091766d6cb57a7b2db5eadcf9d3936da401bc1198c60e

Observation 3f2ee616-3578-4865-9a29-f62fd92bdbb1 · outbound

This paper cites Noise-robust Speech Separation with Fast Generative Correction.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Noise-robust Speech Separation with Fast Generative Correction

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:22:07.817493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:22:07.234297Z digest=sha256:407b84564d83a7af7f4f339a108677a72214e01fc69057c40f10820212f01e7b

Observation 9f531854-3c29-43d1-bf74-867007a691f8 · outbound

This paper cites Speech enhancement and dereverberation with diffusion-based generative models,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Speech enhancement and dereverberation with diffusion-based generative models,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:13.683036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:22:07.239647Z digest=sha256:be75dbaf9d1dc429cf60283d5cc812fc1c51ef76fecb12421972d18096526b55

Observation 1c4af11a-89ad-4505-9ea5-a24441951168 · outbound

This paper cites Diffusion-based generative speech source separation,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Diffusion-based generative speech source separation,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:07.243716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:07.243716Z digest=sha256:f9f2fa611abe37f2a6dbccddbc47cda6b1387a836ff9f1cc81bede6f3755c418

Observation 63742595-d3a9-4838-bd75-3a4f27419613 · outbound

This paper cites Generative pre-training for speech with flow matching,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Generative pre-training for speech with flow matching,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:13.475291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:22:07.248153Z digest=sha256:e36d269d44a56a332e9a38eb9a6c2af90ae58b88e2077c9e2e95d7f7ecc8c608

Observation ffb056b2-5e5a-4357-a714-17287e0219e7 · outbound

This paper cites Metis: A Foundation Speech Generation Model with Masked Generative Pre-training.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Metis: A Foundation Speech Generation Model with Masked Generative Pre-training

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:07.252044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:07.252044Z digest=sha256:a66b8e266b5eb81a350e5115ef5582d03100a295248cb3431d097d50bdce5acd

Observation af0f0304-7d6e-4afa-91ce-5577d88b2cbd · outbound

This paper cites SoloAudio: Target Sound Extraction with Language-oriented Audio Diffusion Transformer.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline SoloAudio: Target Sound Extraction with Language-oriented Audio Diffusion Transformer

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:07.256774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:07.256774Z digest=sha256:0773a4e863fd5676b85195beb00a93515866075eaf83ac7ebf68d37f6c9df146

Observation bf62dc7e-f961-46a4-be01-5f8568be862f · outbound

This paper cites Separate And Diffuse: Using a Pretrained Diffusion Model for Improving Source Separation.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Separate And Diffuse: Using a Pretrained Diffusion Model for Improving Source Separation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:07.261084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:07.261084Z digest=sha256:0b5bab75551962cdaf458e30eea5d32ddd960460fc521e1336d10310811bf2d1

Observation eadbdd15-2a50-446f-88e2-1527b08990b5 · outbound

This paper cites Attention is all you need,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Attention is all you need,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:07.265276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:07.265276Z digest=sha256:fa1d0deae6d7701a562f227f60909a562000de12ee1eac2ee3662eaf79a4c14f

Observation 17e636f9-cabc-4df8-b41f-2cbca8dbe101 · outbound

This paper cites Large language model based generative error correction: A challenge and baselines for speech recognition, speaker tagging, and emotion recognition,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Large language model based generative error correction: A challenge and baselines for speech recognition, speaker tagging, and emotion recognition,

Reference 28

Resolution
verified exact
raw_fallback, observed 2026-08-07T14:22:07.780630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:22:07.269521Z digest=sha256:5b8cb7a44690466c124874bfef3c3f794a918c1eac44523959b60d613a289573

Observation 58f778a2-4cd6-4848-b207-4291a63e5542 · outbound

This paper cites SepALM: Audio Language Models Are Error Correctors for Robust Speech Separation.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline SepALM: Audio Language Models Are Error Correctors for Robust Speech Separation

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:22:07.710817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:22:07.274041Z digest=sha256:8f80138ddbc04aa24db5bf20164196bbd73ef94fdd367b4a04ceeb1db736fa6f

Observation ff992483-89bd-4c28-8710-1079063a1ac4 · outbound

This paper cites LibriMix: An Open-Source Dataset for Generalizable Speech Separation.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline LibriMix: An Open-Source Dataset for Generalizable Speech Separation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:07.278378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:07.278378Z digest=sha256:e7a91de718c1b65a9c23f0adac00fdf352bc953125ab22a45a6d79bcb9bc8505

Observation 335f9c72-14f9-4d0b-9d3d-7ba3e47b1cba · outbound

This paper cites Target speech extraction with conditional diffusion model,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Target speech extraction with conditional diffusion model,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:13.348423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:22:07.282398Z digest=sha256:468e6ecec5fa84d30081729911b0bb78d19886b37009d59778bc5509a576025a

Observation e99c4cc7-f709-4a9a-a5f6-b9cb4bf63e49 · outbound

This paper cites Dpm-tse: A diffusion probabilistic model for target sound extraction,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Dpm-tse: A diffusion probabilistic model for target sound extraction,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:13.098187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:22:07.286482Z digest=sha256:743825e5f513f50dd89799288dd28a4865f0a64a2ada04d0507ac5c0de0c161e

Observation 73ea1afc-061b-4f60-a9f9-019697bca4fb · outbound

This paper cites Diffusion- based generative speech source separation,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Diffusion- based generative speech source separation,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:12.750074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:22:07.290826Z digest=sha256:66c12c014bfc9bfb1433c80a34957ef5fb95787076a9cf8807f8027cd3642063

Observation 246e030c-ca12-4386-8f20-fdd375bd3a4e · outbound

This paper cites Generative Speech Foundation Model Pretraining for High-Quality Speech Extraction and Restoration.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Generative Speech Foundation Model Pretraining for High-Quality Speech Extraction and Restoration

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:22:07.683925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:22:07.295184Z digest=sha256:bf63c19281fe1969087d78fbd8c877fc1b29154860458436df5bda7c96a44c60

Observation aa515a27-f2d8-4c0c-99c9-7d5b713dd5b7 · outbound

This paper cites Generation- based target speech extraction with speech discretization and vocoder,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Generation- based target speech extraction with speech discretization and vocoder,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:12.458042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:22:07.299420Z digest=sha256:32f559f77fa3861ee63547137ab563842f424f1567ef2b1ad066840be52f2e94

Observation 25dbd4bf-1922-457e-abd1-dd3165775919 · outbound

This paper cites Enhancing Intelligibility for Generative Target Speech Extraction via Joint Optimization with Target Speaker ASR.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Enhancing Intelligibility for Generative Target Speech Extraction via Joint Optimization with Target Speaker ASR

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:22:07.667941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:22:07.303374Z digest=sha256:4f3c1262553251c6961a18269eab92cf328022999a4c3472fb12da48b4a7a7eb

Observation 72098b94-b6c8-4c80-a0a5-18864715c414 · outbound

This paper cites Diffusion-based signal refiner for speech separation,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Diffusion-based signal refiner for speech separation,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:07.307054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:07.307054Z digest=sha256:0c1754ac91203f30afaef09daeb33f80b0daae2b9fc3ad797bcfe200037be0fc

Observation 4842be6c-412c-46b7-83ea-3ef0264395a8 · outbound

This paper cites Storm: A diffusion-based stochastic regeneration model for speech enhancement and dereverberation,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Storm: A diffusion-based stochastic regeneration model for speech enhancement and dereverberation,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:12.042351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:22:07.311336Z digest=sha256:6781ca255d6753607178d3577dc2fb1dabfd32300ffb6c2a9fd28a1616462f78

Observation 22f71e89-d785-415d-b410-649fa43d864b · outbound

This paper cites Ddtse: Discriminative diffusion model for target speech extraction,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Ddtse: Discriminative diffusion model for target speech extraction,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:11.826875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:22:07.315289Z digest=sha256:f655e878c10371a831d6965c40eb3e8dd762e44d70de364144f888958060844e

Observation 4c8713c5-9933-4c50-ba86-8f2066bbe735 · outbound

This paper cites Speaker-aware neural network based beamformer for speaker extraction in speech mixtures,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Speaker-aware neural network based beamformer for speaker extraction in speech mixtures,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:11.286750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:22:07.319275Z digest=sha256:dfd83edd82d4a4d6772c59e4d27604abfa0f6bc8f5726b2c3df067f07d9ef5d1

Observation 52a51c3a-7be6-4adb-80a6-4269e55219eb · outbound

This paper cites X- vectors: Robust DNN embeddings for speaker recognition,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline X- vectors: Robust DNN embeddings for speaker recognition,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:10.825851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:22:07.323186Z digest=sha256:097fdc97317c269ee5be253d0760c337289dab813f30f687ee3521bfb5fa1e19

Observation 4f347263-8123-49eb-a228-bd5c259de03c · outbound

This paper cites Probing self-supervised learning models with target speech extraction,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Probing self-supervised learning models with target speech extraction,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:10.383416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:22:07.326526Z digest=sha256:b6d6afaaf4d3f9ab5dab1ae8f99df16ef3a2d24f6eac9999863e756d6b222a2a

Observation 70347c62-4802-4526-b4b1-884eaeda3a8c · outbound

This paper cites Target speech extraction with pre-trained self-supervised learning models,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Target speech extraction with pre-trained self-supervised learning models,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:10.184130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:22:07.329756Z digest=sha256:75bae8b0852a1ff5dc880e36b05a075bb33864e923a92f0b8490ccf8d6a06a7a

Observation a6131e0a-85aa-4405-8e71-6099872bfdb6 · outbound

This paper cites Smma-net: An audio clue-based target speaker extraction network with spectrogram matching and mutual attention,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Smma-net: An audio clue-based target speaker extraction network with spectrogram matching and mutual attention,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:10.039313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:22:07.333165Z digest=sha256:b9da1e9804008a62b32ff68f8038f369bcf3f0c73deb67dfdcf9849565047b98

Observation dc6868b9-d78f-42f0-848f-bb5f43732e07 · outbound

This paper cites Target speaker extraction by directly exploiting contextual information in the time-frequency domain,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Target speaker extraction by directly exploiting contextual information in the time-frequency domain,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:09.902393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:22:07.336782Z digest=sha256:62b7a29b1a76f75d6e0707e15aaf2249bc85644d504f507f1fa033383adedc04

Observation f0f6c734-18fe-4623-9839-4327c25098b4 · outbound

This paper cites Target speaker extraction with ultra-short reference speech by VE-VE framework,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Target speaker extraction with ultra-short reference speech by VE-VE framework,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:09.751317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:22:07.340177Z digest=sha256:e7a3bd485f2defb540dc66097f2bff7f0adf79cb7ed8be4d31e7a85196169ff8

Observation cfcbd0d9-aa6b-4853-b2a3-7dd89ec2c885 · outbound

This paper cites Sef-net: Speaker embedding free target speaker extraction network,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Sef-net: Speaker embedding free target speaker extraction network,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:09.585458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:22:07.344503Z digest=sha256:25467bcb19d576e368f2565d1853eb77d633b65af26e0987c7679317d44792b4

Observation 6dc8bc8b-9d24-4d81-9c76-b40ddc55d071 · outbound

This paper cites Common diffusion noise schedules and sample steps are flawed,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Common diffusion noise schedules and sample steps are flawed,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:09.460733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:22:07.347592Z digest=sha256:6d2dcda70097334c13f08a7719222c3b65defa3e9d3cbd297ea4e4fdee35fac0

Observation 23f2da4c-8b3e-435f-ad39-abb56ef48d70 · outbound

This paper cites Progressive distillation for fast sampling of diffusion models,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Progressive distillation for fast sampling of diffusion models,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:09.360479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:22:07.350993Z digest=sha256:776ae42e718d4ff31f9410802a111271226d3442f27908c02ca8d4e0b1b53989

Observation f157c4e6-1f2e-4093-a974-7438ab5c6e9d · outbound

This paper cites High- fidelity audio compression with improved RVQGAN,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline High- fidelity audio compression with improved RVQGAN,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:09.211137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:22:07.354792Z digest=sha256:e8517f7bf32ccb5e2225c455f186926ff0f2abc611fa20aee397531d75ef9f5e

Observation 010a6aa9-89f3-47cb-9fbb-8b2e41a8ab8f · outbound

This paper cites Stable Audio Open.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Stable Audio Open

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:07.358306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:07.358306Z digest=sha256:e8ce5088a6f4dcf2c5294233322c1c1094afac86720777cd86e4cd37ed52a3ae

Observation 96e4134f-1562-4fc0-96fe-b16912434989 · outbound

This paper cites EzAudio: Enhancing Text-to-Audio Generation with Efficient Diffusion Transformer.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline EzAudio: Enhancing Text-to-Audio Generation with Efficient Diffusion Transformer

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:07.362796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:07.362796Z digest=sha256:13c7d13cdf59a0f28f9665a814f0fda6db943c5c3865681c3fcc2def9bd0f6b5

Observation 15bc5cc7-3c30-4004-a002-2d6a7760ec0b · outbound

This paper cites Tf- gridnet: Integrating full- and sub-band modeling for speech separation,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Tf- gridnet: Integrating full- and sub-band modeling for speech separation,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:09.018402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:22:07.366834Z digest=sha256:8f8af4c93785b0b0fb044f9d763d9902a51760ee239c0e1bac47133d6267e9c8

Observation e6a27d0f-83b0-47a4-b831-3b0e46ed4003 · outbound

This paper cites SPMamba: State-space model is all you need in speech separation.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline SPMamba: State-space model is all you need in speech separation

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:07.370792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:07.370792Z digest=sha256:290d3e04f006647206967c8cbd5503d7afc5e8c06f8a6218dc0c6814290f5342

Observation 11054119-9270-47de-9613-91804548ef4b · outbound

This paper cites Complex ratio masking for monaural speech separation,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Complex ratio masking for monaural speech separation,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:08.871580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:22:07.374906Z digest=sha256:eec94e44b31c2f3e85e3553030d44f519eee78bd6c396c342383cc77e2135872

Observation aadc30e0-17e9-4e6e-ba89-151dfc7f0c01 · outbound

This paper cites auraloss: Audio focused loss functions in pytorch,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline auraloss: Audio focused loss functions in pytorch,

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:07.379255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:07.379255Z digest=sha256:ade8c5fd8298d683592e81bd877b6cd9116a2f517653f651175e1bc881f4ad1e

Observation 54a7bc8f-5781-4c69-b7e1-ad5fe4ce80f4 · outbound

This paper cites High fidelity neural audio compression,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline High fidelity neural audio compression,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:08.674926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:22:07.383128Z digest=sha256:a4746a50f03bbd2283ba0d54c6cc0c54f97d8f885e6358f24305af99bad7978b

Observation de2e9680-5b1a-41a2-9c2f-441ef16b5157 · outbound

This paper cites Scalable diffusion models with transformers,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Scalable diffusion models with transformers,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:08.454405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:22:07.387236Z digest=sha256:a6c6748c772bd9891fcb0e8adaf14a5b288df02967a7de5820880ad67bdfbadd

Observation 31cfd854-1be9-4e3c-804e-60f026a25267 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline An image is worth 16x16 words: Transformers for image recognition at scale,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:08.264101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:22:07.390920Z digest=sha256:d5752862c0b4953185857e77f0e0d82c15f9755c99573dc049c0346f7e12c07c

Observation a0bfd558-bae4-425f-91e7-49c5aed8eebf · outbound

This paper cites Sigmoid-weighted linear units for neural network function approximation in reinforcement learning,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Sigmoid-weighted linear units for neural network function approximation in reinforcement learning,

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:07.395030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:07.395030Z digest=sha256:ef09a8a3126b8d6c69698e08c9c11016f068483e0776557064997d22cc562b53

Observation 2f748628-177d-4a38-b1bd-f2f82efc59c0 · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Roformer: Enhanced transformer with rotary position embedding,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:08.143124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:22:07.399020Z digest=sha256:d2a0f601f754591e8bbd01127ba208991d57da4d15c33e08ab1a94c050ba9b4d

Observation 0edd31f9-8f3b-4bb4-9518-44c1cbd8072d · outbound

This paper cites Single- channel multi-speaker separation using deep clustering,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Single- channel multi-speaker separation using deep clustering,

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:08.040444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:22:07.402914Z digest=sha256:07ce5ae3085db9aea4932825aa3bdbe7207c15634f720f486129f2eced36a98d

Observation 75bdcf37-727d-4132-8a8d-0786387e5952 · outbound

This paper cites Conv-tasnet: Surpassing ideal time-frequency magnitude masking for speech separation,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Conv-tasnet: Surpassing ideal time-frequency magnitude masking for speech separation,

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:08.028706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:22:07.406508Z digest=sha256:43f91f9012a41066d1248ab1e37beb98f0b79bb59c55c36ffd0645ec7d28fe7c

Observation 4918245f-567b-4278-be56-986fac214ce5 · outbound

This paper cites Wham!: Extending speech separation to noisy environments,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Wham!: Extending speech separation to noisy environments,

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:08.017035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:22:07.410086Z digest=sha256:4e394a1854135fae0575aef01a76264293d9602714038eab9f921197026c7cf7

Observation 79627c97-0cab-4265-9437-c08bde14e8bb · outbound

This paper cites Librispeech: An ASR corpus based on public domain audio books,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Librispeech: An ASR corpus based on public domain audio books,

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:08.005422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:22:07.413724Z digest=sha256:a09a79abfb550e5e740bd46e0f71ccaa79e0cb22ec048821ffa69ae0e9d21caa

Observation 73d7aa18-fb73-4828-97c6-8dd42e67a898 · outbound

This paper cites Improving speaker discrimination of target speech extraction with time-domain speakerbeam,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Improving speaker discrimination of target speech extraction with time-domain speakerbeam,

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:07.993118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:22:07.417535Z digest=sha256:be976cd858dfd2f642cd74f2093402b804595aa6b28cccec92bdeb06a1055544

Observation 7deb9100-49d1-4a44-81eb-7d1dac4d1144 · outbound

This paper cites MUSAN: A Music, Speech, and Noise Corpus.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline MUSAN: A Music, Speech, and Noise Corpus

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:07.421318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:07.421318Z digest=sha256:c37fb3ba3b99d2b4dc38556531c9a263dbb3dbf91e1185d3701efb32b15cc677

Observation d5dd377b-9cfe-4767-b06e-9e8dc831bfbf · outbound

This paper cites Multichannel audio database in various acoustic environments,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Multichannel audio database in various acoustic environments,

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:07.979644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:22:07.425128Z digest=sha256:ef03fb15956f5bb8a943ff12774515edbb5be80621774cc8792548a2b67d8067

Observation 6be34bc1-4554-44a6-bbce-60ee64db0d1c · outbound

This paper cites The fifth ’chime’ speech separation and recognition challenge: Dataset, task and baselines,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline The fifth ’chime’ speech separation and recognition challenge: Dataset, task and baselines,

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:07.966005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:22:07.429542Z digest=sha256:d9e59ffcc4787b38c1daceac77dbe625b0df144ce3f749782e8f46068cf2a3f2

Observation b89ca30e-01c2-46bf-b953-d0904c8f36e3 · outbound

This paper cites SonicSim: A customizable simulation platform for speech processing in moving sound source scenarios.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline SonicSim: A customizable simulation platform for speech processing in moving sound source scenarios

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:07.433411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:07.433411Z digest=sha256:fd7e1bb488a15bfc11afc5e361911eace368a766aa5010ad7a84100d64be1d33

Observation 35abf0fb-854c-4bf8-ae44-0209e167cd0d · outbound

This paper cites Perceptual evaluation of speech quality (pesq)-a new method for speech quality assessment of telephone networks and codecs,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Perceptual evaluation of speech quality (pesq)-a new method for speech quality assessment of telephone networks and codecs,

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:07.951031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:22:07.437013Z digest=sha256:72c1450ce0ea9d038a8dd087b1d931d4d3e0fe606abfb32f231525b3b867dffe

Observation e24fe973-e9ee-44e8-88a7-9561bd4b5938 · outbound

This paper cites An algorithm for predicting the intelligibility of speech masked by modulated noise maskers,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline An algorithm for predicting the intelligibility of speech masked by modulated noise maskers,

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:07.931359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:22:07.440926Z digest=sha256:5b3bc12f6b298267d48c3ef28d1380c9da6c36bcdec2f538ddca33d3a28bac97

Observation 9ca04559-6a14-4eb9-95b6-5a04d202f30a · outbound

This paper cites Dnsmos P.835: A non- intrusive perceptual objective speech quality metric to evaluate noise suppressors,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Dnsmos P.835: A non- intrusive perceptual objective speech quality metric to evaluate noise suppressors,

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:07.911881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:22:07.445060Z digest=sha256:ca49c34c53cf42f4269a9ccd3b1fbb620814882ebb9bc29bf46131362abdb4ca

Observation a04353fe-156d-442a-ba94-d5b2718ed1be · outbound

This paper cites Robust speech recognition via large-scale weak supervision,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Robust speech recognition via large-scale weak supervision,

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:07.888087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:22:07.453503Z digest=sha256:071050fd18b362c933782b144d70feb9662ba984536739aa0850b4c65c987947

Observation 7f382b3e-dce3-42ac-8325-5ba18d26efb8 · outbound

This paper cites Wavlm: Large-scale self-supervised pre- training for full stack speech processing,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Wavlm: Large-scale self-supervised pre- training for full stack speech processing,

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:07.457443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:07.457443Z digest=sha256:ee0741cdaab87827ff95e1adee168ce863112954f404b0a4ebf7cfc564040f3a

Observation e5c1acbc-6245-400f-8509-f8751ee5b120 · outbound

This paper cites Improving Generalization of Speech Separation in Real-World Scenarios: Strategies in Simulation, Optimization, and Evaluation.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Improving Generalization of Speech Separation in Real-World Scenarios: Strategies in Simulation, Optimization, and Evaluation

Reference 76

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:22:07.533270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:22:07.460985Z digest=sha256:fd41efb1cdf9a12bdcc5d77c340666ca66a07b72ecd4101622ce85c86a710fff

Observation 31812576-3bd2-4f26-836a-81dda307c00c · outbound

This paper cites Classifier-Free Diffusion Guidance.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Classifier-Free Diffusion Guidance

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:07.466008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:07.466008Z digest=sha256:daaa1b9a79c1eb17a187a940dc81699c49a724b1884244fc67567a2126ecc715

Observation 2da558ad-64b7-47c4-b910-1665ea66ae00 · outbound

This paper cites Photorealistic text-to-image diffusion models with deep language understanding,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Photorealistic text-to-image diffusion models with deep language understanding,

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:07.869297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:22:07.470504Z digest=sha256:51f0f84a259785e2b0836d1c11a8e52e5da0d17b0d0d3d1dcff5864597a4200e

Observation 6f599b9a-c3e8-4b82-9d3d-e8ebea2852e9 · outbound

This paper cites Available: https://openreview.net/forum?id=08Yk-n5l2Al.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Available: https://openreview.net/forum?id=08Yk-n5l2Al

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:07.857925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:22:07.475291Z digest=sha256:1dcc242168b6ce37084cac270326a2aa387c2f321c1fe6becaf5141be84feb10

Observation f9ec7a85-4eca-4bf2-b7a6-cc964b864693 · outbound

This paper cites an unresolved cited work.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Unresolved cited work

Reference 2022

Resolution
parse uncertain
no resolver link, observed 2026-08-07T14:22:07.449434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:07.449434Z digest=sha256:5d8b77deb366749987c9f3d945e4b5f2a02c4531c47ee5632072300bb942dd27

Pith citing papers

Observation cc329c18-7386-4acd-8846-95a8f1672c13 · inbound

GenTSE: Enhancing Target Speaker Extraction via a Coarse-to-Fine Generative Language Model cites this paper.

GenTSE: Enhancing Target Speaker Extraction via a Coarse-to-Fine Generative Language Model SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T14:20:49.132229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:20:49.132229Z digest=sha256:a891d7eefd5ac6c09d37da3de2ffcd2687b8eb2507e5f287ce5dc97ab65493f0

Observation 0293e581-3ff6-4913-8d5f-4960ebfa64ed · inbound

Enroll-on-Wakeup: A First Comparative Study of Target Speech Extraction for Seamless Interaction in Real Noisy Human-Machine Dialogue Scenarios cites this paper.

Enroll-on-Wakeup: A First Comparative Study of Target Speech Extraction for Seamless Interaction in Real Noisy Human-Machine Dialogue Scenarios SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T22:51:20.849476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:51:20.849476Z digest=sha256:7b3c6b4f399981162f871d03e9b0e3dc9a19a6d24809afe148022ddd8ee5a8cb

Observation 29ebe771-cc57-4139-8478-42c13570334e · inbound

Beyond Acoustic Prefixes: Persistent Grounding in Serialized Acoustic Memory for LLM-Based Multi-Talker Speech Recognition cites this paper.

Beyond Acoustic Prefixes: Persistent Grounding in Serialized Acoustic Memory for LLM-Based Multi-Talker Speech Recognition SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-13T17:09:15.281290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T17:09:15.281290Z digest=sha256:4eb0dc9b2bbc0bf1fbf299db4b461883ec60047a8d4eec693393a0651772ba75