Pith. sign in

Paper Citation Record · LEDGER

Beyond Speaker Identity: Text Guided Target Speech Extraction

As of 12 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2501.09169.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.09169 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T20:14:00.754527Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact1
  • verified fuzzy18
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5effb835-1de0-438e-89e6-e1e2f390e16f · outbound

This paper cites V oiceFilter: Targeted V oice Separation by Speaker-Conditioned Spectrogram Masking,.

Beyond Speaker Identity: Text Guided Target Speech Extraction V oiceFilter: Targeted V oice Separation by Speaker-Conditioned Spectrogram Masking,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:01.355088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:14:00.599522Z digest=sha256:6eaef0d9d53adc0558960491eeccaad46ad5efd2281c8c5646321ae35ffb2b83

Observation 943e1499-b812-41fa-bb79-310e48e11aae · outbound

This paper cites Speakerbeam: Speaker aware neural network for target speaker extraction in speech mixtures,.

Beyond Speaker Identity: Text Guided Target Speech Extraction Speakerbeam: Speaker aware neural network for target speaker extraction in speech mixtures,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:00.605275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:14:00.605275Z digest=sha256:34b054d9a611c6b38a5aaf85d993620aa9ff2ad680d2df3bf28167becfda2e9a

Observation 7b4987b4-d832-4ac2-aae7-a985417ee1c3 · outbound

This paper cites Spex+: A complete time domain speaker extraction network,.

Beyond Speaker Identity: Text Guided Target Speech Extraction Spex+: A complete time domain speaker extraction network,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:01.326227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:14:00.610679Z digest=sha256:d7574ed89a56a5a70c74297faa1b352228b7fbc0a03f3749cd811a4e337df1c1

Observation f0fbe5aa-2ab5-461b-9476-76768b5506bd · outbound

This paper cites Looking to listen at the cocktail party: A speaker-independent audio-visual model for speech separation,.

Beyond Speaker Identity: Text Guided Target Speech Extraction Looking to listen at the cocktail party: A speaker-independent audio-visual model for speech separation,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:01.308630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:14:00.615829Z digest=sha256:1aa88ca493858a8f3cdc7ed9714ab86dc8cca2069179462d2f4e39c8b31bf4c5

Observation 76c6c715-309f-4dc5-9c9f-3b9eb7f7f082 · outbound

This paper cites Multimodal SpeakerBeam: Single channel target speech extraction with audio-visual speaker clues.

Beyond Speaker Identity: Text Guided Target Speech Extraction Multimodal SpeakerBeam: Single channel target speech extraction with audio-visual speaker clues

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:01.269058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:14:00.621232Z digest=sha256:0ea5ece9e4807769d356eacd11af63b1883b3f017ed8ac8a43678266aaf42c58

Observation 0d3fb154-6984-4e9f-8660-e71b78b8e072 · outbound

This paper cites Text-driven separation of arbitrary sounds,.

Beyond Speaker Identity: Text Guided Target Speech Extraction Text-driven separation of arbitrary sounds,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:01.241447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:14:00.625979Z digest=sha256:9af9f79017bf246186d1023b54bb82050649b77c1cd6048d9052d17f808846cd

Observation d5f7358f-6ef5-4371-a0e1-12e6f3660e6b · outbound

This paper cites CLIPSep: Learning text-queried sound separation with noisy unlabeled videos,.

Beyond Speaker Identity: Text Guided Target Speech Extraction CLIPSep: Learning text-queried sound separation with noisy unlabeled videos,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:01.213502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:14:00.631961Z digest=sha256:77a717dc92fae9ec87cdc2edbf9dacd8f5b0e0e51d53b394dc1197a3f74497fd

Observation 790022a0-24b9-470c-a242-b5139b912fd3 · outbound

This paper cites Separate Anything You Describe.

Beyond Speaker Identity: Text Guided Target Speech Extraction Separate Anything You Describe

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:00.636550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:14:00.636550Z digest=sha256:ce0d53739474007674f376e9a365775ede29d3cf20c692b014244288f07d49f5

Observation d799943c-7b47-4695-a756-f868eecb56de · outbound

This paper cites Target sound extraction with variable cross-modality clues,.

Beyond Speaker Identity: Text Guided Target Speech Extraction Target sound extraction with variable cross-modality clues,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:01.193632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:14:00.641820Z digest=sha256:685b85cd930030a9523f5323a8095ead2928401d19019c9b3d0990fea5313b36

Observation c8714b0d-7845-4bf9-979a-8fafcb4697ea · outbound

This paper cites CLAPSep: Leveraging Contrastive Pre-trained Model for Multi-Modal Query-Conditioned Target Sound Extraction.

Beyond Speaker Identity: Text Guided Target Speech Extraction CLAPSep: Leveraging Contrastive Pre-trained Model for Multi-Modal Query-Conditioned Target Sound Extraction

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:00.646171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:14:00.646171Z digest=sha256:ed9f9a2af76b3c4cc0719a8e81458c77c3d7db53243be7fbfb1ebb6a62519679

Observation 4577da83-0503-427c-b818-f8e6f8c9f220 · outbound

This paper cites Typing to Listen at the Cocktail Party: Text-Guided Target Speaker Extraction.

Beyond Speaker Identity: Text Guided Target Speech Extraction Typing to Listen at the Cocktail Party: Text-Guided Target Speaker Extraction

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:00.651546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:14:00.651546Z digest=sha256:b024e82b50c88325e8a58117bb6cebf8d41b007a69235d3219b6683ef1d50fb8

Observation 529930c3-ceed-44b9-968d-e6124fc551b3 · outbound

This paper cites Target Speech Diarization with Multimodal Prompts.

Beyond Speaker Identity: Text Guided Target Speech Extraction Target Speech Diarization with Multimodal Prompts

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-10T20:14:00.859643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:14:00.657124Z digest=sha256:9a7361cabbef9ba134634c47d3e6e08c6903a3f231e4818a09e693b58182c222

Observation 41becd24-0ace-429c-ba4f-8606c5eb6231 · outbound

This paper cites Attention is all you need in speech separation,.

Beyond Speaker Identity: Text Guided Target Speech Extraction Attention is all you need in speech separation,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:01.165097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:14:00.662412Z digest=sha256:d5bf12b93fdad3b05deed9f38aa2e20011a9c57ae57a7af54185be86c4aa7856

Observation bdb63dc3-bd02-477b-beed-6ff8d0783781 · outbound

This paper cites SpeechBrain: A General-Purpose Speech Toolkit.

Beyond Speaker Identity: Text Guided Target Speech Extraction SpeechBrain: A General-Purpose Speech Toolkit

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:00.667684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:14:00.667684Z digest=sha256:874c259499461393f41dbaaad59012b04ac4ade458479e53ef4dd8b9cf9c0359

Observation f3eed92a-141a-444f-925f-a158c500a3a3 · outbound

This paper cites Textrolspeech: A text style control speech corpus with codec language text-to-speech models,.

Beyond Speaker Identity: Text Guided Target Speech Extraction Textrolspeech: A text style control speech corpus with codec language text-to-speech models,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:01.147416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:14:00.675548Z digest=sha256:79c0b17ae6f6c0bd016cf048ec13886781f1f34bfe1c742eea6bbbd3ed8077ee

Observation 17418e5a-10de-456d-8c14-58a147cb9a53 · outbound

This paper cites Optimization of speaker extraction neural network with magnitude and temporal spectrum approximation loss,.

Beyond Speaker Identity: Text Guided Target Speech Extraction Optimization of speaker extraction neural network with magnitude and temporal spectrum approximation loss,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:01.124753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:14:00.681693Z digest=sha256:4507b5f70771cf2ffc08b450396d7024c3046518798271b2d42dafac5db2af3d

Observation 2335e1c5-3911-4214-a303-5b7847321c5f · outbound

This paper cites LibriMix: An Open-Source Dataset for Generalizable Speech Separation.

Beyond Speaker Identity: Text Guided Target Speech Extraction LibriMix: An Open-Source Dataset for Generalizable Speech Separation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:00.691235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:14:00.691235Z digest=sha256:4dca9a447e006e261abf216d14cbc01f9adfa07cd428a061870af7824a5fe554

Observation 80c66340-71a3-460d-93f6-0d24113b728a · outbound

This paper cites Dual-path rnn: efficient long sequence modeling for time-domain single-channel speech sepa- ration,.

Beyond Speaker Identity: Text Guided Target Speech Extraction Dual-path rnn: efficient long sequence modeling for time-domain single-channel speech sepa- ration,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:01.103312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:14:00.703745Z digest=sha256:f0d19447ff2715fcbb411fbb31e56dc8810290b54dad88befbaafa67cf8608ef

Observation 4bb5b753-248c-4144-a981-612c9e84d0d8 · outbound

This paper cites Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups,.

Beyond Speaker Identity: Text Guided Target Speech Extraction Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:01.060804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:14:00.709583Z digest=sha256:9c4da2c59b762ee3fbc5b1f13d61518774fffd33fe8da694c2336ee678d3bfea

Observation 540a7d0f-2926-478c-821e-9abe2cbc5e3f · outbound

This paper cites X-vectors: Robust dnn embeddings for speaker recognition,.

Beyond Speaker Identity: Text Guided Target Speech Extraction X-vectors: Robust dnn embeddings for speaker recognition,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:00.715163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:14:00.715163Z digest=sha256:c532886757eb5bb07e9d8247d5dab6ac62e5b236ccf7e9290ecab52d7ff9e64f

Observation 06270b51-d33d-43fc-8a96-84f584383de9 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Beyond Speaker Identity: Text Guided Target Speech Extraction BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:00.721189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:14:00.721189Z digest=sha256:536d6cec832553dbcf2be071b9173edb805bba6c8c8f746253b49ac657dd0dd1

Observation 8491beed-4259-4cc3-a58f-5bb734d19bae · outbound

This paper cites Semi- supervised time domain target speaker extraction with attention,.

Beyond Speaker Identity: Text Guided Target Speech Extraction Semi- supervised time domain target speaker extraction with attention,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:01.032498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:14:00.726748Z digest=sha256:584d9f840fee2f53d74acc8e3d30cf7e80288d640e48e3702241e86835120646

Observation bb48f634-17e7-43d2-ae8f-024f92d5d5b3 · outbound

This paper cites Performance measure- ment in blind audio source separation,.

Beyond Speaker Identity: Text Guided Target Speech Extraction Performance measure- ment in blind audio source separation,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:01.011895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:14:00.731405Z digest=sha256:dbff4e6a80f7d232bb6e9af86e58b5364545cb88bb68aa7662667cbc2dbf07c7

Observation 1ef555f1-f253-4f8b-a894-1524efce4221 · outbound

This paper cites Wavesplit: End-to-end speech sep- aration by speaker clustering,.

Beyond Speaker Identity: Text Guided Target Speech Extraction Wavesplit: End-to-end speech sep- aration by speaker clustering,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:00.985082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:14:00.737069Z digest=sha256:5f56420484b0768abdb3be5aa6906e596412514ec6cfccf99f53a4c3643740b5

Observation 44f33c04-f025-45f8-bbd4-e5dbf6544a93 · outbound

This paper cites Perceptual evaluation of speech quality (PESQ)-a new method for speech quality assessment of telephone networks and codecs,.

Beyond Speaker Identity: Text Guided Target Speech Extraction Perceptual evaluation of speech quality (PESQ)-a new method for speech quality assessment of telephone networks and codecs,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:00.968610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:14:00.742812Z digest=sha256:0e421a696d3f1f11b036b73e58b31cfd4c66895da0d0fc2a1a11518096bccda9

Observation c4e21c3a-685c-4476-af3d-6a1e468f98e7 · outbound

This paper cites Conv-tasnet: Surpassing ideal time– frequency magnitude masking for speech separation,.

Beyond Speaker Identity: Text Guided Target Speech Extraction Conv-tasnet: Surpassing ideal time– frequency magnitude masking for speech separation,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:00.952605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:14:00.749916Z digest=sha256:a1ba2e74a1aa197777158eda9783f626b9221606864c50aa0cc44d6665ceee13

Observation 287ae483-999f-4ee8-ae62-19339118296c · outbound

This paper cites Self-supervised disentangled representation learning for robust target speech extraction,.

Beyond Speaker Identity: Text Guided Target Speech Extraction Self-supervised disentangled representation learning for robust target speech extraction,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:00.939537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:14:00.754527Z digest=sha256:82ee7448dfa429b916e15d33cb38c1c84225b471b5db773de153bf32bb78e1a9

Pith citing papers

No inbound Pith citation observations are available.