Pith. sign in

Paper Citation Record · LEDGER

PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction

As of 8 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 0 inbound Pith citation observations for arXiv:2607.08111.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.08111 v1

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-10T12:54:17.490105Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

28 of 28 outbound references displayed

  • verified exact9
  • verified fuzzy18
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d2f01861-67ed-4bcc-9e2f-22a0fc83e479 · outbound

This paper cites VoiceFilter: Targeted Voice Separation by Speaker-Conditioned Spectrogram Masking.

PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction VoiceFilter: Targeted Voice Separation by Speaker-Conditioned Spectrogram Masking

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-10T12:57:07.591118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T12:54:17.490105Z digest=sha256:f943b5ba6e52262745db251c75004d97fb21a7fe5bd39ff3d0a56f286e6996f9

Observation 97511a37-16b3-4bdb-b48b-cc998df44de2 · outbound

This paper cites Speakerfilter: Deep learning-based target speaker extraction using anchor speech,.

PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction Speakerfilter: Deep learning-based target speaker extraction using anchor speech,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T12:57:07.731870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T12:54:17.490105Z digest=sha256:c4ee212dcf51c7430f9e3ce498e79d9e8a636fe23bd9c606954bdc1fa6e20c33

Observation a6d5d5d7-6341-42d4-945e-dfbd1d71960c · outbound

This paper cites Neural target speech extraction: An overview,.

PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction Neural target speech extraction: An overview,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T12:57:07.754105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T12:54:17.490105Z digest=sha256:c0777c6a28120c5da67070ed74ab1205cb60b22f227d1aea8c7a90f71e2e0dee

Observation e6a1fbaa-405b-4d83-b31e-6448ab2068e1 · outbound

This paper cites Multimodal attention fusion for target speaker extraction,.

PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction Multimodal attention fusion for target speaker extraction,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T12:57:07.730190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T12:54:17.490105Z digest=sha256:921f30357c307b15c3619d5f36b787ce8c37da0ccc3a61113e2a3b81c373d80c

Observation 6c60e289-7a5d-4e25-b9b2-f444a8621ba6 · outbound

This paper cites Improving curriculum learning for target speaker extraction with synthetic speakers,.

PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction Improving curriculum learning for target speaker extraction with synthetic speakers,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T12:57:07.749000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T12:54:17.490105Z digest=sha256:661569de71d885ae7a71b247fb0a2d7ac1139b1a5e301cbc12e8df62a6b638d7

Observation 6ca1e4a0-0ea6-4acb-9374-b6549f135a8c · outbound

This paper cites Conv-tasnet: Surpassing ideal time– frequency magnitude masking for speech separation.

PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction Conv-tasnet: Surpassing ideal time– frequency magnitude masking for speech separation

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T12:57:07.755891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T12:54:17.490105Z digest=sha256:f95a600d77db97b9372f1866dab49dab045796fc538fd4e0c118f0caea843aeb

Observation 7b9a6719-2df7-4be1-9dba-ac3334696b63 · outbound

This paper cites Dual-path rnn: Efficient long sequence modeling for time-domain single-channel speech separation,.

PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction Dual-path rnn: Efficient long sequence modeling for time-domain single-channel speech separation,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T12:57:07.757677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T12:54:17.490105Z digest=sha256:df3f3ab42586aaf0bde1cef81f4ca91eccc1d486956a1ad1518b204cccefe34a

Observation 88fd459b-bb2e-4b91-8cb8-6158bf2d9d6b · outbound

This paper cites Attention is all you need in speech separation,.

PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction Attention is all you need in speech separation,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T12:57:07.733619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T12:54:17.490105Z digest=sha256:89791a08c4746d4b720433035c5af8d6652dab53a2462b0e4caab0ad7aa8da99

Observation 554a92f9-10f1-40d3-aa59-329714754699 · outbound

This paper cites Usef-tse: Universal speaker embedding free target speaker extraction,.

PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction Usef-tse: Universal speaker embedding free target speaker extraction,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T12:57:07.737022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T12:54:17.490105Z digest=sha256:5960069bb946c75896f569d41f9525b36899a66ee2af942b31bf769c61cb372a

Observation b0b89346-19b4-42e1-aa2f-9a012d5cdb37 · outbound

This paper cites Mc-lext: Multi-channel target speaker extraction with onset-prompted speaker conditioning mechanism,.

PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction Mc-lext: Multi-channel target speaker extraction with onset-prompted speaker conditioning mechanism,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T12:57:07.742181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T12:54:17.490105Z digest=sha256:df878ba09f5397b398e127c0f6b23591bcd9d0da5e048faf6e825d0857628943

Observation 3b0eaae2-912c-49b4-9964-ba0ed5944056 · outbound

This paper cites VoxCeleb: a large-scale speaker identification dataset.

PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction VoxCeleb: a large-scale speaker identification dataset

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-10T12:57:07.585404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T12:54:17.490105Z digest=sha256:1ffc075f9321c8fa52e9c908adda0653c26cdde77f42958cf842027bb774bc55

Observation 6526d0f5-f666-43ac-8e52-b41e47a769a1 · outbound

This paper cites Deep clustering: Discriminative embeddings for segmentation and separation,.

PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction Deep clustering: Discriminative embeddings for segmentation and separation,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T12:57:07.750695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T12:54:17.490105Z digest=sha256:faa86ddc7505cf69700878025253fb13844f66f35805bfe0ac1ba6a1f1ff1ec6

Observation d613abed-a7c3-418e-9e39-ea4a1dba0775 · outbound

This paper cites LibriMix: An Open-Source Dataset for Generalizable Speech Separation.

PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction LibriMix: An Open-Source Dataset for Generalizable Speech Separation

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-10T12:57:07.568338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T12:54:17.490105Z digest=sha256:9577dc21aea964025379fbf2093239151543314d99f91c0e9a62c6624a6635f6

Observation a362931b-6417-4c0f-92aa-83ee938b679b · outbound

This paper cites The ami meeting corpus: A pre-announcement,.

PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction The ami meeting corpus: A pre-announcement,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T12:57:07.735283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T12:54:17.490105Z digest=sha256:f019e0f69c8845eb85829d14d6e4fda3240c2808d7249fd19b51d290667568ca

Observation 0fd5958d-17ca-4604-9ac3-e28a0335e2e5 · outbound

This paper cites M2met: The icassp 2022 multi-channel multi- party meeting transcription challenge,.

PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction M2met: The icassp 2022 multi-channel multi- party meeting transcription challenge,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T12:57:07.728511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T12:54:17.490105Z digest=sha256:7693ab25ba80c161bc06966d5d006aa6d2b9fec27c4bd65d29c579414c5cb3a6

Observation 79cf3135-7a32-4bbe-ab95-7073fe312917 · outbound

This paper cites CHiME-6 Challenge:Tackling Multispeaker Speech Recognition for Unsegmented Recordings.

PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction CHiME-6 Challenge:Tackling Multispeaker Speech Recognition for Unsegmented Recordings

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-10T12:57:07.574165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T12:54:17.490105Z digest=sha256:bd50b474a5eae9ae62ae78ef210f2996fed8a7b72acc1de4824eafef86b058d3

Observation 463ff111-9c44-4383-a0b0-a1f9f4f62edb · outbound

This paper cites AISHELL-4: An Open Source Dataset for Speech Enhancement, Separation, Recognition and Speaker Diarization in Conference Scenario.

PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction AISHELL-4: An Open Source Dataset for Speech Enhancement, Separation, Recognition and Speaker Diarization in Conference Scenario

Reference 17

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T12:57:07.582762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T12:54:17.490105Z digest=sha256:92c1dcfbfb7cc065ca58a3f3a3b943665e974c1b6828ba81bb5db660d19842e8

Observation a965f09c-fb7b-4f3b-9762-05ea6ef3bf37 · outbound

This paper cites DiPCo -- Dinner Party Corpus.

PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction DiPCo -- Dinner Party Corpus

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-07-10T12:57:07.579924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T12:54:17.490105Z digest=sha256:60f59e07df92c0629f24031a636e6c94adf62b3741db3a962cdd48c09690b7e6

Observation ad7f40aa-4115-4d27-a109-d67dfbb693c7 · outbound

This paper cites REAL-T: Real Conversational Mixtures for Target Speaker Extraction,.

PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction REAL-T: Real Conversational Mixtures for Target Speaker Extraction,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T12:57:07.747258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T12:54:17.490105Z digest=sha256:cf6db303397a69abd20770a82008a46b0e218d9170fd5ada8ba2d94cd6b40ede

Observation 57a8e247-aafe-4edc-bc59-7bcd318aac71 · outbound

This paper cites Sdr–half-baked or well done?.

PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction Sdr–half-baked or well done?

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T12:57:07.745482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T12:54:17.490105Z digest=sha256:5305186e38ca4e4f44849e416b2f82f62217dae08b1af992a325ea9aa9aad1ff

Observation 1dda5e59-a9cc-494a-a3ed-6354d203b409 · outbound

This paper cites A statistical model-based voice activity detection,.

PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction A statistical model-based voice activity detection,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T12:57:07.743883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T12:54:17.490105Z digest=sha256:9d5fcceea300cc0bad9142cd50300bb2e5297229883e38e94fc4f0b4e68d1a77

Observation a808baec-1e83-4c55-9f59-4a11deff5b02 · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision.

PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction Robust Speech Recognition via Large-Scale Weak Supervision

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-07-10T12:57:07.588117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T12:54:17.490105Z digest=sha256:0d184ed6a8cf8489a54568da32d53b3b31684c5704aed33e6e05b0041eb81776

Observation f2a0fbfe-7974-44d5-b8ca-4a57a5ac6841 · outbound

This paper cites High Fidelity Speech Enhancement with Band-split RNN.

PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction High Fidelity Speech Enhancement with Band-split RNN

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-07-10T12:57:07.565527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T12:54:17.490105Z digest=sha256:76c4437add8759f1f979ca0e96daed4558e55301a290cf89edacef8340a1e4ee

Observation 92e32f92-5497-4c58-a268-ba1c521a3fa8 · outbound

This paper cites ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification.

PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-07-10T12:57:07.596894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T12:54:17.490105Z digest=sha256:7f978f6e8226d5b0c76f22ed696f1ae843287f89968e9bf6b0f9d581809506de

Observation 117291d4-408e-4a68-aa7a-cfefd13098a1 · outbound

This paper cites Cross-entropy loss functions: Theoretical analysis and applications,.

PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction Cross-entropy loss functions: Theoretical analysis and applications,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T12:57:07.740393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T12:54:17.490105Z digest=sha256:c0572f2dc3751de4dc4dc551e148fd6797ce0fcbb4f2f17c4f8e92cf8a97f7f7

Observation a4c1081b-c9f1-41f6-b031-f824b2c33f27 · outbound

This paper cites Wespeaker: A research and production oriented speaker embedding learning toolkit,.

PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction Wespeaker: A research and production oriented speaker embedding learning toolkit,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T12:57:07.738715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T12:54:17.490105Z digest=sha256:8e93d3b8fe8b7d2acbd1de65b286d7d038aacc0fddb06de4c2e2be4faaa6cfd4

Observation 46ebe8a1-5261-487f-ac1b-0622096fc1b0 · outbound

This paper cites Dnsmos: A non-intrusive perceptual objective speech quality metric to evaluate noise suppressors,.

PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction Dnsmos: A non-intrusive perceptual objective speech quality metric to evaluate noise suppressors,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T12:57:07.752411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T12:54:17.490105Z digest=sha256:fec0a72b2f1af6355619c457f0ba07473639d9479fb43c537be24820b95d268e

Observation 855dc388-b15c-4da1-8957-6444778d835e · outbound

This paper cites An Open source Implementation of ITU-T Recommendation P.808 with Validation.

PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction An Open source Implementation of ITU-T Recommendation P.808 with Validation

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-07-10T12:57:07.571418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T12:54:17.490105Z digest=sha256:2f284bba9accb81dd461903a6c19ade7b306f7ab221071d5cd719a77e62ddd0f

Pith citing papers

No inbound Pith citation observations are available.