Pith. sign in

Paper Citation Record · LEDGER

PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction

As of 21 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 1 inbound Pith citation observation for arXiv:2607.08111.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.08111 v1

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-10T12:54:17.490105Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T20:11:17.417789Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-11T20:11:17.767198Z

Reference resolution

28 of 28 outbound references displayed

  • verified exact9
  • verified fuzzy18
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d2f01861-67ed-4bcc-9e2f-22a0fc83e479 · outbound

This paper cites VoiceFilter: Targeted Voice Separation by Speaker-Conditioned Spectrogram Masking.

PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction VoiceFilter: Targeted Voice Separation by Speaker-Conditioned Spectrogram Masking

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-10T12:57:07.591118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T12:54:17.490105Z digest=sha256:1660ef64dcbb4ffab0c6fcc7cb46b69e0c17eca4f58feeea3097dfc5a1d72e26

Observation 97511a37-16b3-4bdb-b48b-cc998df44de2 · outbound

This paper cites Speakerfilter: Deep learning-based target speaker extraction using anchor speech,.

PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction Speakerfilter: Deep learning-based target speaker extraction using anchor speech,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T12:57:07.731870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T12:54:17.490105Z digest=sha256:6af968f03eebed1bb1629ead59d9021620b2a827e97195bc7574457d220ceca9

Observation a6d5d5d7-6341-42d4-945e-dfbd1d71960c · outbound

This paper cites Neural target speech extraction: An overview,.

PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction Neural target speech extraction: An overview,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T12:57:07.754105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T12:54:17.490105Z digest=sha256:e7b233604ed2f0adb0fd20a732163ecd865220179a27c76223ac5b7c9d53109b

Observation e6a1fbaa-405b-4d83-b31e-6448ab2068e1 · outbound

This paper cites Multimodal attention fusion for target speaker extraction,.

PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction Multimodal attention fusion for target speaker extraction,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T12:57:07.730190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T12:54:17.490105Z digest=sha256:5e705f49536f901e5333b58d0b7a1d2c8834d86be1064846f36a108455c5de6b

Observation 6c60e289-7a5d-4e25-b9b2-f444a8621ba6 · outbound

This paper cites Improving curriculum learning for target speaker extraction with synthetic speakers,.

PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction Improving curriculum learning for target speaker extraction with synthetic speakers,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T12:57:07.749000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T12:54:17.490105Z digest=sha256:6c99ece596c5f42dcf2eb3d0d61fdf027d4dbf1405d3c6b7cc5a1e40a7b3abd6

Observation 6ca1e4a0-0ea6-4acb-9374-b6549f135a8c · outbound

This paper cites Conv-tasnet: Surpassing ideal time– frequency magnitude masking for speech separation.

PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction Conv-tasnet: Surpassing ideal time– frequency magnitude masking for speech separation

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T12:57:07.755891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T12:54:17.490105Z digest=sha256:1b01c4f9f0503a28c23c7ecb9eaf7e0990d2d4ba2605c3f4713e72ac1db99f38

Observation 7b9a6719-2df7-4be1-9dba-ac3334696b63 · outbound

This paper cites Dual-path rnn: Efficient long sequence modeling for time-domain single-channel speech separation,.

PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction Dual-path rnn: Efficient long sequence modeling for time-domain single-channel speech separation,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T12:57:07.757677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T12:54:17.490105Z digest=sha256:2fabed305660eae5fc6c4e0eedc0124be333b2c13017b49d4a294e1da2cded9d

Observation 88fd459b-bb2e-4b91-8cb8-6158bf2d9d6b · outbound

This paper cites Attention is all you need in speech separation,.

PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction Attention is all you need in speech separation,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T12:57:07.733619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T12:54:17.490105Z digest=sha256:60a35bdfdc13992684f22b7eeeff73774b91dab4c5ed06c5c6b2b5aa908e2297

Observation 554a92f9-10f1-40d3-aa59-329714754699 · outbound

This paper cites Usef-tse: Universal speaker embedding free target speaker extraction,.

PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction Usef-tse: Universal speaker embedding free target speaker extraction,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T12:57:07.737022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T12:54:17.490105Z digest=sha256:c8ed232b96dc13a223da5d30664748f1d73aaa319b87762e4a2018ff8903c838

Observation b0b89346-19b4-42e1-aa2f-9a012d5cdb37 · outbound

This paper cites Mc-lext: Multi-channel target speaker extraction with onset-prompted speaker conditioning mechanism,.

PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction Mc-lext: Multi-channel target speaker extraction with onset-prompted speaker conditioning mechanism,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T12:57:07.742181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T12:54:17.490105Z digest=sha256:4b6a0f9c84cf5a53e68bc48a889fc070a48c3cee81f958bf6f319d9e2e1c77df

Observation 3b0eaae2-912c-49b4-9964-ba0ed5944056 · outbound

This paper cites VoxCeleb: a large-scale speaker identification dataset.

PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction VoxCeleb: a large-scale speaker identification dataset

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-10T12:57:07.585404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T12:54:17.490105Z digest=sha256:3422a79e1d071292c2b1210ca6526f646ac836455c41ab317a7e3f7eb4da6e03

Observation 6526d0f5-f666-43ac-8e52-b41e47a769a1 · outbound

This paper cites Deep clustering: Discriminative embeddings for segmentation and separation,.

PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction Deep clustering: Discriminative embeddings for segmentation and separation,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T12:57:07.750695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T12:54:17.490105Z digest=sha256:1096780ef1db89ebfd90ab06e3ffbc1759d58d90d9f937cf6c63355c227210a6

Observation d613abed-a7c3-418e-9e39-ea4a1dba0775 · outbound

This paper cites LibriMix: An Open-Source Dataset for Generalizable Speech Separation.

PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction LibriMix: An Open-Source Dataset for Generalizable Speech Separation

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-10T12:57:07.568338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T12:54:17.490105Z digest=sha256:1a0db8b86f2006541b92b812fa7634e53c4572f5ae358521a36acbb39049a2b5

Observation a362931b-6417-4c0f-92aa-83ee938b679b · outbound

This paper cites The ami meeting corpus: A pre-announcement,.

PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction The ami meeting corpus: A pre-announcement,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T12:57:07.735283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T12:54:17.490105Z digest=sha256:18c8ea00397e4cc39090ad890a4664ca4c25789f1a1a08db784dc43fb78c4b98

Observation 0fd5958d-17ca-4604-9ac3-e28a0335e2e5 · outbound

This paper cites M2met: The icassp 2022 multi-channel multi- party meeting transcription challenge,.

PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction M2met: The icassp 2022 multi-channel multi- party meeting transcription challenge,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T12:57:07.728511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T12:54:17.490105Z digest=sha256:bdc80beceb071212757a00e50a9b4a5047822ad6b1c8e9daae4a49ec9cd4c922

Observation 79cf3135-7a32-4bbe-ab95-7073fe312917 · outbound

This paper cites CHiME-6 Challenge:Tackling Multispeaker Speech Recognition for Unsegmented Recordings.

PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction CHiME-6 Challenge:Tackling Multispeaker Speech Recognition for Unsegmented Recordings

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-10T12:57:07.574165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T12:54:17.490105Z digest=sha256:708d91aebd6d6ab31b7aeaf2a544ab51855a704022d9f363962c8b273059717b

Observation 463ff111-9c44-4383-a0b0-a1f9f4f62edb · outbound

This paper cites AISHELL-4: An Open Source Dataset for Speech Enhancement, Separation, Recognition and Speaker Diarization in Conference Scenario.

PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction AISHELL-4: An Open Source Dataset for Speech Enhancement, Separation, Recognition and Speaker Diarization in Conference Scenario

Reference 17

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T12:57:07.582762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T12:54:17.490105Z digest=sha256:cffd3ca0b7426eb357e3d94f2410c3ea234ac15354803d6a5faaeeb2bd802073

Observation a965f09c-fb7b-4f3b-9762-05ea6ef3bf37 · outbound

This paper cites DiPCo -- Dinner Party Corpus.

PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction DiPCo -- Dinner Party Corpus

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-07-10T12:57:07.579924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T12:54:17.490105Z digest=sha256:788502a69bd883c43bb99b3c95b1b311eb78414de26592c0f470bfc031a67215

Observation ad7f40aa-4115-4d27-a109-d67dfbb693c7 · outbound

This paper cites REAL-T: Real Conversational Mixtures for Target Speaker Extraction,.

PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction REAL-T: Real Conversational Mixtures for Target Speaker Extraction,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T12:57:07.747258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T12:54:17.490105Z digest=sha256:77855ff5a3f5dba941abb2c2a06a98c74c6e8924e0ea02da9479afc7ec4ca66f

Observation 57a8e247-aafe-4edc-bc59-7bcd318aac71 · outbound

This paper cites Sdr–half-baked or well done?.

PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction Sdr–half-baked or well done?

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T12:57:07.745482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T12:54:17.490105Z digest=sha256:89050be288ec60a021473cbdbd3b94f16c79d9e53ffafe4c7866d6d2f05bf76b

Observation 1dda5e59-a9cc-494a-a3ed-6354d203b409 · outbound

This paper cites A statistical model-based voice activity detection,.

PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction A statistical model-based voice activity detection,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T12:57:07.743883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T12:54:17.490105Z digest=sha256:307a9d9fa484f8f869bb3855fec0f4d2a60946bb7f3b23d3f74d39ce217bd614

Observation a808baec-1e83-4c55-9f59-4a11deff5b02 · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision.

PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction Robust Speech Recognition via Large-Scale Weak Supervision

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-07-10T12:57:07.588117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T12:54:17.490105Z digest=sha256:0050e60f792839c968b43558d968dd85b42e37720ad16d650580f41fe7eebfc6

Observation f2a0fbfe-7974-44d5-b8ca-4a57a5ac6841 · outbound

This paper cites High Fidelity Speech Enhancement with Band-split RNN.

PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction High Fidelity Speech Enhancement with Band-split RNN

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-07-10T12:57:07.565527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T12:54:17.490105Z digest=sha256:d98f5b56c2eddc9cc05c48e76b9e1240ca54b3f52fdddbe99902aec0a8bb4ab9

Observation 92e32f92-5497-4c58-a268-ba1c521a3fa8 · outbound

This paper cites ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification.

PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-07-10T12:57:07.596894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T12:54:17.490105Z digest=sha256:207f0c705ab64725884e48eb6c98c96c370fd07939aefb5477c117fbf6544e33

Observation 117291d4-408e-4a68-aa7a-cfefd13098a1 · outbound

This paper cites Cross-entropy loss functions: Theoretical analysis and applications,.

PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction Cross-entropy loss functions: Theoretical analysis and applications,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T12:57:07.740393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T12:54:17.490105Z digest=sha256:951e61ffef54292b0d1ef8897958f414702197ce5e9b1e052e16797be7fdcdd1

Observation a4c1081b-c9f1-41f6-b031-f824b2c33f27 · outbound

This paper cites Wespeaker: A research and production oriented speaker embedding learning toolkit,.

PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction Wespeaker: A research and production oriented speaker embedding learning toolkit,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T12:57:07.738715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T12:54:17.490105Z digest=sha256:2fb0fb74520a07470eaa08fd4531b96f9dac5ef5918bd322ff405d63e5bd6180

Observation 46ebe8a1-5261-487f-ac1b-0622096fc1b0 · outbound

This paper cites Dnsmos: A non-intrusive perceptual objective speech quality metric to evaluate noise suppressors,.

PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction Dnsmos: A non-intrusive perceptual objective speech quality metric to evaluate noise suppressors,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T12:57:07.752411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T12:54:17.490105Z digest=sha256:47469a80b2ad42a4603c4a5fab51d55b9477b06bba6371d79d8cf0ac74b1b7a9

Observation 855dc388-b15c-4da1-8957-6444778d835e · outbound

This paper cites An Open source Implementation of ITU-T Recommendation P.808 with Validation.

PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction An Open source Implementation of ITU-T Recommendation P.808 with Validation

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-07-10T12:57:07.571418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T12:54:17.490105Z digest=sha256:5bf0e5fadc4b28bbc5ca040d161b2d014bc9aa02fae19c13496d5c37e5fa17b5

Pith citing papers

Observation 3a2c279c-6dfb-463f-8a7b-c2b2b003c788 · inbound

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation cites this paper.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction

Reference 2026

Resolution
metadata mismatch
local_arxiv, observed 2026-08-11T20:11:17.777692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T20:11:17.417789Z digest=sha256:27b46ac8e2572e58757701e182cd86d49fd879c04c444a7d53864a03b3080a8d