Pith. sign in

Paper Citation Record · LEDGER

Noise-Robust Target-Speaker Voice Activity Detection Through Self-Supervised Pretraining

As of 15 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 0 inbound Pith citation observations for arXiv:2501.03184.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.03184 v1

Coverage vector

measured 53 of 53 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T21:57:02.568478Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

53 of 53 outbound references displayed

  • verified exact0
  • verified fuzzy38
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8cc40807-84df-4175-8b72-6c22edfb4573 · outbound

This paper cites rV AD: An unsupervised segment-based robust voice activity detection method,.

Noise-Robust Target-Speaker Voice Activity Detection Through Self-Supervised Pretraining rV AD: An unsupervised segment-based robust voice activity detection method,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:03.419230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:57:02.311351Z digest=sha256:019458b1f49f04480bf5f82b63b1d276416d67b4dd0ccefa7abf8a2beca13646

Observation 9c0ce093-394f-4588-8194-4bbdd8f3c09f · outbound

This paper cites V oice activity detection in the wild: A data-driven approach using teacher-student training,.

Noise-Robust Target-Speaker Voice Activity Detection Through Self-Supervised Pretraining V oice activity detection in the wild: A data-driven approach using teacher-student training,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T21:57:02.316337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:57:02.316337Z digest=sha256:17b610b97b54143394f922f7d1a34ae5c50951e673a571eee1925af704579db6

Observation e2e559a8-8e5a-4b97-8b68-763bacd4cc8e · outbound

This paper cites Robust voice activity detection using an auditory-inspired masked modulation encoder based convolutional attention network,.

Noise-Robust Target-Speaker Voice Activity Detection Through Self-Supervised Pretraining Robust voice activity detection using an auditory-inspired masked modulation encoder based convolutional attention network,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:03.391416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:57:02.320789Z digest=sha256:933b34f189bc25b16d3b81a27e08e3ceb2bcdc3797873ce26961344db79a98ad

Observation e159f841-f467-47bd-a409-1cb1268db226 · outbound

This paper cites V oice biometric system security: Design and analysis of countermeasures for replay attacks.

Noise-Robust Target-Speaker Voice Activity Detection Through Self-Supervised Pretraining V oice biometric system security: Design and analysis of countermeasures for replay attacks

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:03.373459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:57:02.325687Z digest=sha256:67e4cc399a26d43b6c3fd9f103d39ac8fb6358fda96eeebf2c9f4697954c8f5c

Observation 0fb1f21d-3405-4bcc-b25d-7fb0c59fa4b8 · outbound

This paper cites Audiovisual speaker indexing for web-tv automations,.

Noise-Robust Target-Speaker Voice Activity Detection Through Self-Supervised Pretraining Audiovisual speaker indexing for web-tv automations,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:03.357799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:57:02.331212Z digest=sha256:09fd1fa8b6d40c348735700a2ed20bb61604ae705b80998a3b0fce2c4596fb17

Observation e83e1b8c-0cb6-4aca-a7bc-9bfcb55ca68f · outbound

This paper cites Neural target speech extraction: An overview,.

Noise-Robust Target-Speaker Voice Activity Detection Through Self-Supervised Pretraining Neural target speech extraction: An overview,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:03.343696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:57:02.337966Z digest=sha256:6184ba79ad9f91c2f91d0938b289544fcb7fbda7cabe92b7b4df145b0b213a35

Observation 4f20ab46-882f-4a7b-940e-e82de34eeb48 · outbound

This paper cites Speakerbeam: Speaker aware neural network for target speaker extraction in speech mixtures,.

Noise-Robust Target-Speaker Voice Activity Detection Through Self-Supervised Pretraining Speakerbeam: Speaker aware neural network for target speaker extraction in speech mixtures,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:03.327807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:57:02.343699Z digest=sha256:43b15d8bbe86f1ff225767383426de7319e18fe525625840fd12ea31c977a601

Observation 0e755325-d70e-480c-aaba-4572f718cb3e · outbound

This paper cites Per- sonal V AD: Speaker-conditioned voice activity detection,.

Noise-Robust Target-Speaker Voice Activity Detection Through Self-Supervised Pretraining Per- sonal V AD: Speaker-conditioned voice activity detection,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:03.310965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:57:02.348715Z digest=sha256:7dbabfa9f36b0db5214bafe80ca6794c62d9a62b825f189f661918b7bb0fb5be

Observation 29ac4f61-64a3-46e6-8ff0-58f456621b24 · outbound

This paper cites Personal V AD 2.0: Optimizing personal voice activity detection for on-device speech recognition,.

Noise-Robust Target-Speaker Voice Activity Detection Through Self-Supervised Pretraining Personal V AD 2.0: Optimizing personal voice activity detection for on-device speech recognition,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:03.295892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:57:02.352736Z digest=sha256:62e795befdad72e73113c00037f933756fc1af083dafb06141ca4fc0848951d3

Observation 00acbc08-e568-4731-9611-f1bf0e91a745 · outbound

This paper cites Target speech extraction with pre-trained self-supervised learning models,.

Noise-Robust Target-Speaker Voice Activity Detection Through Self-Supervised Pretraining Target speech extraction with pre-trained self-supervised learning models,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:03.280845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:57:02.357023Z digest=sha256:20d8611f63efbeb2cfefe2d36c8b7cf172315c8f575bf71d4f71f0d5f4f3cecc

Observation 7553fe0a-d2e9-4439-a6dd-6fa360d8f40a · outbound

This paper cites Target-speaker voice activity detection: a novel approach for multi-speaker diarization in a dinner party scenario,.

Noise-Robust Target-Speaker Voice Activity Detection Through Self-Supervised Pretraining Target-speaker voice activity detection: a novel approach for multi-speaker diarization in a dinner party scenario,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:03.264954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:57:02.361696Z digest=sha256:70185c0a7a743b546e6c0ff7113891a942c660f7165c712d1d687f7c115cd158

Observation 79fe2f20-69f2-492b-aa6a-05fdbc71a639 · outbound

This paper cites Target- speaker voice activity detection with improved i-vector estimation for unknown number of speaker,.

Noise-Robust Target-Speaker Voice Activity Detection Through Self-Supervised Pretraining Target- speaker voice activity detection with improved i-vector estimation for unknown number of speaker,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:03.251477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:57:02.366910Z digest=sha256:d77dca1a14121756d73dc760ae728b2f09f31e90f6891d52843e95991ee7d29c

Observation aa4bdbd9-ea5e-4afb-8efe-071424a1de4f · outbound

This paper cites Target speaker voice activity detection with transformers and its integration with end- to-end neural diarization,.

Noise-Robust Target-Speaker Voice Activity Detection Through Self-Supervised Pretraining Target speaker voice activity detection with transformers and its integration with end- to-end neural diarization,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:03.235495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:57:02.370489Z digest=sha256:e03922b80b7ad037df67909830a58d73e89d2da87c18d90c285ecb8397048c7b

Observation 0ccae93b-eab4-46ec-8eea-ace328e26948 · outbound

This paper cites Profile-error-tolerant target-speaker voice activity detection,.

Noise-Robust Target-Speaker Voice Activity Detection Through Self-Supervised Pretraining Profile-error-tolerant target-speaker voice activity detection,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:03.215029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:57:02.374049Z digest=sha256:91c8c48d3eb48aa00e495148a803129cbc3e20a0bb5f0b1068272e708ef06bcf

Observation 7dfb6889-7ba1-4155-b32b-c59f751ca9ea · outbound

This paper cites Multimodal attention fusion for target speaker extraction,.

Noise-Robust Target-Speaker Voice Activity Detection Through Self-Supervised Pretraining Multimodal attention fusion for target speaker extraction,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:03.193356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:57:02.378202Z digest=sha256:698ca8fd5d5f61911401154f92a8ccd2eb6b0530f8aa70003a027e25aeabc7b6

Observation 3b42f450-f794-4c41-b1b3-5deaef20b1dd · outbound

This paper cites Multimodal speakerbeam: Single channel target speech extraction with audio-visual speaker clues.

Noise-Robust Target-Speaker Voice Activity Detection Through Self-Supervised Pretraining Multimodal speakerbeam: Single channel target speech extraction with audio-visual speaker clues

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:03.173474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:57:02.383782Z digest=sha256:001af67cbda1cb7bb5a39da2ee8d3a89d88c4729e86952cc9d961fe4cc523a0b

Observation 1a02336b-b7a5-4672-b898-fe58cb7efdf0 · outbound

This paper cites Av-sepformer: Cross-attention sepformer for audio-visual target speaker extraction,.

Noise-Robust Target-Speaker Voice Activity Detection Through Self-Supervised Pretraining Av-sepformer: Cross-attention sepformer for audio-visual target speaker extraction,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:03.155325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:57:02.388408Z digest=sha256:4f6ab3bb86487ad1c757409d0de775aedaf402781b3c2b967805db7f0beb3bae

Observation 19382b48-c880-40c7-b48c-802cdb5da07f · outbound

This paper cites Usev: Universal speaker extraction with visual cue,.

Noise-Robust Target-Speaker Voice Activity Detection Through Self-Supervised Pretraining Usev: Universal speaker extraction with visual cue,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:03.135132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:57:02.392385Z digest=sha256:f90458c4d8a3d52648f8f89e35c52b87266bb59c048b6537423234af82e1ff88

Observation dc16625e-cfe1-42fe-9122-16d3702c6751 · outbound

This paper cites Brain-informed speech separation (biss) for enhancement of target speaker in multitalker speech perception,.

Noise-Robust Target-Speaker Voice Activity Detection Through Self-Supervised Pretraining Brain-informed speech separation (biss) for enhancement of target speaker in multitalker speech perception,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:03.117893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:57:02.398227Z digest=sha256:ff25f5b47409a309026021395dab1496f8458ebd162158e93a9be7a8099cc260

Observation e29ad883-373e-4a45-bd88-5ebb6c64f633 · outbound

This paper cites Target speaker detection with concealed eeg around the ear,.

Noise-Robust Target-Speaker Voice Activity Detection Through Self-Supervised Pretraining Target speaker detection with concealed eeg around the ear,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:03.100835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:57:02.402985Z digest=sha256:eb4fe0b1527d329cdef16e98f7e7410fe072d7efe86a07972c6ee4e1a66f1084

Observation 0bb8d991-43ab-43f8-8152-c28403f145ba · outbound

This paper cites Eeg decoding of the target speaker in a cocktail party scenario: considerations regarding dynamic switching of talker location,.

Noise-Robust Target-Speaker Voice Activity Detection Through Self-Supervised Pretraining Eeg decoding of the target speaker in a cocktail party scenario: considerations regarding dynamic switching of talker location,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:03.084500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:57:02.408508Z digest=sha256:c1f57b028262141302af97a57b85f57e81d40de669458304f1621d89370ac1a4

Observation 15a3f4c4-2394-448f-9080-29f66d731cb0 · outbound

This paper cites Cortical auditory at- tention decoding during music and speech listening,.

Noise-Robust Target-Speaker Voice Activity Detection Through Self-Supervised Pretraining Cortical auditory at- tention decoding during music and speech listening,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:03.063569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:57:02.414172Z digest=sha256:9dbaf9fa31dc76fbc6c8e953f90b73a75e82973542fe06225cdb1766d4be443b

Observation a8d165d2-e984-48b7-9481-a7e9358b23cc · outbound

This paper cites Self-supervised speech representation learning: A review,.

Noise-Robust Target-Speaker Voice Activity Detection Through Self-Supervised Pretraining Self-supervised speech representation learning: A review,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T21:57:02.419361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:57:02.419361Z digest=sha256:2dd86f54e242efce14d004c63ba155a59eba40e1ae314d6e5545c0bedbfa4d18

Observation 52bd3b3a-6c90-415f-9080-309215240994 · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech representations,.

Noise-Robust Target-Speaker Voice Activity Detection Through Self-Supervised Pretraining wav2vec 2.0: A framework for self-supervised learning of speech representations,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T21:57:02.426095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:57:02.426095Z digest=sha256:8b4a41b8c3be78f39e5ddc5bc95ae9883b03eca2198b8796fa4171b72a602a8f

Observation b2473e1b-e498-4d17-b720-fec15d1a48ac · outbound

This paper cites HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,.

Noise-Robust Target-Speaker Voice Activity Detection Through Self-Supervised Pretraining HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T21:57:02.431390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:57:02.431390Z digest=sha256:c966eea9c188a1266970930f185aba6c6700e1a617d74f35ba1e6557dfd3f6a8

Observation cde589c0-16ba-4cb7-a340-5fdd469f4ed4 · outbound

This paper cites WavLM: Large-scale self- supervised pre-training for full stack speech processing,.

Noise-Robust Target-Speaker Voice Activity Detection Through Self-Supervised Pretraining WavLM: Large-scale self- supervised pre-training for full stack speech processing,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T21:57:02.436697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:57:02.436697Z digest=sha256:74ccac1bc2d75cce47d54a97ef1893841fa238aa459373e506dd14ea2d38237e

Observation d77e2797-a8d8-4b2b-9601-01914e3a64fa · outbound

This paper cites Self-supervised pretraining for robust personalized voice activity detection in adverse conditions,.

Noise-Robust Target-Speaker Voice Activity Detection Through Self-Supervised Pretraining Self-supervised pretraining for robust personalized voice activity detection in adverse conditions,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:02.992252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:57:02.441426Z digest=sha256:ac030e99667d43f439de3e99a1fe713a602cc2a9e381c4cbe78363c932822f72

Observation 2a2b6f54-e75d-4080-8c70-4d593ba5336c · outbound

This paper cites Features for voice activity detection: a comparative analysis,.

Noise-Robust Target-Speaker Voice Activity Detection Through Self-Supervised Pretraining Features for voice activity detection: a comparative analysis,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:02.974156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:57:02.448001Z digest=sha256:24471f45b6c074464e00ff39b26a98f85805ed8db913bd976983615dcee28d96

Observation a6d98d99-5e1d-4057-b6bd-7068c0664229 · outbound

This paper cites I-vector-based speaker adaptation of deep neural networks for french broadcast audio transcription,.

Noise-Robust Target-Speaker Voice Activity Detection Through Self-Supervised Pretraining I-vector-based speaker adaptation of deep neural networks for french broadcast audio transcription,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:02.959356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:57:02.453109Z digest=sha256:3d167a48a7a45f737e1beb4ff9a4abe6b43b8c981e1573d82fccee3b86e5ebf6

Observation 5fee9e3f-d381-4380-8f4f-36db9c9ca7fb · outbound

This paper cites Deep neural networks for small footprint text-dependent speaker verification,.

Noise-Robust Target-Speaker Voice Activity Detection Through Self-Supervised Pretraining Deep neural networks for small footprint text-dependent speaker verification,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:02.943226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:57:02.458119Z digest=sha256:b9c24d42e8e38e9c61ae3095f67c76eef1219cf0a02374f4a41f32489a603e6e

Observation 7054ccb0-e335-4f80-a457-8727123350cf · outbound

This paper cites Generalized end-to-end loss for speaker verification,.

Noise-Robust Target-Speaker Voice Activity Detection Through Self-Supervised Pretraining Generalized end-to-end loss for speaker verification,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:02.922685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:57:02.463382Z digest=sha256:f0e3a4489951ccd0ab2715049800a867bd4b71a618e84525ee650c726de4c7ef

Observation 133c8f00-d408-4709-8214-564c1049c6a8 · outbound

This paper cites X- vectors: Robust dnn embeddings for speaker recognition,.

Noise-Robust Target-Speaker Voice Activity Detection Through Self-Supervised Pretraining X- vectors: Robust dnn embeddings for speaker recognition,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:02.906596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:57:02.468419Z digest=sha256:27fda1dafd4e020af26b6689792ada1c1a3a6304f627365c0ef6ec19686d8928

Observation 27fbdb2d-edc3-49b5-8783-6c286e145c80 · outbound

This paper cites Attention is all you need,.

Noise-Robust Target-Speaker Voice Activity Detection Through Self-Supervised Pretraining Attention is all you need,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:02.890665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:57:02.472674Z digest=sha256:5b57e2eb7bcd2db86a59a4222f82e3c3d79cee3181e15bdf6542fb5298e34c0f

Observation dc9b081e-4a65-40ee-9624-3d2427f665fe · outbound

This paper cites Film: Visual reasoning with a general conditioning layer,.

Noise-Robust Target-Speaker Voice Activity Detection Through Self-Supervised Pretraining Film: Visual reasoning with a general conditioning layer,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T21:57:02.476956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:57:02.476956Z digest=sha256:10b2e7d71925fbbeb2698774a1fe76c058d7ad17858edcf60e7911ebef281da8

Observation 618a3ce4-b723-4b22-80b9-4f777cf353b1 · outbound

This paper cites Conditional conformer: Improving speaker modulation for single and multi-user speech enhancement,.

Noise-Robust Target-Speaker Voice Activity Detection Through Self-Supervised Pretraining Conditional conformer: Improving speaker modulation for single and multi-user speech enhancement,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:02.851599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:57:02.480619Z digest=sha256:11e3ab4eab0677d72321e890011f4064348e06a66143a7593e8ad8f31b4b0f36

Observation bf350042-69c2-402f-9aa1-3dd5cddba11c · outbound

This paper cites Searching for Activation Functions.

Noise-Robust Target-Speaker Voice Activity Detection Through Self-Supervised Pretraining Searching for Activation Functions

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T21:57:02.484788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:57:02.484788Z digest=sha256:7ec87be96527e46f10ebbfe351687b7e8d13ed9d03e65578cc4711ab5a7df43e

Observation c1ef7187-05f6-4974-a845-d5a704bffab0 · outbound

This paper cites Using self- supervised learning can improve model robustness and uncertainty,.

Noise-Robust Target-Speaker Voice Activity Detection Through Self-Supervised Pretraining Using self- supervised learning can improve model robustness and uncertainty,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:02.835380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:57:02.490368Z digest=sha256:30cabea8f60b0792221aa84bd56aaa70a9daf23e80efc0a5024874404530f502

Observation f6c5bddc-983f-4a1f-835e-3687bdfc609f · outbound

This paper cites Noise-Robust Keyword Spotting through Self-supervised Pretraining.

Noise-Robust Target-Speaker Voice Activity Detection Through Self-Supervised Pretraining Noise-Robust Keyword Spotting through Self-supervised Pretraining

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T21:57:02.495311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:57:02.495311Z digest=sha256:c2d30ac81a809e64e4e9acb65cd1b66fb08df17324d6ebfddb9ec8f7fa18f1a7

Observation 695e7c77-3336-4239-9d9c-3e7126522883 · outbound

This paper cites Bootstrap predictive coding: Investigating a non-contrastive self-supervised learning approach,.

Noise-Robust Target-Speaker Voice Activity Detection Through Self-Supervised Pretraining Bootstrap predictive coding: Investigating a non-contrastive self-supervised learning approach,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:02.820023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:57:02.499682Z digest=sha256:8dafa9a72476ce589122e08a38607d2904a95849535ea58f3fe1910d20bbbb26

Observation f179beba-3bba-4b46-b126-8535aa2eb225 · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

Noise-Robust Target-Speaker Voice Activity Detection Through Self-Supervised Pretraining Representation Learning with Contrastive Predictive Coding

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T21:57:02.505721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:57:02.505721Z digest=sha256:b09a9789478e43b5f60da2441dfe4158b859e4fbf4b8012582b3fc442e5a5c6c

Observation 4d18faa0-88ff-4295-a9ab-1d06143b30c0 · outbound

This paper cites Autore- gressive predictive coding: A comprehensive study,.

Noise-Robust Target-Speaker Voice Activity Detection Through Self-Supervised Pretraining Autore- gressive predictive coding: A comprehensive study,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:02.806167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:57:02.509729Z digest=sha256:8df2ebc548d5ec1cfe9469d27cbb352cd8734dcb58ae53549d42319a0bca3253

Observation 79ad0697-058e-4d63-acef-9a9d990bf29c · outbound

This paper cites Generative pre-training for speech with au- toregressive predictive coding,.

Noise-Robust Target-Speaker Voice Activity Detection Through Self-Supervised Pretraining Generative pre-training for speech with au- toregressive predictive coding,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:02.791522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:57:02.512913Z digest=sha256:1c2e7eeb40f92ef95fac69730b1610f5f2c3823a1c40d9f558bd119b3007fe59

Observation 3404282a-bc80-43ef-8552-79bbb936a3bb · outbound

This paper cites Pytorch: An imperative style, high- performance deep learning library,.

Noise-Robust Target-Speaker Voice Activity Detection Through Self-Supervised Pretraining Pytorch: An imperative style, high- performance deep learning library,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:02.777412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:57:02.516729Z digest=sha256:42cd56f51bd366b58d8b9d07aa9fa7321a59e89032f82f512834b6ac0e34eb28

Observation 9cd1009d-abfa-4031-8c74-21b40718d71f · outbound

This paper cites Librispeech: An ASR corpus based on public domain audio books,.

Noise-Robust Target-Speaker Voice Activity Detection Through Self-Supervised Pretraining Librispeech: An ASR corpus based on public domain audio books,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T21:57:02.521885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:57:02.521885Z digest=sha256:80359e69e5f4c9bbdfaf6a5606c92a0251e08cf75b4ab389b9746558debf914c

Observation a2c503d8-5fc7-4ccc-94af-b8c9b1e1d9f9 · outbound

This paper cites Montreal Forced Aligner: Trainable text-speech alignment using kaldi,.

Noise-Robust Target-Speaker Voice Activity Detection Through Self-Supervised Pretraining Montreal Forced Aligner: Trainable text-speech alignment using kaldi,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:02.753804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:57:02.528486Z digest=sha256:375cdc1d207124eb3692cc3ffccfe356e37a050a25c17fe3cbbefce98bdcdadf

Observation b0038722-4240-46ce-aeea-366096f62462 · outbound

This paper cites Speech enhancement using long short-term memory based recurrent neural networks for noise robust speaker verification,.

Noise-Robust Target-Speaker Voice Activity Detection Through Self-Supervised Pretraining Speech enhancement using long short-term memory based recurrent neural networks for noise robust speaker verification,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T21:57:02.534690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:57:02.534690Z digest=sha256:4502d7856a9944694cc597660a05f84c2485bda03103399b8452398183ba2be1

Observation a877476e-2ed0-400f-ac3c-8164a73e86a5 · outbound

This paper cites A study on data augmentation of reverberant speech for robust speech recognition,.

Noise-Robust Target-Speaker Voice Activity Detection Through Self-Supervised Pretraining A study on data augmentation of reverberant speech for robust speech recognition,

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T21:57:02.540473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:57:02.540473Z digest=sha256:5071b77b4fb3f462b2f4edacf6b7c1982aa18da72177dd89451d8167b7d71b4e

Observation 6eab3ae7-99dd-4477-9f57-2f4b8aeda8ea · outbound

This paper cites audiomentations,.

Noise-Robust Target-Speaker Voice Activity Detection Through Self-Supervised Pretraining audiomentations,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:02.718608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:57:02.545669Z digest=sha256:657a800d3287291f78613ecca87367877d39c7f22c70a3a8946366b94682b9e9

Observation 3a36b517-3e24-4d07-b0c7-f3adb31b7768 · outbound

This paper cites Transformer-xl: Attentive language models beyond a fixed- length context,.

Noise-Robust Target-Speaker Voice Activity Detection Through Self-Supervised Pretraining Transformer-xl: Attentive language models beyond a fixed- length context,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:02.705433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:57:02.550209Z digest=sha256:fc2bfe00c604dbc0d2601c5c44a5b851910acca46a5f6ea5a29daedc656ca4c6

Observation 98366df6-51dd-4fe5-b9b1-a9bf787c0af0 · outbound

This paper cites V oxceleb: Large- scale speaker verification in the wild,.

Noise-Robust Target-Speaker Voice Activity Detection Through Self-Supervised Pretraining V oxceleb: Large- scale speaker verification in the wild,

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T21:57:02.554906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:57:02.554906Z digest=sha256:2e0f049415f0ceda9f3df67217ac1bc4c227b5cce901be15b63cb023b590e9f2

Observation 97196bad-743d-42b8-8748-deab98cc94f0 · outbound

This paper cites SGDR: Stochastic gradient descent with warm restarts,.

Noise-Robust Target-Speaker Voice Activity Detection Through Self-Supervised Pretraining SGDR: Stochastic gradient descent with warm restarts,

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T21:57:02.560104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:57:02.560104Z digest=sha256:f816eabf90e833ae0828f81b143452284166982a0dad0f755cf569bacb91dc62

Observation 1eda67bf-5213-425e-9552-17f1af2fa819 · outbound

This paper cites On batching variable size inputs for training end-to-end speech enhancement systems,.

Noise-Robust Target-Speaker Voice Activity Detection Through Self-Supervised Pretraining On batching variable size inputs for training end-to-end speech enhancement systems,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:02.676959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:57:02.564382Z digest=sha256:17fd1d07c6effb00c20e3212678dbfc26d727ca3c2e3e698faf2f9fe32381308

Observation 51e4370d-cac2-4dbc-b46d-14bd02734cc1 · outbound

This paper cites Visualizing data using t-sne,.

Noise-Robust Target-Speaker Voice Activity Detection Through Self-Supervised Pretraining Visualizing data using t-sne,

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T21:57:02.568478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:57:02.568478Z digest=sha256:1d52dbd3091e28ab3ebf0c22cd64dc68cabb389532e86b22f1a0c9b9b9262138

Pith citing papers

No inbound Pith citation observations are available.