Pith. sign in

Paper Citation Record · LEDGER

SpeechBrain: A General-Purpose Speech Toolkit

As of 5 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 48 inbound Pith citation observations for arXiv:2106.04624.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2106.04624 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 48 of 48 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T15:52:34.585187Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-07T20:34:09.733636Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 03ff4e08-5398-4fa5-8d4f-0ee3c7af1055 · inbound

DASB - Discrete Audio and Speech Benchmark cites this paper.

DASB - Discrete Audio and Speech Benchmark SpeechBrain: A General-Purpose Speech Toolkit

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-24T00:28:39.555237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T00:26:57.419537Z digest=sha256:5e98180dc56a50b09bc74da9b073a81f848e226ab1c9930412175c67d6b8ed52

Observation ab84775f-5968-4bab-96fd-a80e98d4b669 · inbound

MECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding Tasks cites this paper.

MECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding Tasks SpeechBrain: A General-Purpose Speech Toolkit

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-19T02:41:59.647937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T02:41:52.996457Z digest=sha256:ed738d22acad500db3efefe7b6016b805962446456d8fb2fd7af86ffffb0fd40

Observation 4b6c3108-644b-42f8-982b-fcde1c237a11 · inbound

Benchmarking Large Pretrained Multilingual Models on Qu\'ebec French Speech Recognition cites this paper.

Benchmarking Large Pretrained Multilingual Models on Qu\'ebec French Speech Recognition SpeechBrain: A General-Purpose Speech Toolkit

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T14:33:21.048857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:33:21.048857Z digest=sha256:931a8dcde6d879d3309f7419db38867aa98248111a22bf54a990376a1a707070

Observation fa2b14c0-3adf-407a-8f9c-8fcc845c9604 · inbound

Speaker-Conditioned Phrase Break Prediction for Text-to-Speech with Phoneme-Level Pre-trained Language Model cites this paper.

Speaker-Conditioned Phrase Break Prediction for Text-to-Speech with Phoneme-Level Pre-trained Language Model SpeechBrain: A General-Purpose Speech Toolkit

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-05T13:26:17.225948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:26:17.225948Z digest=sha256:8f6719e0a7044dff0342e823642618f89272e41936bb9843766e33a559aa0402

Observation 8795520f-d940-4ae1-8b92-aa907a3fd6ca · inbound

Analysis of Speaker Verification Performance Trade-offs with Neural Audio Codec Transmission cites this paper.

Analysis of Speaker Verification Performance Trade-offs with Neural Audio Codec Transmission SpeechBrain: A General-Purpose Speech Toolkit

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T11:29:03.289501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:29:03.289501Z digest=sha256:0360ef13e3722291fcbebefd72bc53c6295f11ce9b0a27247d2702e6c4595a32

Observation 7815c48d-833d-4958-98a9-37ff63980bb0 · inbound

DarkStream: real-time speech anonymization with low latency cites this paper.

DarkStream: real-time speech anonymization with low latency SpeechBrain: A General-Purpose Speech Toolkit

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T06:01:28.516189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T06:01:28.516189Z digest=sha256:c88e1cf4dcb6b7ddfee218fe0352cf4bb8da6cabdb5c53d22c48ec2358747f89

Observation 8ba76ccf-b79e-42c4-a5ee-9e40c665cd94 · inbound

Efficient Trie-based Biasing using K-step Prediction for Rare Word Recognition cites this paper.

Efficient Trie-based Biasing using K-step Prediction for Rare Word Recognition SpeechBrain: A General-Purpose Speech Toolkit

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T19:36:44.191622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:36:44.191622Z digest=sha256:9fd2f22ecc6f52e99f6d2ea4379a0d7e0d9f6e30d772b097232067603a01a8ac

Observation c449d607-5843-4983-9e02-a7dfbb3cc2d1 · inbound

Improving Synthetic Data Training for Contextual Biasing Models with a Keyword-Aware Cost Function cites this paper.

Improving Synthetic Data Training for Contextual Biasing Models with a Keyword-Aware Cost Function SpeechBrain: A General-Purpose Speech Toolkit

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T19:35:50.642548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:35:50.642548Z digest=sha256:e9a9b7e106ce7869bead5db0fed25c9fcaf63885349e6e2af4880f42e54b024b

Observation 1bdf1c66-9918-442e-9e88-34e8c11aa282 · inbound

FAC-FACodec: Controllable Zero-Shot Foreign Accent Conversion with Factorized Speech Codec cites this paper.

FAC-FACodec: Controllable Zero-Shot Foreign Accent Conversion with Factorized Speech Codec SpeechBrain: A General-Purpose Speech Toolkit

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-18T08:01:06.813118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T07:56:46.302918Z digest=sha256:8d0c23949c43791331225d105926c426d177bde93eaf372d2199800f3eddf319

Observation efe285bf-bfab-4787-a307-57fb14c1594f · inbound

A Study of Data Selection Strategies for Pre-training Self-Supervised Speech Models cites this paper.

A Study of Data Selection Strategies for Pre-training Self-Supervised Speech Models SpeechBrain: A General-Purpose Speech Toolkit

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:17:43.810929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T10:17:15.726726Z digest=sha256:a580e8d0945e0564769daae2ec4731b864a1550f86fafd7a2a18f2794d200dfa

Observation d4d3ea68-d443-4092-9532-321246a79cd2 · inbound

Precision-Varying Prediction (PVP): Robustifying ASR systems against adversarial attacks cites this paper.

Precision-Varying Prediction (PVP): Robustifying ASR systems against adversarial attacks SpeechBrain: A General-Purpose Speech Toolkit

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T17:37:37.865879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:37:37.865879Z digest=sha256:33e38725b00de5901e695ece00596eeba07aefafa7c33b15312db2f87830ff61

Observation 398b8ad4-9035-4187-9ac8-b4c826dd6fc3 · inbound

DeepFense: A Unified, Modular, and Extensible Framework for Robust Deepfake Audio Detection cites this paper.

DeepFense: A Unified, Modular, and Extensible Framework for Robust Deepfake Audio Detection SpeechBrain: A General-Purpose Speech Toolkit

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:45:57.571202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T17:00:05.187070Z digest=sha256:95363d74018ce54cc66a1fb52e019a5f63048c4764acf756b3413941ff65b58d

Observation a2f19927-ea26-4e36-a4bc-0ddffeaa8bfe · inbound

Contextual Biasing for ASR in Speech LLM with Common Word Cues and Bias Word Position Prediction cites this paper.

Contextual Biasing for ASR in Speech LLM with Common Word Cues and Bias Word Position Prediction SpeechBrain: A General-Purpose Speech Toolkit

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:35:33.967335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T14:35:20.515352Z digest=sha256:bb8fcac1f8f8268102c6cf4a50c26a61221cceb0d8921edf606a60c24b5515bb

Observation f7a336ad-8814-4316-b4a5-d50dc6e9a686 · inbound

SEDTalker: Emotion-Aware 3D Facial Animation Using Frame-Level Speech Emotion Diarization cites this paper.

SEDTalker: Emotion-Aware 3D Facial Animation Using Frame-Level Speech Emotion Diarization SpeechBrain: A General-Purpose Speech Toolkit

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:21:04.487867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T14:59:44.064137Z digest=sha256:24ba5f17f7cd804cafadd453a84a9a2069ce3f1a29f966df67145e720bb5d433

Observation 2210251c-a48b-4864-8bbb-3f5aa96affb6 · inbound

Hierarchical Codec Diffusion for Video-to-Speech Generation cites this paper.

Hierarchical Codec Diffusion for Video-to-Speech Generation SpeechBrain: A General-Purpose Speech Toolkit

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:12:26.275104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T08:12:01.260833Z digest=sha256:158400b672d845bbcd801980c851ca84ba399c494eec650047c3aa998cca4ac8

Observation a20f70c0-7572-414a-8a3b-02936d9d4387 · inbound

Learning Posterior Predictive Distributions for Node Classification from Synthetic Graph Priors cites this paper.

Learning Posterior Predictive Distributions for Node Classification from Synthetic Graph Priors SpeechBrain: A General-Purpose Speech Toolkit

Reference 224

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T12:21:04.729303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-10T03:50:44.626261Z digest=sha256:9b5c260c5d77cbfb527e7fdc336b4fad63ca9f33ddb20c5c4d8ce3ba9fbd9209

Observation 356dd397-1fc8-48cd-93e2-f80740ceef87 · inbound

Voice of India: A Large-Scale Benchmark for Real-World Speech Recognition in India cites this paper.

Voice of India: A Large-Scale Benchmark for Real-World Speech Recognition in India SpeechBrain: A General-Purpose Speech Toolkit

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:56:05.596946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T02:36:06.825753Z digest=sha256:fbeca5f1a64acb4cbffbd6143a9591a2fbede0a47d1e4070f51a2c7de0691339

Observation 9682e527-aec6-4734-a843-d164e549f27b · inbound

Text-To-Speech with Chain-of-Details: modeling temporal dynamics in speech generation cites this paper.

Text-To-Speech with Chain-of-Details: modeling temporal dynamics in speech generation SpeechBrain: A General-Purpose Speech Toolkit

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-10T01:04:50.145569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T01:01:06.094276Z digest=sha256:89b4a6ed4e0e648c98c6532791a81f74fbc35c0b91cbc1bf3af1409c199c9c75

Observation 19ea0026-1b6e-4764-916d-f7c5a07eac0a · inbound

Enhancing Speaker Verification with Whispered Speech via Post-Processing cites this paper.

Enhancing Speaker Verification with Whispered Speech via Post-Processing SpeechBrain: A General-Purpose Speech Toolkit

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-09T23:29:43.958668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-09T23:28:05.344284Z digest=sha256:d11ea7a72fc0884177f9d272af1a7b434685cdba026cfcb6007baf938cad5e66

Observation ae3991df-2278-4d19-9f79-fa697650c224 · inbound

A Toolkit for Detecting Spurious Correlations in Speech Datasets cites this paper.

A Toolkit for Detecting Spurious Correlations in Speech Datasets SpeechBrain: A General-Purpose Speech Toolkit

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:11:27.405133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T12:04:36.906876Z digest=sha256:6e56ba1e71bb3d97cb92621618a04c95526545dee2ad41f322489e0a204c579d

Observation f62378f7-5d2d-48fe-8df3-4a476d1e214d · inbound

HATS: An Open data set Integrating Human Perception Applied to the Evaluation of Automatic Speech Recognition Metrics cites this paper.

HATS: An Open data set Integrating Human Perception Applied to the Evaluation of Automatic Speech Recognition Metrics SpeechBrain: A General-Purpose Speech Toolkit

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:46:27.031807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T09:19:31.900807Z digest=sha256:eb3aa216c66d044d7f5565b2fb733ed313f48b371508fa98a61fd59679ea74c2

Observation 337f6e71-ec55-4aa5-877a-c2bace00e8b2 · inbound

A Paradigm for Interpreting Metrics and Identifying Critical Errors in Automatic Speech Recognition cites this paper.

A Paradigm for Interpreting Metrics and Identifying Critical Errors in Automatic Speech Recognition SpeechBrain: A General-Purpose Speech Toolkit

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:36:31.492928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:39:02.065010Z digest=sha256:539d7bbda9dbc4bbac94006afa8e6d17424589dc9c2ac2c2b27efb3382208af3

Observation 702ef76d-b5a0-4a21-9237-3a00d4fd7144 · inbound

Evaluating voice anonymisation using similarity rank disclosure cites this paper.

Evaluating voice anonymisation using similarity rank disclosure SpeechBrain: A General-Purpose Speech Toolkit

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:00:55.849557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T00:55:20.324372Z digest=sha256:87a35e03b7f52c9d0f66782c9075cf327818edda287966b124760fe6eab77503

Observation 88532c32-d578-47f6-80fb-34aa42460307 · inbound

Mechanisms of Misgeneralization in Physical Sequence Modeling cites this paper.

Mechanisms of Misgeneralization in Physical Sequence Modeling SpeechBrain: A General-Purpose Speech Toolkit

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:44:48.708253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-21T07:44:37.810511Z digest=sha256:95ab02cc536de574ae5062b8ba918410884516f68e5e1ff3704128405f550a0d

Observation b20d2510-fb91-4fed-84da-c5a819cf4ae3 · inbound

Phonetic Modeling of Dialectal Variation in Vietnamese Speech cites this paper.

Phonetic Modeling of Dialectal Variation in Vietnamese Speech SpeechBrain: A General-Purpose Speech Toolkit

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T13:34:40.334031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T13:29:59.049839Z digest=sha256:8fb618a96512a414b6f7bcac8ab5545150caa0bc3c897e1ecc933729da8bade6

Observation 6b9e549b-3f59-4f58-a5b8-cf140b364a33 · inbound

PashtoTTS-Bench: automated screening for low-resource non-Latin-script text-to-speech cites this paper.

PashtoTTS-Bench: automated screening for low-resource non-Latin-script text-to-speech SpeechBrain: A General-Purpose Speech Toolkit

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:53:51.906976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T18:44:12.111291Z digest=sha256:8aa6d5fdd8432f110a4c48e7ae243e421ddd979ac79c5e181d07cee3d8178cc5

Observation d4e9fddf-723c-4e89-b04a-2551a4a4a1dd · inbound

Syllabic-Structure Decoder for Automatic Speech Recognition in Vietnamese cites this paper.

Syllabic-Structure Decoder for Automatic Speech Recognition in Vietnamese SpeechBrain: A General-Purpose Speech Toolkit

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T13:33:28.505625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T13:24:28.184435Z digest=sha256:590c2bc4a25a68f229981ae2afe12abf2d03841bda3bb9ee9bed67b77e5c016d

Observation 3d596dbb-32e0-41bb-b76d-1342e3e3cee9 · inbound

Spiking and Event-driven Neuromorphic Mamba Models for Efficient Speech Recognition cites this paper.

Spiking and Event-driven Neuromorphic Mamba Models for Efficient Speech Recognition SpeechBrain: A General-Purpose Speech Toolkit

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-07-01T21:46:15.482506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T16:22:07.001549Z digest=sha256:df74cf06d944f0154fac439f1e455ebb1f9a4ca83648f859682e06901535c5be

Observation d6bee9e5-98fd-412b-b6d1-96bf8548f102 · inbound

KIT's Submission to Cross-Lingual Voice Cloning in IWSLT 2026 cites this paper.

KIT's Submission to Cross-Lingual Voice Cloning in IWSLT 2026 SpeechBrain: A General-Purpose Speech Toolkit

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T19:07:18.239939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-27T21:39:11.265338Z digest=sha256:f2b0489ca2efefef4ad27bcb56031899df4f3a83059f1a41125cfbeb66d8e70d

Observation e01fcc82-81d9-4ba4-82fc-1eff38a70c69 · inbound

SpeechDx: A Multi-Task Benchmark for Clinical Speech AI cites this paper.

SpeechDx: A Multi-Task Benchmark for Clinical Speech AI SpeechBrain: A General-Purpose Speech Toolkit

Reference 115

Resolution
verified exact
arxiv_id, observed 2026-07-03T18:28:48.484005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T03:09:07.295810Z digest=sha256:8333ad97c5ba53dfc02f469357c4c9733ea99cdfed560e98307f73bcb22fe701

Observation 8f53b2f8-a5e8-44c8-b12a-dd8716ec0836 · inbound

Montreal Forced Aligner and the state of speech-to-text alignment in 2026 cites this paper.

Montreal Forced Aligner and the state of speech-to-text alignment in 2026 SpeechBrain: A General-Purpose Speech Toolkit

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-07-03T21:28:59.352659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T00:27:25.550404Z digest=sha256:1556d03949b55cedbabdfc4e7df7f6011cb10f5b8329c4e0208bc7448dfd8d64

Observation 141b9598-06ad-49f2-af8d-e8f0ba02889d · inbound

Beyond ROC-AUC: Operating-Point Performance Reporting for Biometric Verification cites this paper.

Beyond ROC-AUC: Operating-Point Performance Reporting for Biometric Verification SpeechBrain: A General-Purpose Speech Toolkit

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-07-03T17:28:44.523417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T04:08:21.510476Z digest=sha256:7d7f6b3688c0242cd32ea553e2be33a75103df96db1fbf3dbaa40b053c726905

Observation 546f19b6-79d6-4f92-92af-af6e65361b18 · inbound

LISE : Listenable Interpretable Speaker Embeddings cites this paper.

LISE : Listenable Interpretable Speaker Embeddings SpeechBrain: A General-Purpose Speech Toolkit

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:39:39.082204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T13:10:03.421453Z digest=sha256:93161f44bffbb01374cac230c7c487f71d1d34771dbdf310bc6b57622b784832

Observation f657bdcd-718b-4de8-8622-20664f22c5e1 · inbound

ESPnet3: Infrastructure for Scalable Speech and Audio Research in the Foundation Model Era cites this paper.

ESPnet3: Infrastructure for Scalable Speech and Audio Research in the Foundation Model Era SpeechBrain: A General-Purpose Speech Toolkit

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:19:44.211104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T11:54:08.457573Z digest=sha256:57b556543b9c792bf9a857ab4b3e13387036240cacd967fd80dcec7d1a9f7880

Observation 15644453-f388-4335-a515-62b9e3980227 · inbound

What Counts as an Error? Dual-Reference Benchmarking for Atypical ASR cites this paper.

What Counts as an Error? Dual-Reference Benchmarking for Atypical ASR SpeechBrain: A General-Purpose Speech Toolkit

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-01T10:05:41.138932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T05:52:30.228009Z digest=sha256:9974cc5134286968581a7b9f267049adfe2b80c6c5d5c63a98dacc142926c93d

Observation 5beeb03a-ed44-42de-afd5-42e41bcd8edc · inbound

LuxEmo: Expressive Text-to-Speech Corpus for Luxembourgish cites this paper.

LuxEmo: Expressive Text-to-Speech Corpus for Luxembourgish SpeechBrain: A General-Purpose Speech Toolkit

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-07-01T10:15:45.051640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T05:38:16.224617Z digest=sha256:47bda76205d150d9d3351e3bd6c2e96924342aefea87c8d88a09bb7ca2e58676

Observation 015a7276-ef5e-4851-a4d3-0633fc1ed2f8 · inbound

LuxEmo: Expressive Text-to-Speech Corpus for Luxembourgish cites this paper.

LuxEmo: Expressive Text-to-Speech Corpus for Luxembourgish SpeechBrain: A General-Purpose Speech Toolkit

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-03T21:58:58.322839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-03T21:56:41.303269Z digest=sha256:587925c19dec983a8876dc7553c3fe089e334383d6195e3f0148784190fe30c5

Observation 448946c4-88cc-401e-9a9b-9e68b73e0e91 · inbound

From Monolingual to Multilingual: Evaluating Mamba for ASR in South African Languages cites this paper.

From Monolingual to Multilingual: Evaluating Mamba for ASR in South African Languages SpeechBrain: A General-Purpose Speech Toolkit

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:58:57.286676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-03T20:54:42.216660Z digest=sha256:ee65203d898710ec27dd49c457c258387025ac4132a58d3c893cae90dc408a6d

Observation f1a3da29-8750-4303-98f5-833a79ab54ee · inbound

Conversational Human Audio-visual Talking Dialogue Generation cites this paper.

Conversational Human Audio-visual Talking Dialogue Generation SpeechBrain: A General-Purpose Speech Toolkit

Reference 52

Resolution
unresolved
no resolver link, observed 2026-07-12T06:59:35.258176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:59:35.258176Z digest=sha256:086a9d997245d428f7a23297c079314a9c5aa5ba2a4ac8406325d75863e370f6

Observation d832546d-4020-4aff-9a35-390c916bd3ea · inbound

ProPS: Prompted Profile Synthesis for Natural Language-Conditioned Speaker Embedding Distributions cites this paper.

ProPS: Prompted Profile Synthesis for Natural Language-Conditioned Speaker Embedding Distributions SpeechBrain: A General-Purpose Speech Toolkit

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-07-07T20:34:09.742292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-07T20:30:53.546267Z digest=sha256:6b331f5ddec13f0223611d315ec87d0f436744a37e09bb724303114f05d6dea1

Observation 02c95dd8-0ba1-4995-b20c-5122bb44e192 · inbound

Gender Gap Analysis in News and Talk Online Radio Broadcast cites this paper.

Gender Gap Analysis in News and Talk Online Radio Broadcast SpeechBrain: A General-Purpose Speech Toolkit

Reference 57

Resolution
unresolved
no resolver link, observed 2026-07-14T18:13:32.966335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:13:32.966335Z digest=sha256:bddbd542021844ecee736b226baed10b32f510af3202c6afb82e31ca2c8f131d

Observation 24215b28-f5b4-4b50-b952-df5106a09a86 · inbound

AMECxSV: Adaptive Metadata-Driven Embedding-Fusion Calibration for X-Lingual Speaker Verification cites this paper.

AMECxSV: Adaptive Metadata-Driven Embedding-Fusion Calibration for X-Lingual Speaker Verification SpeechBrain: A General-Purpose Speech Toolkit

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-01T20:47:01.794889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T20:47:01.794889Z digest=sha256:bf89c4f0d77ffa96f55f9219d9482f6ca68fbb5a805866d5a2ec8d8be5fa0ecb

Observation 836c1f18-2ec3-4ce6-b6b3-f52efceae5e4 · inbound

From Read Speech to Spoken Digits: A Task-Specific Evaluation of Speech Privacy With Informed Attackers cites this paper.

From Read Speech to Spoken Digits: A Task-Specific Evaluation of Speech Privacy With Informed Attackers SpeechBrain: A General-Purpose Speech Toolkit

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T07:37:18.250178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:37:18.250178Z digest=sha256:18433e278f3ab57096df963bd777f7cf67ecb1da5fd44d9ddc4bc4c1a14f28d8

Observation 43581c37-f2a4-4509-babf-e337a3a127d3 · inbound

Leveraging Gradient Reversal Loss and Multitask Learning for Datasets-Aware Audio Deepfake Detection cites this paper.

Leveraging Gradient Reversal Loss and Multitask Learning for Datasets-Aware Audio Deepfake Detection SpeechBrain: A General-Purpose Speech Toolkit

Reference 70

Resolution
unresolved
no resolver link, observed 2026-07-31T23:27:05.730485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:27:05.730485Z digest=sha256:d6d910e22f0352318e8af5e613206c957cab8537ad927aefcf42ab12a4c809b8

Observation d5a6f97a-d528-47aa-9b47-c52229f6fb69 · inbound

Face and Voice Cross-modal Association with Learning Convex Feature Embedding cites this paper.

Face and Voice Cross-modal Association with Learning Convex Feature Embedding SpeechBrain: A General-Purpose Speech Toolkit

Reference 74

Resolution
unresolved
no resolver link, observed 2026-07-31T16:58:29.335057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:58:29.335057Z digest=sha256:b77955ed4943a52e7ef4fa7758c63ad79444aa3134ce8c3178b73f8d93d32e6c

Observation 21b7b645-460b-4af0-bb54-72671c7b5029 · inbound

Leveraging Beam Search Information for Confidence Estimation in E2E ASR cites this paper.

Leveraging Beam Search Information for Confidence Estimation in E2E ASR SpeechBrain: A General-Purpose Speech Toolkit

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T09:46:28.964740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:46:28.964740Z digest=sha256:6e1b9072df9f1994ce030a15a4cbbd99944584228dda556bd303d040b968b953

Observation b7040812-3e7f-4077-a96f-54d202f54a14 · inbound

Multi-Backbone Self-Supervised Ensembles for Audio Deepfake Detection and a Cross-Track Analysis of Generation-Detection Asymmetry cites this paper.

Multi-Backbone Self-Supervised Ensembles for Audio Deepfake Detection and a Cross-Track Analysis of Generation-Detection Asymmetry SpeechBrain: A General-Purpose Speech Toolkit

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T20:44:56.274017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:44:56.274017Z digest=sha256:5eeea05991c18169e195cad65fe6f711f0efa2ea50b7ced388831d22fe3fe822

Observation bcaad713-9f58-4df8-a603-773df50a24d9 · inbound

Speaker Verification Under Real Classroom Conditions for English Speech cites this paper.

Speaker Verification Under Real Classroom Conditions for English Speech SpeechBrain: A General-Purpose Speech Toolkit

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T15:52:34.585187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:52:34.585187Z digest=sha256:e73bb572e433088e2d53869e025f909f70af5810ff495556a3b2d9f1ea5a74c0