Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T00:50:17.272359Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 1 inbound Pith citation observation for arXiv:2506.12672.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T00:50:17.272359Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T00:50:17.073394Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-07T00:50:17.584368Z
40 of 40 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 1ad7751f-54f4-4b45-964d-2f03033599d5 · outbound
SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 405de8c9-4626-4240-a43a-a36fe2cf999c · outbound
SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Unresolved cited work
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation cb0e5e3b-925e-4fc2-b02a-54fdf9ad3b7b · outbound
SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Unresolved cited work
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7807dc91-3110-48ba-93a3-5fe6d82d29ee · outbound
SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Dataset The training set consists of LibriSpeech-360h [27], Libri2Mix, and Libri3Mix [28]
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 119c431a-5298-420b-bf3a-0621ff9ce989 · outbound
SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Our proposed SC-SOT models, particularly when combined with multi-task learning (MTL), demonstrate promis- ing performance gains over the conventional SOT-based model
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation be4c3326-ae81-42a4-97e6-6d78eb03ccf4 · outbound
SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Our initial analy- sis, through visualization of source-target attention weights in an SOT-based model, revealed that the decoder performs a de- gree of speaker separation
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f73ff6b2-0cf5-4938-bc88-b1d7c2627bd6 · outbound
SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Unresolved cited work
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4af809b9-9acb-46bb-8a38-6b1edd2a03d4 · outbound
SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Sequence to multi-sequence learning via conditional chain mapping for mixture signals,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8814306e-e349-45ff-9567-2150e2b26cc3 · outbound
SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Recent advances in end-to-end automatic speech recognition,
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f515b9f0-5c58-44cc-9342-61b398a66887 · outbound
SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition End-to-end speech recognition: A survey,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 951bc0d5-095d-4279-ae7a-28ca126fdbea · outbound
SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Robust speech recognition via large-scale weak supervision,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4ed4b4e-fe46-4103-a45f-71342bddb5d0 · outbound
SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Anatomy of Industrial Scale Multilingual ASR
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86e970ca-86a7-4fdb-8db1-89a9f125a60d · outbound
SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Less is More: Accurate Speech Recognition & Translation without Web-Scale Data
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed548814-085e-400f-9406-86d11342b4be · outbound
SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Recognizing Multi-talker Speech with Permutation Invariant Training
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d14078c9-7074-4944-a84a-b28828a3ee89 · outbound
SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Serialized Output Training for End-to-End Overlapped Speech Recognition
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba27bf0b-135b-4fbe-bee4-cf7f94d8c084 · outbound
SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Speakerbeam: Speaker aware neural network for target speaker extraction in speech mixtures,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4ba06cd0-8d43-483b-b401-36c1d57ef1ef · outbound
SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Streaming Multi-Talker ASR with Token-Level Serialized Output Training
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4376e618-bfce-4786-8f37-1467ab96e3f7 · outbound
SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Surt 2.0: Advances in transducer-based multi-talker speech recognition,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d553db92-cce3-4a02-8e06-0f4d2aa68ca7 · outbound
SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Permutation invari- ant training of deep models for speaker-independent multi-talker speech separation,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 2952a336-5482-45c4-bda7-03457cb27a32 · outbound
SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition A Purely End-to-end System for Multi-speaker Speech Recognition
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ecf7c578-90f1-41d9-8383-166d59ebc84f · outbound
SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Mimo-speech: End-to-end multi-channel multi-speaker speech recognition,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e6c902b3-0659-48a7-a629-effaa25cc403 · outbound
SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition End-to-End Speaker Diarization for an Unknown Number of Speakers with Encoder-Decoder Based Attractors
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4017e837-3f20-4e2f-bc3e-58ec3e399815 · outbound
SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Neural target speech extraction: An overview,
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 380de5a2-eb45-4886-8511-bc397d5cc1be · outbound
SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Transcribe-to-diarize: Neural speaker diariza- tion for unlimited number of speakers using end-to-end speaker- attributed asr,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3c19fef3-ff92-49fa-a295-6b1fcd445f7a · outbound
SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Adaptive blind audio source extraction supervised by dominant speaker identification using x-vectors,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 60b19474-c3ea-4af3-92ef-e7ba9ad5b911 · outbound
SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Target-Speaker Voice Activity Detection: a Novel Approach for Multi-Speaker Diarization in a Dinner Party Scenario
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 866845d2-be7c-46f1-ad3c-b8423cd743d0 · outbound
SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Tar- get speaker voice activity detection with transformers and its in- tegration with end-to-end neural diarization,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7715e151-8ace-474f-a08c-23f071d90dab · outbound
SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Adapting self-supervised models to multi-talker speech recognition using speaker embeddings,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 35cb56ab-2885-4e35-916f-a01a5f63ea7d · outbound
SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition The Conformer encoder has 12 layers, while the Transformer decoder is structured with 6 layers
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 09c22c77-82f4-48cd-a98f-8a49be71a03b · outbound
SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Con- nectionist temporal classification: labelling unsegmented se- quence data with recurrent neural networks,
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1154dfe7-8571-4985-8f49-c0b593f47eeb · outbound
SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Target Speaker ASR with Whisper
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 33eeac7b-ef79-4c00-9143-2aa993c6071f · outbound
SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition DiCoW: Diarization-Conditioned Whisper for Target Speaker Automatic Speech Recognition
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61ec0b16-5040-4d03-b1d6-95f7c08db808 · outbound
SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition But/jhu system descrip- tion for chime-8 notsofar-1 challenge,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 068f07fe-08e1-42f7-8d45-08c3a86ec4d0 · outbound
SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition ESPnet: End-to-End Speech Processing Toolkit
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ddde7f5-7f1b-4edd-a0be-6f94237742de · outbound
SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Lib- rispeech: an asr corpus based on public domain audio books,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 0123293d-e901-4c82-8787-7cce00d2debe · outbound
SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition LibriMix: An Open-Source Dataset for Generalizable Speech Separation
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0fa73f01-3f40-4b90-a87c-e5527b34ac03 · outbound
SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Conformer: Convolution-augmented Transformer for Speech Recognition
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b39777e2-7e61-493e-9d59-8ffc4ff45661 · outbound
SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Attention is all you need,
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38baa58f-d614-4712-8dd3-d8bce09ac4e2 · outbound
SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Wavlm: Large-scale self- supervised pre-training for full stack speech processing,
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63a2af0e-2c8b-4091-8761-28d663944392 · outbound
SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Adam: A Method for Stochastic Optimization
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ad7751f-54f4-4b45-964d-2f03033599d5 · inbound
SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.