Pith. sign in

Paper Citation Record · LEDGER

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition

As of 9 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 1 inbound Pith citation observation for arXiv:2506.12672.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.12672 v1

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:50:17.272359Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:50:17.073394Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T00:50:17.584368Z

Reference resolution

40 of 40 outbound references displayed

  • verified exact7
  • verified fuzzy15
  • unresolved17
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1ad7751f-54f4-4b45-964d-2f03033599d5 · outbound

This paper cites SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:50:17.589525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:50:17.073394Z digest=sha256:4013a85743b327d4fe6d4bdd03f2167b6319f7d6cec0ed6fd6d8daba5ab33364

Observation 405de8c9-4626-4240-a43a-a36fe2cf999c · outbound

This paper cites an unresolved cited work.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:50:17.942252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:50:17.078666Z digest=sha256:212f63d27875f2d5d956151da6289c653de1d3ba1ac1e0b0801064a11bf9d622

Observation cb0e5e3b-925e-4fc2-b02a-54fdf9ad3b7b · outbound

This paper cites an unresolved cited work.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:50:17.924431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:50:17.083516Z digest=sha256:ef25591a00cd145c8a843413ac7d3421d881b520229e1b5941812b96884a04d7

Observation 7807dc91-3110-48ba-93a3-5fe6d82d29ee · outbound

This paper cites Dataset The training set consists of LibriSpeech-360h [27], Libri2Mix, and Libri3Mix [28].

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Dataset The training set consists of LibriSpeech-360h [27], Libri2Mix, and Libri3Mix [28]

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:17.909059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:50:17.088616Z digest=sha256:cc637c3f7235ebb2a6437a7c880d359ac3a1ffce90ca487c1a2c1bf8daf3b1c1

Observation 119c431a-5298-420b-bf3a-0621ff9ce989 · outbound

This paper cites Our proposed SC-SOT models, particularly when combined with multi-task learning (MTL), demonstrate promis- ing performance gains over the conventional SOT-based model.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Our proposed SC-SOT models, particularly when combined with multi-task learning (MTL), demonstrate promis- ing performance gains over the conventional SOT-based model

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:17.876339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:50:17.099198Z digest=sha256:4b538e81030819271bce60983222098bb55cf54ec13bf19a46c95a2cbf97a820

Observation be4c3326-ae81-42a4-97e6-6d78eb03ccf4 · outbound

This paper cites Our initial analy- sis, through visualization of source-target attention weights in an SOT-based model, revealed that the decoder performs a de- gree of speaker separation.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Our initial analy- sis, through visualization of source-target attention weights in an SOT-based model, revealed that the decoder performs a de- gree of speaker separation

Reference 6

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T00:50:17.859960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:50:17.104148Z digest=sha256:c3ff25e78916e31887d20bdc53fe8ca298a96e188c9eb158af4c141ef563ae1e

Observation f73ff6b2-0cf5-4938-bc88-b1d7c2627bd6 · outbound

This paper cites an unresolved cited work.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:17.110271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:17.110271Z digest=sha256:d8bd42c1241c50f80567f52181bc972818cc1db80e9869e531da14ab143e9ed2

Observation 4af809b9-9acb-46bb-8a38-6b1edd2a03d4 · outbound

This paper cites Sequence to multi-sequence learning via conditional chain mapping for mixture signals,.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Sequence to multi-sequence learning via conditional chain mapping for mixture signals,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:17.800507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:50:17.151926Z digest=sha256:aeb561d45a783ef7cb8cfe5a3fd9efa947d7fea54ff2ad0ce57449a3fcc0a8ab

Observation 8814306e-e349-45ff-9567-2150e2b26cc3 · outbound

This paper cites Recent advances in end-to-end automatic speech recognition,.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Recent advances in end-to-end automatic speech recognition,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:17.115011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:17.115011Z digest=sha256:b74e6ec24033488778e39a5835e8d29bb3a668a5e1e16e0b4c8c9e21041b8b6a

Observation f515b9f0-5c58-44cc-9342-61b398a66887 · outbound

This paper cites End-to-end speech recognition: A survey,.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition End-to-end speech recognition: A survey,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:17.824737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:50:17.120316Z digest=sha256:af354004241ddf7d515cd21c3dca171af45e94f510351851fb7146b5375d00b0

Observation 951bc0d5-095d-4279-ae7a-28ca126fdbea · outbound

This paper cites Robust speech recognition via large-scale weak supervision,.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Robust speech recognition via large-scale weak supervision,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:17.125516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:17.125516Z digest=sha256:446a9bcbe3d40678f1bb2db6a19ff87f285eee714e9b6016dd5f40e8c56a3be9

Observation a4ed4b4e-fe46-4103-a45f-71342bddb5d0 · outbound

This paper cites Anatomy of Industrial Scale Multilingual ASR.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Anatomy of Industrial Scale Multilingual ASR

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:17.130311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:17.130311Z digest=sha256:0f34c80b7c088cbd89d825d944965eac5d47d30c1997a3e968711bbfd54d9f4f

Observation 86e970ca-86a7-4fdb-8db1-89a9f125a60d · outbound

This paper cites Less is More: Accurate Speech Recognition & Translation without Web-Scale Data.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Less is More: Accurate Speech Recognition & Translation without Web-Scale Data

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:17.135316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:17.135316Z digest=sha256:1ce5dd0983dedd6544c5e3b51e943c97489f6f833fb87325b3355dbffcf48286

Observation ed548814-085e-400f-9406-86d11342b4be · outbound

This paper cites Recognizing Multi-talker Speech with Permutation Invariant Training.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Recognizing Multi-talker Speech with Permutation Invariant Training

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:50:17.529996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:50:17.141353Z digest=sha256:f76682798f93c64d50b09f2578b03d31009309449cc9c0e22f21e21e3f0669af

Observation d14078c9-7074-4944-a84a-b28828a3ee89 · outbound

This paper cites Serialized Output Training for End-to-End Overlapped Speech Recognition.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Serialized Output Training for End-to-End Overlapped Speech Recognition

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:17.146989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:17.146989Z digest=sha256:ed9b7ca18d66566cdc1832c063a615f4a44b5dbf7dacff41fb292f9c3861cea4

Observation ba27bf0b-135b-4fbe-bee4-cf7f94d8c084 · outbound

This paper cites Speakerbeam: Speaker aware neural network for target speaker extraction in speech mixtures,.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Speakerbeam: Speaker aware neural network for target speaker extraction in speech mixtures,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:17.727622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:50:17.191684Z digest=sha256:6210eb764f1590a2f571013ae8255dd550c7d77d1ef8ab5673f35be45950fcd7

Observation 4ba06cd0-8d43-483b-b401-36c1d57ef1ef · outbound

This paper cites Streaming Multi-Talker ASR with Token-Level Serialized Output Training.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Streaming Multi-Talker ASR with Token-Level Serialized Output Training

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:50:17.490013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:50:17.156569Z digest=sha256:2e506aa2652fc35daf59feb58e908b407678a4ea9156011b7096e8c92a1a6ae8

Observation 4376e618-bfce-4786-8f37-1467ab96e3f7 · outbound

This paper cites Surt 2.0: Advances in transducer-based multi-talker speech recognition,.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Surt 2.0: Advances in transducer-based multi-talker speech recognition,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:17.784366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:50:17.161468Z digest=sha256:cd60d2585a91d8e55dea2b4daba15e3389d7423465eee921355f0759dcfb6f32

Observation d553db92-cce3-4a02-8e06-0f4d2aa68ca7 · outbound

This paper cites Permutation invari- ant training of deep models for speaker-independent multi-talker speech separation,.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Permutation invari- ant training of deep models for speaker-independent multi-talker speech separation,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:17.768604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:50:17.166023Z digest=sha256:6e541738ca2e13a2560207dfded6d1e5fb0c9776f369d5fc7747aaee91c873d1

Observation 2952a336-5482-45c4-bda7-03457cb27a32 · outbound

This paper cites A Purely End-to-end System for Multi-speaker Speech Recognition.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition A Purely End-to-end System for Multi-speaker Speech Recognition

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:50:17.464512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:50:17.170341Z digest=sha256:8788212e84f2248dc74208db567766a1fdea513623a5fad28076f13b2d6f098c

Observation ecf7c578-90f1-41d9-8383-166d59ebc84f · outbound

This paper cites Mimo-speech: End-to-end multi-channel multi-speaker speech recognition,.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Mimo-speech: End-to-end multi-channel multi-speaker speech recognition,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:17.752648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:50:17.176328Z digest=sha256:056b6f68556386cbdc1fbb5eaed6cecc6c00fe6b3520281fff29251208dfb181

Observation e6c902b3-0659-48a7-a629-effaa25cc403 · outbound

This paper cites End-to-End Speaker Diarization for an Unknown Number of Speakers with Encoder-Decoder Based Attractors.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition End-to-End Speaker Diarization for an Unknown Number of Speakers with Encoder-Decoder Based Attractors

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:50:17.443388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:50:17.181573Z digest=sha256:6d1cf1ff902362ad507a8ab9f184b4242b6d9d5c6bdc0b4e5793f27d41b6f030

Observation 4017e837-3f20-4e2f-bc3e-58ec3e399815 · outbound

This paper cites Neural target speech extraction: An overview,.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Neural target speech extraction: An overview,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:17.186901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:17.186901Z digest=sha256:5c0b2bdeecee700e2facfb46352e0ca39f47f4e5e89117b4988f98767c795cec

Observation 380de5a2-eb45-4886-8511-bc397d5cc1be · outbound

This paper cites Transcribe-to-diarize: Neural speaker diariza- tion for unlimited number of speakers using end-to-end speaker- attributed asr,.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Transcribe-to-diarize: Neural speaker diariza- tion for unlimited number of speakers using end-to-end speaker- attributed asr,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:17.657506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:50:17.232416Z digest=sha256:a44992e707d32c6621f92784b1f14f9b30537faa7861f7fa47878f7311b45909

Observation 3c19fef3-ff92-49fa-a295-6b1fcd445f7a · outbound

This paper cites Adaptive blind audio source extraction supervised by dominant speaker identification using x-vectors,.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Adaptive blind audio source extraction supervised by dominant speaker identification using x-vectors,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:17.711415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:50:17.196306Z digest=sha256:cd6eef0753c87185db5d382584713f94dbed28e4edd538e509d11d433bdbebc6

Observation 60b19474-c3ea-4af3-92ef-e7ba9ad5b911 · outbound

This paper cites Target-Speaker Voice Activity Detection: a Novel Approach for Multi-Speaker Diarization in a Dinner Party Scenario.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Target-Speaker Voice Activity Detection: a Novel Approach for Multi-Speaker Diarization in a Dinner Party Scenario

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:50:17.422697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:50:17.200864Z digest=sha256:abab03040862a02219d8bc81432dee11b1a023a4b04060704f53217aacfea9c8

Observation 866845d2-be7c-46f1-ad3c-b8423cd743d0 · outbound

This paper cites Tar- get speaker voice activity detection with transformers and its in- tegration with end-to-end neural diarization,.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Tar- get speaker voice activity detection with transformers and its in- tegration with end-to-end neural diarization,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:17.696412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:50:17.206891Z digest=sha256:4411926761f053f05c24eaf3c31bb0dd6dd58337ab3ae985896ab19022b208e0

Observation 7715e151-8ace-474f-a08c-23f071d90dab · outbound

This paper cites Adapting self-supervised models to multi-talker speech recognition using speaker embeddings,.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Adapting self-supervised models to multi-talker speech recognition using speaker embeddings,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:17.682041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:50:17.211578Z digest=sha256:15aae46d030bd41ec35fe7be6ba0ad8e13d7822af5ec920bc0209c6eadddb9fc

Observation 35cb56ab-2885-4e35-916f-a01a5f63ea7d · outbound

This paper cites The Conformer encoder has 12 layers, while the Transformer decoder is structured with 6 layers.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition The Conformer encoder has 12 layers, while the Transformer decoder is structured with 6 layers

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:17.892600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:50:17.093997Z digest=sha256:6137610582dfe7973295cf31cdc5c59c1d1e4519e124af5360042ec193f4c5ff

Observation 09c22c77-82f4-48cd-a98f-8a49be71a03b · outbound

This paper cites Con- nectionist temporal classification: labelling unsegmented se- quence data with recurrent neural networks,.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Con- nectionist temporal classification: labelling unsegmented se- quence data with recurrent neural networks,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:17.216547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:17.216547Z digest=sha256:76032a7309491b6e2590b0626c79b0d492a8e9136330eba6fc4a68b805c23e5b

Observation 1154dfe7-8571-4985-8f49-c0b593f47eeb · outbound

This paper cites Target Speaker ASR with Whisper.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Target Speaker ASR with Whisper

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:50:17.399004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:50:17.222508Z digest=sha256:a531a8551dac68f111658b40f0b466a31a2c67ff1d4322b45cdd65c4b46baa7b

Observation 33eeac7b-ef79-4c00-9143-2aa993c6071f · outbound

This paper cites DiCoW: Diarization-Conditioned Whisper for Target Speaker Automatic Speech Recognition.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition DiCoW: Diarization-Conditioned Whisper for Target Speaker Automatic Speech Recognition

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:17.227570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:17.227570Z digest=sha256:4a297951eea4e42b4c95663a1860f6e0fb42ba6c5ed74241a874b204ce351389

Observation 61ec0b16-5040-4d03-b1d6-95f7c08db808 · outbound

This paper cites But/jhu system descrip- tion for chime-8 notsofar-1 challenge,.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition But/jhu system descrip- tion for chime-8 notsofar-1 challenge,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:17.640268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:50:17.237094Z digest=sha256:2e7ad26c38d608566509c4a5ed5163e085eb97896f3f43dabea7faa33c33ec8a

Observation 068f07fe-08e1-42f7-8d45-08c3a86ec4d0 · outbound

This paper cites ESPnet: End-to-End Speech Processing Toolkit.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition ESPnet: End-to-End Speech Processing Toolkit

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:17.242039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:17.242039Z digest=sha256:74fe473aa62b52459ab27c322d77057140509c6cc16f8225594a275ebf91c300

Observation 4ddde7f5-7f1b-4edd-a0be-6f94237742de · outbound

This paper cites Lib- rispeech: an asr corpus based on public domain audio books,.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Lib- rispeech: an asr corpus based on public domain audio books,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:17.625398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:50:17.246975Z digest=sha256:c51c3488a9e01b52f77b8542ab41579c00a26eb4c4e2eb7911577967c874b988

Observation 0123293d-e901-4c82-8787-7cce00d2debe · outbound

This paper cites LibriMix: An Open-Source Dataset for Generalizable Speech Separation.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition LibriMix: An Open-Source Dataset for Generalizable Speech Separation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:17.251656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:17.251656Z digest=sha256:517667c22d6cc87babe10d9e883e86008a6184110094dc913ebf468271f40f3a

Observation 0fa73f01-3f40-4b90-a87c-e5527b34ac03 · outbound

This paper cites Conformer: Convolution-augmented Transformer for Speech Recognition.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:17.256472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:17.256472Z digest=sha256:8eb6e431a4a6fbf3f8ef0a228653c35ddc0d639004b61ed72d1a9d2e1224cc68

Observation b39777e2-7e61-493e-9d59-8ffc4ff45661 · outbound

This paper cites Attention is all you need,.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Attention is all you need,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:17.262241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:17.262241Z digest=sha256:a52ed28ae70cf2921af10414272883c06ee7b45d0231eda1b64fee3d7055bc5d

Observation 38baa58f-d614-4712-8dd3-d8bce09ac4e2 · outbound

This paper cites Wavlm: Large-scale self- supervised pre-training for full stack speech processing,.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Wavlm: Large-scale self- supervised pre-training for full stack speech processing,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:17.267393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:17.267393Z digest=sha256:e34a7ea60def5714f36e88a6e335f563af5d8ed2b205f600e0da201b7be3f0bd

Observation 63a2af0e-2c8b-4091-8761-28d663944392 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Adam: A Method for Stochastic Optimization

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:17.272359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:17.272359Z digest=sha256:0096e851bf3b2b48c4fb569eff46243e889e66fd02b5a7bb75946607f994eeae

Pith citing papers

Observation 1ad7751f-54f4-4b45-964d-2f03033599d5 · inbound

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition cites this paper.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:50:17.589525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:50:17.073394Z digest=sha256:4013a85743b327d4fe6d4bdd03f2167b6319f7d6cec0ed6fd6d8daba5ab33364