Pith. sign in

Paper Citation Record · LEDGER

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition

As of 9 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 1 inbound Pith citation observation for arXiv:2506.12672.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.12672 v1

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:50:17.272359Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:50:17.073394Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T00:50:17.584368Z

Reference resolution

40 of 40 outbound references displayed

  • verified exact7
  • verified fuzzy15
  • unresolved17
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1ad7751f-54f4-4b45-964d-2f03033599d5 · outbound

This paper cites SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:50:17.589525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:50:17.073394Z digest=sha256:a6e583f6691418f28666675983c159de5aebc45ead2e952e0e51b8fe9fd3bbab

Observation 405de8c9-4626-4240-a43a-a36fe2cf999c · outbound

This paper cites an unresolved cited work.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:50:17.942252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:50:17.078666Z digest=sha256:86e14d8ab2a7465c8558a02b11f8c1b94e2d1bb8e5dfa7b224bbf68495cf77c6

Observation cb0e5e3b-925e-4fc2-b02a-54fdf9ad3b7b · outbound

This paper cites an unresolved cited work.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:50:17.924431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:50:17.083516Z digest=sha256:6036f25a75360eb9356e47ce1f73a7468dfdbb4174c9bad046beb51aa0e9c452

Observation 7807dc91-3110-48ba-93a3-5fe6d82d29ee · outbound

This paper cites Dataset The training set consists of LibriSpeech-360h [27], Libri2Mix, and Libri3Mix [28].

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Dataset The training set consists of LibriSpeech-360h [27], Libri2Mix, and Libri3Mix [28]

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:17.909059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:50:17.088616Z digest=sha256:45ee8fec7568fb03a304ee261eaaae96e9cc6c1dbb484712628d3cc1ac5a92fd

Observation 119c431a-5298-420b-bf3a-0621ff9ce989 · outbound

This paper cites Our proposed SC-SOT models, particularly when combined with multi-task learning (MTL), demonstrate promis- ing performance gains over the conventional SOT-based model.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Our proposed SC-SOT models, particularly when combined with multi-task learning (MTL), demonstrate promis- ing performance gains over the conventional SOT-based model

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:17.876339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:50:17.099198Z digest=sha256:b37fad34f4f2dca869c3973761066c96b17a25ba92a1a3324cc7a3bd991fee03

Observation be4c3326-ae81-42a4-97e6-6d78eb03ccf4 · outbound

This paper cites Our initial analy- sis, through visualization of source-target attention weights in an SOT-based model, revealed that the decoder performs a de- gree of speaker separation.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Our initial analy- sis, through visualization of source-target attention weights in an SOT-based model, revealed that the decoder performs a de- gree of speaker separation

Reference 6

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T00:50:17.859960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:50:17.104148Z digest=sha256:06b47ad2a5d191f796a60f1fbb67c23a4e24e7645e59d5526a3d1cec3be5e4e9

Observation f73ff6b2-0cf5-4938-bc88-b1d7c2627bd6 · outbound

This paper cites an unresolved cited work.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:17.110271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:17.110271Z digest=sha256:1c98183ee5c05ab2fdc714bbc9f726727b347cc18f3edfe5437bf6d575213d1d

Observation 4af809b9-9acb-46bb-8a38-6b1edd2a03d4 · outbound

This paper cites Sequence to multi-sequence learning via conditional chain mapping for mixture signals,.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Sequence to multi-sequence learning via conditional chain mapping for mixture signals,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:17.800507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:50:17.151926Z digest=sha256:8009b12f7c6f6576a75d51b06fb174f88a0375da07c512e3288a891f53a978a3

Observation 8814306e-e349-45ff-9567-2150e2b26cc3 · outbound

This paper cites Recent advances in end-to-end automatic speech recognition,.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Recent advances in end-to-end automatic speech recognition,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:17.115011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:17.115011Z digest=sha256:5941bde21d84857ea672cc9f5bc6313df6f5e249fd5d70f50739c803d0aeec4f

Observation f515b9f0-5c58-44cc-9342-61b398a66887 · outbound

This paper cites End-to-end speech recognition: A survey,.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition End-to-end speech recognition: A survey,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:17.824737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:50:17.120316Z digest=sha256:9855e04ee2d3e8cab03019bd3e043c99f755eeca0fb01571523cc8a9dbdad2e4

Observation 951bc0d5-095d-4279-ae7a-28ca126fdbea · outbound

This paper cites Robust speech recognition via large-scale weak supervision,.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Robust speech recognition via large-scale weak supervision,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:17.125516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:17.125516Z digest=sha256:636495cd69452b421d75a95dbf8fa705f539ac020d95bf68708859369c8ab855

Observation a4ed4b4e-fe46-4103-a45f-71342bddb5d0 · outbound

This paper cites Anatomy of Industrial Scale Multilingual ASR.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Anatomy of Industrial Scale Multilingual ASR

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:17.130311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:17.130311Z digest=sha256:157713072ae1b5ee3a44ad12ae72b9b1d9495eedde5bf2be34822196d646d5ab

Observation 86e970ca-86a7-4fdb-8db1-89a9f125a60d · outbound

This paper cites Less is More: Accurate Speech Recognition & Translation without Web-Scale Data.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Less is More: Accurate Speech Recognition & Translation without Web-Scale Data

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:17.135316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:17.135316Z digest=sha256:b0aa1431ec84cd7bf84b1b7e328efdaa3eb39589d5bb606ca59e3fe1d19bed48

Observation ed548814-085e-400f-9406-86d11342b4be · outbound

This paper cites Recognizing Multi-talker Speech with Permutation Invariant Training.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Recognizing Multi-talker Speech with Permutation Invariant Training

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:50:17.529996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:50:17.141353Z digest=sha256:d155c2138e951966801adcac985a0a2abdb2b0ec0f15c92cdbe8e02c1fe1fa87

Observation d14078c9-7074-4944-a84a-b28828a3ee89 · outbound

This paper cites Serialized Output Training for End-to-End Overlapped Speech Recognition.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Serialized Output Training for End-to-End Overlapped Speech Recognition

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:17.146989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:17.146989Z digest=sha256:cefe157cd3b588f30738bbc7b62289a6673e21199b007294422e4ea1e7d1afd9

Observation ba27bf0b-135b-4fbe-bee4-cf7f94d8c084 · outbound

This paper cites Speakerbeam: Speaker aware neural network for target speaker extraction in speech mixtures,.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Speakerbeam: Speaker aware neural network for target speaker extraction in speech mixtures,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:17.727622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:50:17.191684Z digest=sha256:df2637b1d2bf23c578b19b70aad9a42500c78261885b4babd6c053e909f7d0eb

Observation 4ba06cd0-8d43-483b-b401-36c1d57ef1ef · outbound

This paper cites Streaming Multi-Talker ASR with Token-Level Serialized Output Training.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Streaming Multi-Talker ASR with Token-Level Serialized Output Training

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:50:17.490013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:50:17.156569Z digest=sha256:5e02d69277b711b6568a3cc6c875ab3a586e6511edbd5047c1e771e62905f207

Observation 4376e618-bfce-4786-8f37-1467ab96e3f7 · outbound

This paper cites Surt 2.0: Advances in transducer-based multi-talker speech recognition,.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Surt 2.0: Advances in transducer-based multi-talker speech recognition,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:17.784366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:50:17.161468Z digest=sha256:37c70f52de8bea7bcefa64da857356b9d75e1a864c567e877f67d7a6ca179341

Observation d553db92-cce3-4a02-8e06-0f4d2aa68ca7 · outbound

This paper cites Permutation invari- ant training of deep models for speaker-independent multi-talker speech separation,.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Permutation invari- ant training of deep models for speaker-independent multi-talker speech separation,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:17.768604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:50:17.166023Z digest=sha256:b2db78c4d9d583342c5f02894c2559734ece5cbcd554fbd42a1dc7b54ae90eaf

Observation 2952a336-5482-45c4-bda7-03457cb27a32 · outbound

This paper cites A Purely End-to-end System for Multi-speaker Speech Recognition.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition A Purely End-to-end System for Multi-speaker Speech Recognition

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:50:17.464512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:50:17.170341Z digest=sha256:439e13087afb43841774147a232cb35f12af28f2d0a847e6613648b621a7d679

Observation ecf7c578-90f1-41d9-8383-166d59ebc84f · outbound

This paper cites Mimo-speech: End-to-end multi-channel multi-speaker speech recognition,.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Mimo-speech: End-to-end multi-channel multi-speaker speech recognition,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:17.752648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:50:17.176328Z digest=sha256:d551a468542d41cf67dd5be9e3ba64ff662eb9ac410312c6a13ee8b7f96faa00

Observation e6c902b3-0659-48a7-a629-effaa25cc403 · outbound

This paper cites End-to-End Speaker Diarization for an Unknown Number of Speakers with Encoder-Decoder Based Attractors.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition End-to-End Speaker Diarization for an Unknown Number of Speakers with Encoder-Decoder Based Attractors

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:50:17.443388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:50:17.181573Z digest=sha256:6a077d11e371831d25849fe20a7d32d4b32a49afd56ce855602068ab3d245eb6

Observation 4017e837-3f20-4e2f-bc3e-58ec3e399815 · outbound

This paper cites Neural target speech extraction: An overview,.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Neural target speech extraction: An overview,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:17.186901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:17.186901Z digest=sha256:3611fff6fe338b0af613da57f95453c7fdee50f1fe1ea07f068ede75fc5b7f04

Observation 380de5a2-eb45-4886-8511-bc397d5cc1be · outbound

This paper cites Transcribe-to-diarize: Neural speaker diariza- tion for unlimited number of speakers using end-to-end speaker- attributed asr,.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Transcribe-to-diarize: Neural speaker diariza- tion for unlimited number of speakers using end-to-end speaker- attributed asr,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:17.657506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:50:17.232416Z digest=sha256:c456e39607b0738a3ff09eb3820695ced51d69c0633b5a517285aa2663ee57d4

Observation 3c19fef3-ff92-49fa-a295-6b1fcd445f7a · outbound

This paper cites Adaptive blind audio source extraction supervised by dominant speaker identification using x-vectors,.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Adaptive blind audio source extraction supervised by dominant speaker identification using x-vectors,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:17.711415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:50:17.196306Z digest=sha256:aaa1bdb84b2784cf77d83b3d9ea528b5cd098da4a03eb0cd91c7d204ac858185

Observation 60b19474-c3ea-4af3-92ef-e7ba9ad5b911 · outbound

This paper cites Target-Speaker Voice Activity Detection: a Novel Approach for Multi-Speaker Diarization in a Dinner Party Scenario.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Target-Speaker Voice Activity Detection: a Novel Approach for Multi-Speaker Diarization in a Dinner Party Scenario

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:50:17.422697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:50:17.200864Z digest=sha256:37a7f7dd2ca27bf456683c30cb5499271c49954b897d31a96449198f5af61461

Observation 866845d2-be7c-46f1-ad3c-b8423cd743d0 · outbound

This paper cites Tar- get speaker voice activity detection with transformers and its in- tegration with end-to-end neural diarization,.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Tar- get speaker voice activity detection with transformers and its in- tegration with end-to-end neural diarization,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:17.696412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:50:17.206891Z digest=sha256:45e613cce2bfa6892d1673c172329a8d34d5e6b50772a9630ae53d7f6a4af00a

Observation 7715e151-8ace-474f-a08c-23f071d90dab · outbound

This paper cites Adapting self-supervised models to multi-talker speech recognition using speaker embeddings,.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Adapting self-supervised models to multi-talker speech recognition using speaker embeddings,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:17.682041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:50:17.211578Z digest=sha256:123c7f7b30cebb8402b0ccf163dd0f8431a9b9388de835939c1a33e9c073d3a3

Observation 35cb56ab-2885-4e35-916f-a01a5f63ea7d · outbound

This paper cites The Conformer encoder has 12 layers, while the Transformer decoder is structured with 6 layers.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition The Conformer encoder has 12 layers, while the Transformer decoder is structured with 6 layers

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:17.892600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:50:17.093997Z digest=sha256:53d8521781247d545f4dc566556c5213b32600ad14362c57fd73570dcc798df3

Observation 09c22c77-82f4-48cd-a98f-8a49be71a03b · outbound

This paper cites Con- nectionist temporal classification: labelling unsegmented se- quence data with recurrent neural networks,.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Con- nectionist temporal classification: labelling unsegmented se- quence data with recurrent neural networks,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:17.216547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:17.216547Z digest=sha256:6330d58a9585fbec06c79c88b087cf8073af096ca9bf4a550eb514dae7bd1e2b

Observation 1154dfe7-8571-4985-8f49-c0b593f47eeb · outbound

This paper cites Target Speaker ASR with Whisper.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Target Speaker ASR with Whisper

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:50:17.399004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:50:17.222508Z digest=sha256:3776fba3eb49471c78622c720363c4b78199f68da56e8da42fec81927fd35df1

Observation 33eeac7b-ef79-4c00-9143-2aa993c6071f · outbound

This paper cites DiCoW: Diarization-Conditioned Whisper for Target Speaker Automatic Speech Recognition.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition DiCoW: Diarization-Conditioned Whisper for Target Speaker Automatic Speech Recognition

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:17.227570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:17.227570Z digest=sha256:03160821d80c106234b1c1c676524d95a7ac966af68daeee6058885f8be32610

Observation 61ec0b16-5040-4d03-b1d6-95f7c08db808 · outbound

This paper cites But/jhu system descrip- tion for chime-8 notsofar-1 challenge,.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition But/jhu system descrip- tion for chime-8 notsofar-1 challenge,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:17.640268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:50:17.237094Z digest=sha256:7cda51931241e49753d3713274dcb6ae82f9b31b86a923f672c6f6a90c026cda

Observation 068f07fe-08e1-42f7-8d45-08c3a86ec4d0 · outbound

This paper cites ESPnet: End-to-End Speech Processing Toolkit.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition ESPnet: End-to-End Speech Processing Toolkit

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:17.242039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:17.242039Z digest=sha256:90c029c672ef4b934881b0f3f64448a2d29c4df5d849f6c3b6db82ccfb7c0c3e

Observation 4ddde7f5-7f1b-4edd-a0be-6f94237742de · outbound

This paper cites Lib- rispeech: an asr corpus based on public domain audio books,.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Lib- rispeech: an asr corpus based on public domain audio books,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:17.625398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:50:17.246975Z digest=sha256:d54dba496e334899e561129ef5ac1801d1a2e97c1e6128dd4c351ee3941c5fe1

Observation 0123293d-e901-4c82-8787-7cce00d2debe · outbound

This paper cites LibriMix: An Open-Source Dataset for Generalizable Speech Separation.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition LibriMix: An Open-Source Dataset for Generalizable Speech Separation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:17.251656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:17.251656Z digest=sha256:baca356b9a1f777d210349c2842ba8814cc9e897bbea0c8447b8c5e7a4855edd

Observation 0fa73f01-3f40-4b90-a87c-e5527b34ac03 · outbound

This paper cites Conformer: Convolution-augmented Transformer for Speech Recognition.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:17.256472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:17.256472Z digest=sha256:5a37e4dfa6a52fe5fbf7dd24c366b50bc81fc220034aa822e8c78a7122703e29

Observation b39777e2-7e61-493e-9d59-8ffc4ff45661 · outbound

This paper cites Attention is all you need,.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Attention is all you need,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:17.262241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:17.262241Z digest=sha256:59ed23bc270f104af2a8729a40cad80f93ac0ab7d7558eeb16f2f0637cbe2258

Observation 38baa58f-d614-4712-8dd3-d8bce09ac4e2 · outbound

This paper cites Wavlm: Large-scale self- supervised pre-training for full stack speech processing,.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Wavlm: Large-scale self- supervised pre-training for full stack speech processing,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:17.267393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:17.267393Z digest=sha256:c3de0afe4500e688586a21e74710785064eefeb2d6a30b8dbde8b0eaf20688c0

Observation 63a2af0e-2c8b-4091-8761-28d663944392 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Adam: A Method for Stochastic Optimization

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:17.272359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:17.272359Z digest=sha256:56f76a5d153aa6cb00f7c9f777c1e37f6527ec32ecf2740929d6513fa316c76b

Pith citing papers

Observation 1ad7751f-54f4-4b45-964d-2f03033599d5 · inbound

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition cites this paper.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:50:17.589525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:50:17.073394Z digest=sha256:a6e583f6691418f28666675983c159de5aebc45ead2e952e0e51b8fe9fd3bbab