Pith. sign in

Paper Citation Record · LEDGER

Joint ASR and Speaker Role Tagging with Serialized Output Training

As of 8 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 0 inbound Pith citation observations for arXiv:2506.10349.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.10349 v1

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:32:30.776033Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

30 of 30 outbound references displayed

  • verified exact2
  • verified fuzzy13
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a56eb7fb-e67d-4b8a-ba72-b8af49fc8c6e · outbound

This paper cites End-to-end speech recognition: A survey,.

Joint ASR and Speaker Role Tagging with Serialized Output Training End-to-end speech recognition: A survey,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:32:32.220372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:32:27.650243Z digest=sha256:8513467bed8af83a063271d532ca17e926d6f9fbe9284580babf3a65c6e6cfa1

Observation a499c853-59cd-44d6-95fc-4aa43855ff14 · outbound

This paper cites A review of speaker diarization: Recent advances with deep learning,.

Joint ASR and Speaker Role Tagging with Serialized Output Training A review of speaker diarization: Recent advances with deep learning,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T04:32:27.749181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:32:27.749181Z digest=sha256:cf63ddebb5fa9a0bea1b040e9175ae67b6fb222af7bde6c96389b0aa0ccfe340

Observation c6aa7b7a-10db-4bbd-a7ea-4ac0cfda46f3 · outbound

This paper cites Joint vs sequential speaker- role detection and automatic speech recognition for air-traffic control,.

Joint ASR and Speaker Role Tagging with Serialized Output Training Joint vs sequential speaker- role detection and automatic speech recognition for air-traffic control,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:32:32.201669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:32:27.859759Z digest=sha256:d1f0c34f8c4d7e7a7d3248ae814f8830fcfc4d16a4e23ce9ad455f37a6eb8c1d

Observation d9b4c47a-ad4d-405a-b808-72a72c24843f · outbound

This paper cites Joint Speech Recognition and Speaker Diarization via Sequence Transduction.

Joint ASR and Speaker Role Tagging with Serialized Output Training Joint Speech Recognition and Speaker Diarization via Sequence Transduction

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T04:32:27.969261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:32:27.969261Z digest=sha256:0f34cfc6ecab47f53f3cd2505ad8a5a2e89523eb19b1b1c53ed4d8bd29771196

Observation d2e40b0a-8d3b-402b-a795-1c6e9919c25c · outbound

This paper cites One model to rule them all? towards end-to-end joint speaker diarization and speech recognition,.

Joint ASR and Speaker Role Tagging with Serialized Output Training One model to rule them all? towards end-to-end joint speaker diarization and speech recognition,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:32:32.191045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:32:28.088900Z digest=sha256:d74e8ca9fca20a9d5de761a8171e7d24fc6d5ee72417a0f16f9683af687cef34

Observation df57b59b-7deb-4786-a949-d47497047a53 · outbound

This paper cites Sortformer: A Novel Approach for Permutation-Resolved Speaker Supervision in Speech-to-Text Systems.

Joint ASR and Speaker Role Tagging with Serialized Output Training Sortformer: A Novel Approach for Permutation-Resolved Speaker Supervision in Speech-to-Text Systems

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T04:32:28.219916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:32:28.219916Z digest=sha256:d4d2a952eda8b6ab23e0d54f1dea9d973478947f5f943861732a808d3570e714

Observation 88a4f7d2-0faa-4bff-8e09-ef6eadc0959e · outbound

This paper cites Serialized Output Training for End-to-End Overlapped Speech Recognition.

Joint ASR and Speaker Role Tagging with Serialized Output Training Serialized Output Training for End-to-End Overlapped Speech Recognition

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T04:32:28.346093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:32:28.346093Z digest=sha256:31f0719c0ce61adb9bd1c5520bad609562c42adc40f2cc3b1ce56ac1ed9fe46c

Observation 02be410a-d15b-4189-acfc-7e06fdf39b44 · outbound

This paper cites Robust speech recognition via large-scale weak supervi- sion,.

Joint ASR and Speaker Role Tagging with Serialized Output Training Robust speech recognition via large-scale weak supervi- sion,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T04:32:28.501651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:32:28.501651Z digest=sha256:f7cb86153f2f8a801dbc9c97d1a58bacb20f1aa024b3044a21c3e8f4d87b58f5

Observation 037f3b84-3a70-477b-82b3-41105d6fdbd0 · outbound

This paper cites Large Language Models based ASR Error Correction for Child Conversations.

Joint ASR and Speaker Role Tagging with Serialized Output Training Large Language Models based ASR Error Correction for Child Conversations

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:32:31.149009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:32:28.582864Z digest=sha256:7798e5af4c047c7cd933ad6cc2dfd4081ac70c82bce784b2cc379f8a2c397422

Observation c398d687-7946-4a49-88de-b288647c4b63 · outbound

This paper cites Whislu: End-to-end spoken language under- standing with whisper,.

Joint ASR and Speaker Role Tagging with Serialized Output Training Whislu: End-to-end spoken language under- standing with whisper,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:32:32.174849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:32:28.696382Z digest=sha256:bbd34d8bb6d5454ba29f3b9f8bb4753070faef940e054c4b2bebe325d70dab5d

Observation 1590c83d-fced-4424-82f2-2587dd72ea34 · outbound

This paper cites Attention is all you need,.

Joint ASR and Speaker Role Tagging with Serialized Output Training Attention is all you need,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T04:32:28.862742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:32:28.862742Z digest=sha256:c94fb2a1ce40a219064b8bb64d0d9c7ea9464e8340843251869007b8e14bc3c9

Observation dd031743-beb8-4c7e-ae96-72d9ab3117c5 · outbound

This paper cites Directional speech recognition for speaker disambiguation and cross- talk suppression,.

Joint ASR and Speaker Role Tagging with Serialized Output Training Directional speech recognition for speaker disambiguation and cross- talk suppression,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:32:32.158696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:32:28.959051Z digest=sha256:abc012808a5b627ed81520cd0b7556fcffa155c387b63137cae7d04709dcb979

Observation f85097fe-00ee-45bf-aa30-c361fc1584e9 · outbound

This paper cites Agadir: Towards array-geometry agnostic directional speech recognition,.

Joint ASR and Speaker Role Tagging with Serialized Output Training Agadir: Towards array-geometry agnostic directional speech recognition,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:32:32.148804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:32:29.137476Z digest=sha256:1272fff17cb08f14d15d5974e609b14669ef4360bc9abb962fb9656b5f273a79

Observation a856c219-40db-438d-b3ae-ed6df1be813d · outbound

This paper cites Directional source separation for robust speech recognition on smart glasses,.

Joint ASR and Speaker Role Tagging with Serialized Output Training Directional source separation for robust speech recognition on smart glasses,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:32:32.138485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:32:29.257872Z digest=sha256:b9345a1512561c6f5035e108027b269607ac43b38d6b49b65b769a35dc99f17c

Observation 09ba88e8-626d-4d8d-8e53-1a54f04deab1 · outbound

This paper cites Streaming Multi-Talker ASR with Token-Level Serialized Output Training.

Joint ASR and Speaker Role Tagging with Serialized Output Training Streaming Multi-Talker ASR with Token-Level Serialized Output Training

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T04:32:29.388286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:32:29.388286Z digest=sha256:e94848f0db1aa3f30766995291279fa6521fc0361cef3dd01f19ed41482bbfdc

Observation e8f2ff60-b511-4cd0-b851-ba330a7c4caa · outbound

This paper cites Bertraffic: Bert-based joint speaker role and speaker change detection for air traffic control com- munications,.

Joint ASR and Speaker Role Tagging with Serialized Output Training Bertraffic: Bert-based joint speaker role and speaker change detection for air traffic control com- munications,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:32:32.126102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:32:29.500842Z digest=sha256:e4fed8b46b93a3940a77af1684f8ac9418485511497a9a74ee22cd8976179aca

Observation 22877d80-cfea-492a-94bf-01c534ad243c · outbound

This paper cites Bert: Pre-training of deep bidirectional transformers for language understanding,.

Joint ASR and Speaker Role Tagging with Serialized Output Training Bert: Pre-training of deep bidirectional transformers for language understanding,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T04:32:29.676078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:32:29.676078Z digest=sha256:3f2b5af383f4a13c24b6911ac63cb22ee0bb0f404012b31eda9e8bfea94983f9

Observation 10c32fc3-2a65-47b7-ba7d-e6d474f1b1c4 · outbound

This paper cites Who said what wsw 2.0? enhanced automated analysis of preschool classroom speech,.

Joint ASR and Speaker Role Tagging with Serialized Output Training Who said what wsw 2.0? enhanced automated analysis of preschool classroom speech,

Reference 18

Resolution
verified exact
raw_fallback, observed 2026-08-07T04:32:31.012212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:32:29.756704Z digest=sha256:683d01d5674ca8cdba819d60cbf2bb9f32ba8813a2cfc83bd7e12d49dbb8cb00

Observation 4d38ee8c-7cfe-404d-86ed-04b746ea24ed · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech representations,.

Joint ASR and Speaker Role Tagging with Serialized Output Training wav2vec 2.0: A framework for self-supervised learning of speech representations,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T04:32:29.844757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:32:29.844757Z digest=sha256:71ca3465c1175db11c14ed7a0a74a3b3fffe2e43a1b1f536b70c6d73834b0c5b

Observation 4b44b5da-5781-4bf6-a0fe-b697476cb03e · outbound

This paper cites Exploring speech foundation models for speaker diariza- tion in child-adult dyadic interactions,.

Joint ASR and Speaker Role Tagging with Serialized Output Training Exploring speech foundation models for speaker diariza- tion in child-adult dyadic interactions,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:32:32.023994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:32:29.972442Z digest=sha256:442c61ce110ab2d4733b04175f5458a0f62bcbd62b3a95ac6b5a883c1fad7a37

Observation a6e47252-1fb9-4101-80f5-121cfbd47230 · outbound

This paper cites Data efficient child-adult speaker diarization with simulated conversations,.

Joint ASR and Speaker Role Tagging with Serialized Output Training Data efficient child-adult speaker diarization with simulated conversations,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:32:31.877891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:32:30.007711Z digest=sha256:147d33fe9bef5d93d0dbcc1a668a0727de5f77f3d14ed475dcf6f6d716897360

Observation 65de492a-1c5c-4f55-874d-43f4866d19d2 · outbound

This paper cites Vox-Profile: A Speech Foundation Model Benchmark for Characterizing Diverse Speaker and Speech Traits.

Joint ASR and Speaker Role Tagging with Serialized Output Training Vox-Profile: A Speech Foundation Model Benchmark for Characterizing Diverse Speaker and Speech Traits

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T04:32:30.081659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:32:30.081659Z digest=sha256:9afcc9d1db4e8241553dcf9580a0182f7af135045236c5770b64574785aa3317

Observation 74354199-cc9d-4234-8519-b698d0098208 · outbound

This paper cites The chime-8 mmcsg chal- lenge: Multi-modal conversations in smart glasses,.

Joint ASR and Speaker Role Tagging with Serialized Output Training The chime-8 mmcsg chal- lenge: Multi-modal conversations in smart glasses,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:32:31.637652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:32:30.179128Z digest=sha256:564b98cadc84519946c99049314e80d371337c365478bc4b033fb23bd0ae0478

Observation 97d95895-7355-4b31-83b6-d50e6d76d3c4 · outbound

This paper cites XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale.

Joint ASR and Speaker Role Tagging with Serialized Output Training XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T04:32:30.227311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:32:30.227311Z digest=sha256:3c9aa87d0216402a65a203bfd25570ad7811ce3f3e677bae1b9303a8b86f4d11

Observation d2271d5a-ea78-494c-9b26-a5c82d23040f · outbound

This paper cites Wavlm: Large-scale self-supervised pre- training for full stack speech processing,.

Joint ASR and Speaker Role Tagging with Serialized Output Training Wavlm: Large-scale self-supervised pre- training for full stack speech processing,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T04:32:30.318273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:32:30.318273Z digest=sha256:7fd05b4a1549724581b04655b13879a94a7c1470e44279beb0bc384c28546398

Observation d8de2a0e-6168-4d88-ab94-a6a625aaacae · outbound

This paper cites SUPERB: Speech processing Universal PERformance Benchmark.

Joint ASR and Speaker Role Tagging with Serialized Output Training SUPERB: Speech processing Universal PERformance Benchmark

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T04:32:30.437423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:32:30.437423Z digest=sha256:d75041a61fde56d15e0115b912148511b5623e40170aa5ed8d5abad5bc7dacd7

Observation 353ba2e7-ac54-4706-b7d4-d03b74d90228 · outbound

This paper cites Playlogue: Dataset and benchmarks for analyzing adult- child conversations during play,.

Joint ASR and Speaker Role Tagging with Serialized Output Training Playlogue: Dataset and benchmarks for analyzing adult- child conversations during play,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:32:31.412456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:32:30.521621Z digest=sha256:9df3ab302064515ddfc46f70bc56b7d9b392bef6026d92ddaa5fbd2523f55339

Observation 5ddb2afe-44ea-4564-848f-49f535c1b8ef · outbound

This paper cites The talkbank project,.

Joint ASR and Speaker Role Tagging with Serialized Output Training The talkbank project,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:32:31.273720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:32:30.571547Z digest=sha256:2af673fd3daaed3c8b1bb526a1b53d4813e39cd7bfb481cf82d062a76cf6bc7e

Observation 3c41ea32-f349-4914-95bf-a58fa6aec58b · outbound

This paper cites NeMo: a toolkit for building AI applications using Neural Modules.

Joint ASR and Speaker Role Tagging with Serialized Output Training NeMo: a toolkit for building AI applications using Neural Modules

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T04:32:30.682646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:32:30.682646Z digest=sha256:d7733caa5c28b62c8fc9c8d5d42ff5549e8e8c4b4efc1ce4ef4dd4b05e88efdc

Observation 455c2cf4-d61f-4173-84a1-f0b760fedc38 · outbound

This paper cites HuggingFace's Transformers: State-of-the-art Natural Language Processing.

Joint ASR and Speaker Role Tagging with Serialized Output Training HuggingFace's Transformers: State-of-the-art Natural Language Processing

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T04:32:30.776033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:32:30.776033Z digest=sha256:a8294b43b915812fda168ea2de719a7aff82b7eed64004ca5188eac1b11b3c45

Pith citing papers

No inbound Pith citation observations are available.