Pith. sign in

Paper Citation Record · LEDGER

Joint ASR and Speaker Role Tagging with Serialized Output Training

As of 8 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 0 inbound Pith citation observations for arXiv:2506.10349.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.10349 v1

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:32:30.776033Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

30 of 30 outbound references displayed

  • verified exact2
  • verified fuzzy13
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a56eb7fb-e67d-4b8a-ba72-b8af49fc8c6e · outbound

This paper cites End-to-end speech recognition: A survey,.

Joint ASR and Speaker Role Tagging with Serialized Output Training End-to-end speech recognition: A survey,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:32:32.220372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:32:27.650243Z digest=sha256:93f8318d0478e3d2174cdbb226f26369c692408424301f2783544e1946c32ddd

Observation a499c853-59cd-44d6-95fc-4aa43855ff14 · outbound

This paper cites A review of speaker diarization: Recent advances with deep learning,.

Joint ASR and Speaker Role Tagging with Serialized Output Training A review of speaker diarization: Recent advances with deep learning,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T04:32:27.749181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:32:27.749181Z digest=sha256:1368ce19da8e60d1fa92ed8d45de37586bb348fd474eebda45fe47595b82dd69

Observation c6aa7b7a-10db-4bbd-a7ea-4ac0cfda46f3 · outbound

This paper cites Joint vs sequential speaker- role detection and automatic speech recognition for air-traffic control,.

Joint ASR and Speaker Role Tagging with Serialized Output Training Joint vs sequential speaker- role detection and automatic speech recognition for air-traffic control,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:32:32.201669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:32:27.859759Z digest=sha256:ccaa848d10b9963da22f6d51e4a2e8dae9d53d3f4766d0e32069c86b2dcccf3c

Observation d9b4c47a-ad4d-405a-b808-72a72c24843f · outbound

This paper cites Joint Speech Recognition and Speaker Diarization via Sequence Transduction.

Joint ASR and Speaker Role Tagging with Serialized Output Training Joint Speech Recognition and Speaker Diarization via Sequence Transduction

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T04:32:27.969261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:32:27.969261Z digest=sha256:736a787c5151034dbeeccaa6bf2a9102b57db4a48a9824a45452309170914a10

Observation d2e40b0a-8d3b-402b-a795-1c6e9919c25c · outbound

This paper cites One model to rule them all? towards end-to-end joint speaker diarization and speech recognition,.

Joint ASR and Speaker Role Tagging with Serialized Output Training One model to rule them all? towards end-to-end joint speaker diarization and speech recognition,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:32:32.191045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:32:28.088900Z digest=sha256:eb11fe913d47105147b741b7c5811b62e6c08975b6f3b214dd78f5f9b786c398

Observation df57b59b-7deb-4786-a949-d47497047a53 · outbound

This paper cites Sortformer: A Novel Approach for Permutation-Resolved Speaker Supervision in Speech-to-Text Systems.

Joint ASR and Speaker Role Tagging with Serialized Output Training Sortformer: A Novel Approach for Permutation-Resolved Speaker Supervision in Speech-to-Text Systems

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T04:32:28.219916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:32:28.219916Z digest=sha256:f0796f658dc78394201c5f188ccb622f140ffff8cfa0ff9207225354f21c56ac

Observation 88a4f7d2-0faa-4bff-8e09-ef6eadc0959e · outbound

This paper cites Serialized Output Training for End-to-End Overlapped Speech Recognition.

Joint ASR and Speaker Role Tagging with Serialized Output Training Serialized Output Training for End-to-End Overlapped Speech Recognition

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T04:32:28.346093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:32:28.346093Z digest=sha256:086889204341863317cf900f423a9dfe595a1482a9f65a1ec3b54fd82f4184a3

Observation 02be410a-d15b-4189-acfc-7e06fdf39b44 · outbound

This paper cites Robust speech recognition via large-scale weak supervi- sion,.

Joint ASR and Speaker Role Tagging with Serialized Output Training Robust speech recognition via large-scale weak supervi- sion,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T04:32:28.501651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:32:28.501651Z digest=sha256:9e555920717702ba4f3eb8c6e67341cf7b50d993d839910cd67e57f7e9d0a4c6

Observation 037f3b84-3a70-477b-82b3-41105d6fdbd0 · outbound

This paper cites Large Language Models based ASR Error Correction for Child Conversations.

Joint ASR and Speaker Role Tagging with Serialized Output Training Large Language Models based ASR Error Correction for Child Conversations

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:32:31.149009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:32:28.582864Z digest=sha256:b85d16d3e2336b07db9ff299f0f467f9cdd1080c5b7edd7a014ea82c16c18062

Observation c398d687-7946-4a49-88de-b288647c4b63 · outbound

This paper cites Whislu: End-to-end spoken language under- standing with whisper,.

Joint ASR and Speaker Role Tagging with Serialized Output Training Whislu: End-to-end spoken language under- standing with whisper,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:32:32.174849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:32:28.696382Z digest=sha256:8defc8508f38a2459f46063477cd9793185e8c0ad716938750a406db56eac6aa

Observation 1590c83d-fced-4424-82f2-2587dd72ea34 · outbound

This paper cites Attention is all you need,.

Joint ASR and Speaker Role Tagging with Serialized Output Training Attention is all you need,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T04:32:28.862742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:32:28.862742Z digest=sha256:397e253e4e70b1ce84295f7a9a6906cc6d7be3e2483f289ccaa37763eb5702d9

Observation dd031743-beb8-4c7e-ae96-72d9ab3117c5 · outbound

This paper cites Directional speech recognition for speaker disambiguation and cross- talk suppression,.

Joint ASR and Speaker Role Tagging with Serialized Output Training Directional speech recognition for speaker disambiguation and cross- talk suppression,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:32:32.158696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:32:28.959051Z digest=sha256:9a6e21a5ae821219ee22100db0a35630a8b37d58fd08cf2c1a75560e9ff90c9d

Observation f85097fe-00ee-45bf-aa30-c361fc1584e9 · outbound

This paper cites Agadir: Towards array-geometry agnostic directional speech recognition,.

Joint ASR and Speaker Role Tagging with Serialized Output Training Agadir: Towards array-geometry agnostic directional speech recognition,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:32:32.148804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:32:29.137476Z digest=sha256:bc127712962d5cff3ae2cc8c8d841a03987e1b236dd398ff920cb6181874d91a

Observation a856c219-40db-438d-b3ae-ed6df1be813d · outbound

This paper cites Directional source separation for robust speech recognition on smart glasses,.

Joint ASR and Speaker Role Tagging with Serialized Output Training Directional source separation for robust speech recognition on smart glasses,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:32:32.138485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:32:29.257872Z digest=sha256:1dc538d3fb170e84bac1ee28929eb2d57df6e226e3c67ca36aa87fc5bea87b7a

Observation 09ba88e8-626d-4d8d-8e53-1a54f04deab1 · outbound

This paper cites Streaming Multi-Talker ASR with Token-Level Serialized Output Training.

Joint ASR and Speaker Role Tagging with Serialized Output Training Streaming Multi-Talker ASR with Token-Level Serialized Output Training

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T04:32:29.388286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:32:29.388286Z digest=sha256:304b6d27feaf21d4c981de4e0ccee51cc14ea0c48fe89c5b854b9e8a0c24114b

Observation e8f2ff60-b511-4cd0-b851-ba330a7c4caa · outbound

This paper cites Bertraffic: Bert-based joint speaker role and speaker change detection for air traffic control com- munications,.

Joint ASR and Speaker Role Tagging with Serialized Output Training Bertraffic: Bert-based joint speaker role and speaker change detection for air traffic control com- munications,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:32:32.126102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:32:29.500842Z digest=sha256:ff43c3b67795c230a872422e1b5230c8d3c129c3fe105943f7f90a19cac8d090

Observation 22877d80-cfea-492a-94bf-01c534ad243c · outbound

This paper cites Bert: Pre-training of deep bidirectional transformers for language understanding,.

Joint ASR and Speaker Role Tagging with Serialized Output Training Bert: Pre-training of deep bidirectional transformers for language understanding,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T04:32:29.676078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:32:29.676078Z digest=sha256:fb01b0561644822ab2d31fb207ebce92c79a098fdedb497beee924a879a1a515

Observation 10c32fc3-2a65-47b7-ba7d-e6d474f1b1c4 · outbound

This paper cites Who said what wsw 2.0? enhanced automated analysis of preschool classroom speech,.

Joint ASR and Speaker Role Tagging with Serialized Output Training Who said what wsw 2.0? enhanced automated analysis of preschool classroom speech,

Reference 18

Resolution
verified exact
raw_fallback, observed 2026-08-07T04:32:31.012212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:32:29.756704Z digest=sha256:5f42d5b5235007e499cf428bb602e4b3a229aa62b77707333e864089e71cbd86

Observation 4d38ee8c-7cfe-404d-86ed-04b746ea24ed · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech representations,.

Joint ASR and Speaker Role Tagging with Serialized Output Training wav2vec 2.0: A framework for self-supervised learning of speech representations,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T04:32:29.844757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:32:29.844757Z digest=sha256:636463371746b0c34f3f76047e56eafe1dc38e58f485fb76d48c80106217fab0

Observation 4b44b5da-5781-4bf6-a0fe-b697476cb03e · outbound

This paper cites Exploring speech foundation models for speaker diariza- tion in child-adult dyadic interactions,.

Joint ASR and Speaker Role Tagging with Serialized Output Training Exploring speech foundation models for speaker diariza- tion in child-adult dyadic interactions,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:32:32.023994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:32:29.972442Z digest=sha256:3697322e0126b9648dd7e7172aee965d8baf1836c48ea25df5cbe5ff1ada523b

Observation a6e47252-1fb9-4101-80f5-121cfbd47230 · outbound

This paper cites Data efficient child-adult speaker diarization with simulated conversations,.

Joint ASR and Speaker Role Tagging with Serialized Output Training Data efficient child-adult speaker diarization with simulated conversations,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:32:31.877891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:32:30.007711Z digest=sha256:df1443b7ff3b9ed4efe404fd3e2fb6c8c995e08a4cc6d9408352ee697bff6c5c

Observation 65de492a-1c5c-4f55-874d-43f4866d19d2 · outbound

This paper cites Vox-Profile: A Speech Foundation Model Benchmark for Characterizing Diverse Speaker and Speech Traits.

Joint ASR and Speaker Role Tagging with Serialized Output Training Vox-Profile: A Speech Foundation Model Benchmark for Characterizing Diverse Speaker and Speech Traits

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T04:32:30.081659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:32:30.081659Z digest=sha256:c06fefc7b68001e0e0b66e18d1562516d49fa08c620d84d12eeb18118206189a

Observation 74354199-cc9d-4234-8519-b698d0098208 · outbound

This paper cites The chime-8 mmcsg chal- lenge: Multi-modal conversations in smart glasses,.

Joint ASR and Speaker Role Tagging with Serialized Output Training The chime-8 mmcsg chal- lenge: Multi-modal conversations in smart glasses,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:32:31.637652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:32:30.179128Z digest=sha256:096a417cc71fed69b0d1772f597c35a3abd86b6008a5c1f419fb49ae3ae6fb47

Observation 97d95895-7355-4b31-83b6-d50e6d76d3c4 · outbound

This paper cites XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale.

Joint ASR and Speaker Role Tagging with Serialized Output Training XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T04:32:30.227311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:32:30.227311Z digest=sha256:eeb5c8b40635e5e26e7e0ed96d74e5e8e2ec8c2f6fdce46abd65d112be047a6a

Observation d2271d5a-ea78-494c-9b26-a5c82d23040f · outbound

This paper cites Wavlm: Large-scale self-supervised pre- training for full stack speech processing,.

Joint ASR and Speaker Role Tagging with Serialized Output Training Wavlm: Large-scale self-supervised pre- training for full stack speech processing,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T04:32:30.318273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:32:30.318273Z digest=sha256:f488361439a5c869f1baaaa4488941de058b1f7f4840e6d5371b0d6903ada4ad

Observation d8de2a0e-6168-4d88-ab94-a6a625aaacae · outbound

This paper cites SUPERB: Speech processing Universal PERformance Benchmark.

Joint ASR and Speaker Role Tagging with Serialized Output Training SUPERB: Speech processing Universal PERformance Benchmark

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T04:32:30.437423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:32:30.437423Z digest=sha256:ba2ffc68eb3bed9519db18e1be3c95595fcea958e389b2e5583eb384ae39fe3c

Observation 353ba2e7-ac54-4706-b7d4-d03b74d90228 · outbound

This paper cites Playlogue: Dataset and benchmarks for analyzing adult- child conversations during play,.

Joint ASR and Speaker Role Tagging with Serialized Output Training Playlogue: Dataset and benchmarks for analyzing adult- child conversations during play,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:32:31.412456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:32:30.521621Z digest=sha256:e341af76329de4ed2b3588c9f6d0c04c76d7ab78c7709ad7d8f52c33560f73e8

Observation 5ddb2afe-44ea-4564-848f-49f535c1b8ef · outbound

This paper cites The talkbank project,.

Joint ASR and Speaker Role Tagging with Serialized Output Training The talkbank project,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:32:31.273720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:32:30.571547Z digest=sha256:c50cffdc4abc1c8559bab23463553acd8b25f4fd3040643af3c93bbae02e5d89

Observation 3c41ea32-f349-4914-95bf-a58fa6aec58b · outbound

This paper cites NeMo: a toolkit for building AI applications using Neural Modules.

Joint ASR and Speaker Role Tagging with Serialized Output Training NeMo: a toolkit for building AI applications using Neural Modules

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T04:32:30.682646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:32:30.682646Z digest=sha256:73d8130bcb1d6d925848cb3491c691e51e039640ae8e057a276ab4a649f553db

Observation 455c2cf4-d61f-4173-84a1-f0b760fedc38 · outbound

This paper cites HuggingFace's Transformers: State-of-the-art Natural Language Processing.

Joint ASR and Speaker Role Tagging with Serialized Output Training HuggingFace's Transformers: State-of-the-art Natural Language Processing

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T04:32:30.776033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:32:30.776033Z digest=sha256:f6507ac7a2dec8d0de5891f91cca52285d12856d3540be7784bc6d8ead043bc4

Pith citing papers

No inbound Pith citation observations are available.