Pith. sign in

Paper Citation Record · LEDGER

Unifying Diarization, Separation, and ASR with Multi-Speaker Encoder

As of 12 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 3 inbound Pith citation observations for arXiv:2508.20474.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.20474 v2

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-18T21:20:00.286888Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-14T18:22:25.235866Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-02T13:06:59.588978Z

Reference resolution

50 of 50 outbound references displayed

  • verified exact0
  • verified fuzzy50
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cc1e516e-daf8-414e-affa-3b6d5cbd1200 · outbound

This paper cites A review of speaker diarization: Recent advances with deep learning.

Unifying Diarization, Separation, and ASR with Multi-Speaker Encoder A review of speaker diarization: Recent advances with deep learning

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T21:21:51.646257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T21:20:00.286888Z digest=sha256:aede64c1115145d18134431356ae88ff15d2f429c5060283c3269279ff5df9f0

Observation 1abf1080-e2d4-4d74-ab43-8d8e1c77d440 · outbound

This paper cites Encoder-decoder based attractors for end-to-end neural diarization.

Unifying Diarization, Separation, and ASR with Multi-Speaker Encoder Encoder-decoder based attractors for end-to-end neural diarization

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T21:21:51.639033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T21:20:00.286888Z digest=sha256:4a3eeb96e4db598b5224f0a5bcefdd117022714ac5e339d607006529fc86dcf3

Observation a60e7931-d5e1-418e-8b2b-0d2664e07d68 · outbound

This paper cites Powerset multi-class cross entropy loss for neural speaker diarization.

Unifying Diarization, Separation, and ASR with Multi-Speaker Encoder Powerset multi-class cross entropy loss for neural speaker diarization

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T21:21:51.522706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T21:20:00.286888Z digest=sha256:7bd04f78c625c1fb8f8a5e557386f2223c837988fdd2b82f769d7c04d2771bf2

Observation b39f67f9-37f3-4af8-8a4b-384bd4efc6f9 · outbound

This paper cites Supervised speech separation based on deep learning: An overview.

Unifying Diarization, Separation, and ASR with Multi-Speaker Encoder Supervised speech separation based on deep learning: An overview

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T21:21:51.618746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T21:20:00.286888Z digest=sha256:00a77ce618ea2661cfee6330a72a5a69530d700a70bc0ab87ec31689df480771

Observation f67f2e83-9d22-4a3e-b71e-2f14b0e089eb · outbound

This paper cites Conv-TasNet: Surpassing ideal time–frequency magnitude masking for speech separation.

Unifying Diarization, Separation, and ASR with Multi-Speaker Encoder Conv-TasNet: Surpassing ideal time–frequency magnitude masking for speech separation

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T21:21:51.622467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T21:20:00.286888Z digest=sha256:5563262cdf56af8eecb7e2c90a9e6ca6944095f557f7c01cb40c298ad9603f7d

Observation 0c6da9fd-a07e-4e07-82b8-c2fb332aabaa · outbound

This paper cites TF-GRIDNET: Making time- frequency domain models great again for monaural speaker separation.

Unifying Diarization, Separation, and ASR with Multi-Speaker Encoder TF-GRIDNET: Making time- frequency domain models great again for monaural speaker separation

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T21:21:51.519854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T21:20:00.286888Z digest=sha256:b94e6e82274a689e19b18a90686ea991711424b5b1f1c6607cdd3d861338daff

Observation 4c02c8ad-f4f3-419d-bdd3-221c81e49f65 · outbound

This paper cites Single-channel multi-talker speech recognition with permutation invariant training.

Unifying Diarization, Separation, and ASR with Multi-Speaker Encoder Single-channel multi-talker speech recognition with permutation invariant training

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T21:21:51.525725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T21:20:00.286888Z digest=sha256:15384cdaedfa81448aa1ea15c47a15290cbb73c730f58531cafae989c4b63ea5

Observation 45c099f3-3f00-47d2-bc4f-d66815e2e024 · outbound

This paper cites A purely end-to-end system for multi-speaker speech recognition.

Unifying Diarization, Separation, and ASR with Multi-Speaker Encoder A purely end-to-end system for multi-speaker speech recognition

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T21:21:51.615878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T21:20:00.286888Z digest=sha256:989cac3824c5f7729aaf19560ee16512cee6bb4a4419e4b40b2be67b6774d881

Observation ed1b6dfe-6b36-4689-8d4b-50422846b25a · outbound

This paper cites End-to-end multi-speaker speech recognition with transformer.

Unifying Diarization, Separation, and ASR with Multi-Speaker Encoder End-to-end multi-speaker speech recognition with transformer

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T21:21:51.528530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T21:20:00.286888Z digest=sha256:dbbc9e53f44beb2deafb6b728daf598048c7117c042ed6217efda31aca008079

Observation c52fda73-b706-4475-84dc-9bfab6b708ac · outbound

This paper cites Serialized output training for end-to-end overlapped speech recognition.

Unifying Diarization, Separation, and ASR with Multi-Speaker Encoder Serialized output training for end-to-end overlapped speech recognition

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T21:21:51.625527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T21:20:00.286888Z digest=sha256:a02fc4c657edc1ce1e7a2617ae0f6056d4f0bb067a009fc86b3800ea7896b6cb

Observation 722ac973-82dc-4934-a935-e7be7630ad27 · outbound

This paper cites Integration of speech separation, diarization, and recog- nition for multi-speaker meetings: System description, comparison, and analysis.

Unifying Diarization, Separation, and ASR with Multi-Speaker Encoder Integration of speech separation, diarization, and recog- nition for multi-speaker meetings: System description, comparison, and analysis

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T21:21:51.510118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T21:20:00.286888Z digest=sha256:04ea1c4a61cb8d242c2d882ec4f5bed5e0dca86642402171183b5cb26cf8f2db

Observation 1ec9cd09-1348-4340-aceb-c9a52307fb26 · outbound

This paper cites Continuous speech separation: Dataset and analysis.

Unifying Diarization, Separation, and ASR with Multi-Speaker Encoder Continuous speech separation: Dataset and analysis

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T21:21:51.512680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T21:20:00.286888Z digest=sha256:15204b99f380c6d31373120698184001dfb12882548e040f86f789a702223538

Observation 67da9005-dfa6-4c8b-b8b9-5b543763bf9a · outbound

This paper cites CHiME-6 Challenge: Tackling multispeaker speech recognition for unsegmented recordings.

Unifying Diarization, Separation, and ASR with Multi-Speaker Encoder CHiME-6 Challenge: Tackling multispeaker speech recognition for unsegmented recordings

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T21:21:51.628622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T21:20:00.286888Z digest=sha256:715846097babc849d469974834db6a2d0b60c8896c4430796ceccd4f89b702f2

Observation bc4b6450-3be1-4b7e-8761-49fc6e2debab · outbound

This paper cites Tandem multitask training of speaker diarisation and speech recognition for meeting transcription.

Unifying Diarization, Separation, and ASR with Multi-Speaker Encoder Tandem multitask training of speaker diarisation and speech recognition for meeting transcription

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T21:21:51.506479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T21:20:00.286888Z digest=sha256:308d363d6efe741b5243783d6a10a742d584f7d1ddeac9de76a856441f9a565a

Observation 762a0fee-3104-4434-9b49-a94fb67e76c3 · outbound

This paper cites pyannote.audio 2.1 speaker diarization pipeline: principle, benchmark, and recipe.

Unifying Diarization, Separation, and ASR with Multi-Speaker Encoder pyannote.audio 2.1 speaker diarization pipeline: principle, benchmark, and recipe

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T21:21:51.515654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T21:20:00.286888Z digest=sha256:ea7b819680468624ed6cbf4bcbd43416a4f143cec47dd259aa7aa018fe0e0b5b

Observation df07468f-a45d-4196-8c01-690fdcf01e29 · outbound

This paper cites TS-SEP: Joint di- arization and separation conditioned on estimated speaker embeddings.

Unifying Diarization, Separation, and ASR with Multi-Speaker Encoder TS-SEP: Joint di- arization and separation conditioned on estimated speaker embeddings

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T21:21:51.534509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T21:20:00.286888Z digest=sha256:1597f92b3aa4daa0e70d0ed0354b56de6314f48b60d697141250adf099d5754d

Observation 50689e6e-e261-4bb3-a9b4-77eb36952239 · outbound

This paper cites PixIT: Joint training of speaker diarization and speech separation from real-world multi-speaker recordings.

Unifying Diarization, Separation, and ASR with Multi-Speaker Encoder PixIT: Joint training of speaker diarization and speech separation from real-world multi-speaker recordings

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T21:21:51.558302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T21:20:00.286888Z digest=sha256:97756d1c156fcd2e73995a368b966b1dd346a464d67f5e7276c8590922390dac

Observation f1bb6545-f1ad-499e-93bd-24ef735f654e · outbound

This paper cites Adapting multi-lingual asr models for handling multiple talkers.

Unifying Diarization, Separation, and ASR with Multi-Speaker Encoder Adapting multi-lingual asr models for handling multiple talkers

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T21:21:51.595289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T21:20:00.286888Z digest=sha256:53f4de62812a1738e9e7806e0dbff76b65e83741a1a23e9e3e9ff8f23e66de6b

Observation 960184b7-bad8-465d-8ef3-12012ca61b0a · outbound

This paper cites Speech recog- nition and multi-speaker diarization of long conversations.

Unifying Diarization, Separation, and ASR with Multi-Speaker Encoder Speech recog- nition and multi-speaker diarization of long conversations

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T21:21:51.598407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T21:20:00.286888Z digest=sha256:5f859c4b58b6ca073c673d705a6289cd2494395753ee9661e7a6d81d53f7f6b6

Observation c2e1733f-1936-42c1-bd33-400d120e93d1 · outbound

This paper cites One model to rule them all ? towards end-to-end joint speaker diarization and speech recognition.

Unifying Diarization, Separation, and ASR with Multi-Speaker Encoder One model to rule them all ? towards end-to-end joint speaker diarization and speech recognition

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T21:21:51.601475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T21:20:00.286888Z digest=sha256:c6a49c77ae888611bc65686868a3479cb292c5ad93ad84d392285428209eb12f

Observation 76cb90c1-5a35-40e0-ba7b-2d067aaae32f · outbound

This paper cites Streaming speaker-attributed ASR with token-level speaker embeddings.

Unifying Diarization, Separation, and ASR with Multi-Speaker Encoder Streaming speaker-attributed ASR with token-level speaker embeddings

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T21:21:51.607693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T21:20:00.286888Z digest=sha256:2ff858b539e25f1793f9c65fe54bff5749e01473e425160626ee970bfddd6cf5

Observation 236d8abc-7075-4136-87f8-49b22738ab0d · outbound

This paper cites MIMO-Speech: End-to-end multi-channel multi- speaker speech recognition.

Unifying Diarization, Separation, and ASR with Multi-Speaker Encoder MIMO-Speech: End-to-end multi-channel multi- speaker speech recognition

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T21:21:51.545023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T21:20:00.286888Z digest=sha256:6f468058702c58515b4bec6139136e50bb89df294e906ea5309844f02b6d19a8

Observation f8614a95-9226-4d32-9ae6-05d74efd289c · outbound

This paper cites Multi-talker ASR for an unknown number of sources: Joint training of source counting, separation and ASR.

Unifying Diarization, Separation, and ASR with Multi-Speaker Encoder Multi-talker ASR for an unknown number of sources: Joint training of source counting, separation and ASR

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T21:21:51.547777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T21:20:00.286888Z digest=sha256:98632651d2619ce429d1bbdcb7d1a405b4e68c83efd4630bdf325e590417be10

Observation cf027948-31a7-4b95-997e-cd92bcc796b3 · outbound

This paper cites All-neural online source separation, counting, and diarization for meeting analysis.

Unifying Diarization, Separation, and ASR with Multi-Speaker Encoder All-neural online source separation, counting, and diarization for meeting analysis

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T21:21:51.553118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T21:20:00.286888Z digest=sha256:5a6292c2e7f2ba8a2c19466aed24abebd47c1e8a793fff9ba6489be96be21474

Observation 1f10084c-135a-4091-a0c4-fc7efeb7c4ff · outbound

This paper cites Neural blind source separa- tion and diarization for distant speech recognition.

Unifying Diarization, Separation, and ASR with Multi-Speaker Encoder Neural blind source separa- tion and diarization for distant speech recognition

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T21:21:51.550143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T21:20:00.286888Z digest=sha256:cec5b1bc619bfe94aeff2c962a26832130c0c560f2984a0d12b6c44618284cd6

Observation 7cab4107-4143-4bef-a3cc-505c4e95f314 · outbound

This paper cites Stcon system for the chime-8 challenge.

Unifying Diarization, Separation, and ASR with Multi-Speaker Encoder Stcon system for the chime-8 challenge

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T21:21:51.580282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T21:20:00.286888Z digest=sha256:b3d1f7a30aac76ea128c545d7b144294921ff44fc221d52bc192cda437eeaf60

Observation 0bb1b868-855a-4ca0-92f0-b9e0733d44bb · outbound

This paper cites BUT/JHU system description for CHiME-8 NOTSOFAR-1 challenge.

Unifying Diarization, Separation, and ASR with Multi-Speaker Encoder BUT/JHU system description for CHiME-8 NOTSOFAR-1 challenge

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T21:21:51.582915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T21:20:00.286888Z digest=sha256:052100000141e169f3bbaa04848303161e4f29ef9a1c40b7a36dc09d6133ac9d

Observation 606b2f6a-6520-46ee-b716-b15d4e638f9b · outbound

This paper cites The USTC-NERCSLIP systems for the CHiME-8 NOTSOFAR-1 challenge.

Unifying Diarization, Separation, and ASR with Multi-Speaker Encoder The USTC-NERCSLIP systems for the CHiME-8 NOTSOFAR-1 challenge

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T21:21:51.542070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T21:20:00.286888Z digest=sha256:02982bdc1be363c1111a241d264fd14cd1e0d010f7c21d3836dbdb11831a4136

Observation f7bdd670-a1ac-4663-a1fa-9f2266ec0e9f · outbound

This paper cites NTT multi-speaker asr system for the DASR task of CHiME-8 challenge.

Unifying Diarization, Separation, and ASR with Multi-Speaker Encoder NTT multi-speaker asr system for the DASR task of CHiME-8 challenge

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T21:21:51.588753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T21:20:00.286888Z digest=sha256:8b7876889fbc16bd35c02e30f2eaddeb771a7ffac17897af6171c0fbfd9bd0f8

Observation ab504eb1-b0bf-4581-a620-6f8fad373239 · outbound

This paper cites wav2vec 2.0: a framework for self-supervised learning of speech representations.

Unifying Diarization, Separation, and ASR with Multi-Speaker Encoder wav2vec 2.0: a framework for self-supervised learning of speech representations

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T21:21:51.570585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T21:20:00.286888Z digest=sha256:5606436ff9d5630e98e1d0b51ea3e88d32de4f229f4c2d2eec1d9e10290f1011

Observation 343e67f3-df0d-41cb-b46d-05d4a14b0da1 · outbound

This paper cites HuBERT: Self-supervised speech representation learning by masked prediction of hidden units.

Unifying Diarization, Separation, and ASR with Multi-Speaker Encoder HuBERT: Self-supervised speech representation learning by masked prediction of hidden units

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T21:21:51.560928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T21:20:00.286888Z digest=sha256:5783eaaab13d930260b6b7155559d32467555f4f8be6dbb00507d8973716f95e

Observation 4858d18b-79bb-4bdc-964d-49e7c49d7d8c · outbound

This paper cites WavLM: Large-scale self-supervised pre-training for full stack speech processing.

Unifying Diarization, Separation, and ASR with Multi-Speaker Encoder WavLM: Large-scale self-supervised pre-training for full stack speech processing

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T21:21:51.612811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T21:20:00.286888Z digest=sha256:cfce87f59a7d6e78ff17cafdda194a36dad63356e91af366085b75b76ec7ea6c

Observation 20417b24-cf3d-49ab-bbbb-f36cea69aadc · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

Unifying Diarization, Separation, and ASR with Multi-Speaker Encoder Robust speech recognition via large-scale weak supervision

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T21:21:51.538557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T21:20:00.286888Z digest=sha256:80d40f7a2952e297a756242c3b8098d8730475900dee4eb90ad8e739bcf7085a

Observation 05ade4e2-ba05-4679-b0c2-5cca8698f5e5 · outbound

This paper cites OWSM-CTC: An open encoder-only speech foundation model for speech recognition, translation, and language identification.

Unifying Diarization, Separation, and ASR with Multi-Speaker Encoder OWSM-CTC: An open encoder-only speech foundation model for speech recognition, translation, and language identification

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T21:21:51.568408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T21:20:00.286888Z digest=sha256:f7784437d710f80090af62b262f90ea8e9607ac09d5921313895df36cd3e69ff

Observation a4e2f819-fca3-42b5-b06a-7089112f9394 · outbound

This paper cites SUPERB: Speech processing universal performance benchmark.

Unifying Diarization, Separation, and ASR with Multi-Speaker Encoder SUPERB: Speech processing universal performance benchmark

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T21:21:51.573026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T21:20:00.286888Z digest=sha256:a90d69347beb60738985e5178dfbd8d9c672890ba56286b532aaee13e01abb91

Observation e0559b20-3460-4a95-9f9c-c98b6852082c · outbound

This paper cites OWSM v3.1: Better and faster open whisper-style speech models based on e-branchformer.

Unifying Diarization, Separation, and ASR with Multi-Speaker Encoder OWSM v3.1: Better and faster open whisper-style speech models based on e-branchformer

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T21:21:51.565979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T21:20:00.286888Z digest=sha256:a1462d5f274b911502b427f7af69b76ddc52ab7c2bc22f9391efb4da7e6d38c7

Observation a776efa8-d7a7-4d1d-ab6e-595abfbe18a7 · outbound

This paper cites LibriMix: An open-source dataset for generalizable speech separation.

Unifying Diarization, Separation, and ASR with Multi-Speaker Encoder LibriMix: An open-source dataset for generalizable speech separation

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T21:21:51.563365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T21:20:00.286888Z digest=sha256:5a81c242aa14e48714f5548d08ded0f6a21c84125a603384425285597808dee3

Observation e8d22c95-1384-41b7-9533-7b1c6676ca72 · outbound

This paper cites E-Branchformer: Branchformer with enhanced merging for speech recognition.

Unifying Diarization, Separation, and ASR with Multi-Speaker Encoder E-Branchformer: Branchformer with enhanced merging for speech recognition

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T21:21:51.592039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T21:20:00.286888Z digest=sha256:c2c92626bf28df6b8019cf4fe59ae843829591ad4db47ca206b3fc6f7d551580

Observation 1aa5f4f6-74fd-42a4-aafb-56df01fc1f70 · outbound

This paper cites End-to-end training of time domain audio separation and recognition.

Unifying Diarization, Separation, and ASR with Multi-Speaker Encoder End-to-end training of time domain audio separation and recognition

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T21:21:51.575655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T21:20:00.286888Z digest=sha256:15fb30a295f2edbb3ea5374ea86f9ce9d5422b29263e13d4652fed9baf886239

Observation bb4fc01a-c0d9-48fc-92eb-a61912afbfeb · outbound

This paper cites The AMI meeting corpus: A pre-announcement.

Unifying Diarization, Separation, and ASR with Multi-Speaker Encoder The AMI meeting corpus: A pre-announcement

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T21:21:51.531268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T21:20:00.286888Z digest=sha256:305ddf5508ac2e6c8570c6097655704cd0ed806fb4fffb1805748cb26767cf3f

Observation 836dae3d-0366-450f-be64-6d9cd2690d65 · outbound

This paper cites The Hitachi-JHU DIHARD III System: Competitive end-to-end neural diarization and x-vector clustering systems combined by dover-lap.

Unifying Diarization, Separation, and ASR with Multi-Speaker Encoder The Hitachi-JHU DIHARD III System: Competitive end-to-end neural diarization and x-vector clustering systems combined by dover-lap

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T21:21:51.577956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T21:20:00.286888Z digest=sha256:0611b14f9d392bfeface7d574b17cebdb20cc7af7763fa3b758063b90f04a768

Observation 36842d5e-d6fc-45d0-9273-c58a1a0edbf1 · outbound

This paper cites The rich transcription 2006 spring meeting recognition evaluation.

Unifying Diarization, Separation, and ASR with Multi-Speaker Encoder The rich transcription 2006 spring meeting recognition evaluation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T21:21:51.585929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T21:20:00.286888Z digest=sha256:f3d5263cd45f112905415723084e63bbc8f6230c73d16845bf68e8d42441c712

Observation a05a1210-cd17-4742-9c24-84c1aa0bf6e8 · outbound

This paper cites A short- time objective intelligibility measure for time-frequency weighted noisy speech.

Unifying Diarization, Separation, and ASR with Multi-Speaker Encoder A short- time objective intelligibility measure for time-frequency weighted noisy speech

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T21:21:51.610432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T21:20:00.286888Z digest=sha256:a867ca6edde110b99bc8433529b175a93861e84fd50209a6b989dfc45e468cb6

Observation f2cf1b82-09e7-42b0-b476-1ec39cfd2ef8 · outbound

This paper cites Performance measurement in blind audio source separation.

Unifying Diarization, Separation, and ASR with Multi-Speaker Encoder Performance measurement in blind audio source separation

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T21:21:51.555861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T21:20:00.286888Z digest=sha256:fb6eb1daaaae63350dc76119183f177164c476385cf5b442fbd2f5d0c6712fe5

Observation 14982b27-c6e9-44cc-b12f-ea1681c0252c · outbound

This paper cites Streaming end-to-end multi-talker speech recognition.

Unifying Diarization, Separation, and ASR with Multi-Speaker Encoder Streaming end-to-end multi-talker speech recognition

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T21:21:51.604338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T21:20:00.286888Z digest=sha256:ddd7c7c68701245312f3ad7fe17eda325cd55610e52e11c7198879649038fbf9

Observation 680035f5-dd40-42cb-b75e-5581cbbcf8d5 · outbound

This paper cites End-to-end Speaker-Attributed ASR with transformer.

Unifying Diarization, Separation, and ASR with Multi-Speaker Encoder End-to-end Speaker-Attributed ASR with transformer

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T21:21:51.634012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T21:20:00.286888Z digest=sha256:05d230890839f31a82c4072e0e55c8024b314c4ca76c9af92b398ca77ecde6d9

Observation b2996ebe-c4bb-49ca-8f35-5bd1017e589c · outbound

This paper cites Empowering whisper as a joint multi- talker and target-talker speech recognition system.

Unifying Diarization, Separation, and ASR with Multi-Speaker Encoder Empowering whisper as a joint multi- talker and target-talker speech recognition system

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T21:21:51.631198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T21:20:00.286888Z digest=sha256:d397ae7eabd16b57a5fad7d654d24dff9c84d6fec62336215d5470b12301b504

Observation 8ad9e094-13de-4499-87bc-facd18ada5c5 · outbound

This paper cites ESPnet: End-to-end speech processing toolkit.

Unifying Diarization, Separation, and ASR with Multi-Speaker Encoder ESPnet: End-to-end speech processing toolkit

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T21:21:51.636553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T21:20:00.286888Z digest=sha256:4608262b518a3184b60af520c5876dd014aeabfc09463aaf15d07dd4acbc37e9

Observation 28581f29-c463-401d-a942-75821bf5a48a · outbound

This paper cites The power of the weighted sum scalarization for approximating multiobjective optimization problems.

Unifying Diarization, Separation, and ASR with Multi-Speaker Encoder The power of the weighted sum scalarization for approximating multiobjective optimization problems

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T21:21:51.641625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T21:20:00.286888Z digest=sha256:07c1ebe4e9b7b6351c554e9a3bb1b04481efd69e012f7151bfc94feb372047b2

Observation f23eaa7c-20d3-47d0-8c89-3111cd76295b · outbound

This paper cites Joint beam search integrating CTC, attention, and trans- ducer decoders.

Unifying Diarization, Separation, and ASR with Multi-Speaker Encoder Joint beam search integrating CTC, attention, and trans- ducer decoders

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T21:21:51.644161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T21:20:00.286888Z digest=sha256:95a64966ff7640d3385f905ecd53f30ac47834d5a5c7f77965b1667cd12962af

Pith citing papers

Observation f5605c3b-0e95-4a0b-9f50-bdb78413ef8a · inbound

Beyond Acoustic Prefixes: Persistent Grounding in Serialized Acoustic Memory for LLM-Based Multi-Talker Speech Recognition cites this paper.

Beyond Acoustic Prefixes: Persistent Grounding in Serialized Acoustic Memory for LLM-Based Multi-Talker Speech Recognition Unifying Diarization, Separation, and ASR with Multi-Speaker Encoder

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-13T17:09:15.281290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T17:09:15.281290Z digest=sha256:88bb8817582fdb54ff8f03251f87028457821993014d684c4fe4b186da80579f

Observation f8dff840-eb26-4084-ae2e-2bdd1e2ebf5e · inbound

Multi-task Learning is Not Enough: Representational Entanglement in Dual-output Second Language Speech Recognition cites this paper.

Multi-task Learning is Not Enough: Representational Entanglement in Dual-output Second Language Speech Recognition Unifying Diarization, Separation, and ASR with Multi-Speaker Encoder

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T13:06:59.590126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-28T01:31:37.223846Z digest=sha256:32f0ca8c9b2bd47e7e4dc9c5e4562d7cc3082495bc15eda8c719146113727b47

Observation 21835f3a-9d1d-4932-b954-8e1d145f9110 · inbound

Multi-task Learning is Not Enough: Representational Entanglement in Dual-output Second Language Speech Recognition cites this paper.

Multi-task Learning is Not Enough: Representational Entanglement in Dual-output Second Language Speech Recognition Unifying Diarization, Separation, and ASR with Multi-Speaker Encoder

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-14T18:22:25.235866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:22:25.235866Z digest=sha256:0dc0ecc459d2027a3a73674d0d6ec8963b362c9c2da777987f79881e38f7bfc8