Pith. sign in

Paper Citation Record · LEDGER

MOSS Transcribe Diarize Technical Report

As of 21 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 6 inbound Pith citation observations for arXiv:2601.01554.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2601.01554 v7

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T12:49:51.849785Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:23:07.516057Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T23:19:03.363492Z

Reference resolution

22 of 22 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 42cb439c-8ddc-40a9-b932-859f225b9d11 · outbound

This paper cites WhisperX: Time-Accurate Speech Transcription of Long-Form Audio.

MOSS Transcribe Diarize Technical Report WhisperX: Time-Accurate Speech Transcription of Long-Form Audio

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T12:49:49.594090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:49:49.594090Z digest=sha256:1c9d364c335229b88d34290f900c0e9e73500039917b71a4159272c2d7f99fae

Observation 067aece8-d695-45e0-ae5e-06ac08f32994 · outbound

This paper cites Pyannote.audio: neural building blocks for speaker diarization.

MOSS Transcribe Diarize Technical Report Pyannote.audio: neural building blocks for speaker diarization

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T12:49:49.630792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:49:49.630792Z digest=sha256:c2b46091d91448d2805bb86912dd371f383064f052e902584fa35a846e1fde53

Observation 6dea2322-2d44-4ca5-8590-bddba16ce66c · outbound

This paper cites The ami meeting corpus: A pre-announcement.

MOSS Transcribe Diarize Technical Report The ami meeting corpus: A pre-announcement

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T12:49:49.703260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:49:49.703260Z digest=sha256:d1f3ba9891595ed55a380c646329ab852070ff19ee54d5e7ba4c10d82249cfc3

Observation 0694c304-05df-4846-a266-05e98b75860c · outbound

This paper cites an unresolved cited work.

MOSS Transcribe Diarize Technical Report Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T12:49:49.749861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:49:49.749861Z digest=sha256:bbb6112169590c9dbffe166dde4f9e613bf7f80da73458c0b57e02e1c9a33d1e

Observation 2c0f0358-b194-4b02-b27d-d509d06188f5 · outbound

This paper cites TimeMarker: A Versatile Video-LLM for Long and Short Video Understanding with Superior Temporal Localization Ability.

MOSS Transcribe Diarize Technical Report TimeMarker: A Versatile Video-LLM for Long and Short Video Understanding with Superior Temporal Localization Ability

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T12:49:49.871069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:49:49.871069Z digest=sha256:eba9785d40eace44eed83ed54a639b23bfa48c997b6ed261e33dede2b4b87652

Observation 938d5d20-137b-438b-92bb-7c6fc392fde4 · outbound

This paper cites Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models.

MOSS Transcribe Diarize Technical Report Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T12:49:50.056919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:49:50.056919Z digest=sha256:8d6669385556be01597411dbd9fc15dedaee5cbd1ce8920948ceb0e0c46af0d8

Observation eaf5114c-d9a4-43ef-baa9-573290076e85 · outbound

This paper cites AISHELL-4: An open source dataset for speech enhancement, separation, recognition and speaker diarization in conference scenario.

MOSS Transcribe Diarize Technical Report AISHELL-4: An open source dataset for speech enhancement, separation, recognition and speaker diarization in conference scenario

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T12:49:50.210239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:49:50.210239Z digest=sha256:a61622de659cbdaf8abdc7488763802eb80263419ce01e5dbdb7d217229bbb25

Observation 2f8090ef-0180-44f1-a70c-00850c8157d8 · outbound

This paper cites End-to-end neural speaker diarization with self-attention.

MOSS Transcribe Diarize Technical Report End-to-end neural speaker diarization with self-attention

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T12:49:50.358730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:49:50.358730Z digest=sha256:3601e9afdc98fcda955e3cc8de04344aff66df6bd2c812c0fb315a32af504e2d

Observation 7e6fb4fc-c6ea-49ea-8cf6-e1b5c4f08327 · outbound

This paper cites The icsi meeting corpus.

MOSS Transcribe Diarize Technical Report The icsi meeting corpus

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T12:49:50.517796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:49:50.517796Z digest=sha256:f949c9b7c4320ff51f772f5d33258277d9b1e8a2c25bacfe2018fce44a334c56

Observation 31092241-eebc-4311-935d-ffd325c13f13 · outbound

This paper cites Serialized output training for end-to-end overlapping speech recognition.

MOSS Transcribe Diarize Technical Report Serialized output training for end-to-end overlapping speech recognition

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T12:49:50.661043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:49:50.661043Z digest=sha256:9392443c90ac78bd3fe62a611230586feb5cdd847d075a2ff934231db913fd60

Observation 61fe783c-c228-4bd6-844c-6fc7fd63dca8 · outbound

This paper cites From simulated mixtures to simulated conversations as training data for neural diarization.

MOSS Transcribe Diarize Technical Report From simulated mixtures to simulated conversations as training data for neural diarization

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T12:49:50.777521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:49:50.777521Z digest=sha256:0093e78bb97d1b3f6b03e19611bd8b6e8765ddf5faeeba6c5410af7f27f6b599

Observation 152a581f-4627-4d58-968b-b5c2b9a59f32 · outbound

This paper cites Levenshtein.

MOSS Transcribe Diarize Technical Report Levenshtein

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T12:49:50.884954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:49:50.884954Z digest=sha256:25cace66073f523ac75b5d453b85b04a3aa24c721c9a3e7f78cb95dac9d46f9c

Observation c4ee19bc-f34d-4c68-a5a1-e7cf26ed7185 · outbound

This paper cites Montreal forced aligner: A trainable text-speech alignment system.

MOSS Transcribe Diarize Technical Report Montreal forced aligner: A trainable text-speech alignment system

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T12:49:51.022176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:49:51.022176Z digest=sha256:a5e632e06a35da2380d00ff2dd5f5e079ec00bf61cd89012fb68c9bf05bbfd89

Observation 394aad98-ec97-490e-b5ca-9c09da050ff6 · outbound

This paper cites an unresolved cited work.

MOSS Transcribe Diarize Technical Report Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T12:49:51.143941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:49:51.143941Z digest=sha256:f279d7c184e29d15ecc69a2574d242fe281251dce3486f68477bd11c4bdbfe8b

Observation 23dc1063-51ed-4618-baae-cdf9f4306937 · outbound

This paper cites an unresolved cited work.

MOSS Transcribe Diarize Technical Report Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T12:49:51.215298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:49:51.215298Z digest=sha256:c5432e6787fb9f93b0a3db5a14b21f4a719e9349eaa4d00165c008326d22aad2

Observation 6dfc0ed7-2352-46f4-ad63-69e67442f542 · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

MOSS Transcribe Diarize Technical Report Robust speech recognition via large-scale weak supervision

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T12:49:51.300528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:49:51.300528Z digest=sha256:1daa556b2fd48f9fdac989c9c8467aa87786a70397b0362a4b4edafab7967519

Observation 511e6b07-de34-4c5b-acce-d578aae04e1c · outbound

This paper cites Train short, infer long: Speech-llm enables zero- shot streamable joint asr and diarization on long audio.

MOSS Transcribe Diarize Technical Report Train short, infer long: Speech-llm enables zero- shot streamable joint asr and diarization on long audio

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T12:49:51.399679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:49:51.399679Z digest=sha256:f7f4036fe2bfeb147f70462d410b96829975b69902fed29dcd416d7878a73d9d

Observation baa9dd3c-36aa-4476-9450-3f69b1b8eb9d · outbound

This paper cites X-vectors: Robust dnn embeddings for speaker recognition.

MOSS Transcribe Diarize Technical Report X-vectors: Robust dnn embeddings for speaker recognition

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T12:49:51.494444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:49:51.494444Z digest=sha256:071c152cfc670052681398f45b6907f12d4e8e675658ebcaaedda00ceabd1ce1

Observation ef4877bf-2f9a-403b-991d-dbc37df16664 · outbound

This paper cites SALMONN: Towards Generic Hearing Abilities for Large Language Models.

MOSS Transcribe Diarize Technical Report SALMONN: Towards Generic Hearing Abilities for Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T12:49:51.562851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:49:51.562851Z digest=sha256:dceb47217ad40fb9476a6def062c76477e2b0861e46858a0089e86b740e6fca4

Observation b14172fb-6ba3-4c3c-b0bd-afe046228234 · outbound

This paper cites Wang et al.

MOSS Transcribe Diarize Technical Report Wang et al

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T12:49:51.599482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:49:51.599482Z digest=sha256:2a5da6d3e508c13079a656c1cb41e88c35c0153d3558af81813fe29d3e28d55d

Observation 2edd7df7-ec02-4127-8efc-b1deba10863d · outbound

This paper cites an unresolved cited work.

MOSS Transcribe Diarize Technical Report Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T12:49:51.739851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:49:51.739851Z digest=sha256:dcc5d8b92d9c7f8ac3f0f7ec0b5c8930319e8a2eb81884112435ba5438f5d8dc

Observation 4ce3c9b1-4159-4cfb-a349-c5e50394cb19 · outbound

This paper cites Speechgpt: Empow- ering large language models with intrinsic cross-modal conversational abilities.

MOSS Transcribe Diarize Technical Report Speechgpt: Empow- ering large language models with intrinsic cross-modal conversational abilities

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T12:49:51.849785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:49:51.849785Z digest=sha256:02145707d144ecc7987e142ee6271d409ba7ac380dc43915bd191f2af501566b

Pith citing papers

Observation c16340d9-16ca-4317-88b6-b50f734ed571 · inbound

DM-ASR: Diarization-aware Multi-speaker ASR with Large Language Models cites this paper.

DM-ASR: Diarization-aware Multi-speaker ASR with Large Language Models MOSS Transcribe Diarize Technical Report

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-07-09T02:19:49.762568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-08T09:18:53.285951Z digest=sha256:7667ad48c4d0925ae47e3504adbf4ebf46b58c7f6883f754a384556ffe72529e

Observation c031f99c-4eee-4423-9a11-79cc20a86d16 · inbound

MOSS-Audio Technical Report cites this paper.

MOSS-Audio Technical Report MOSS Transcribe Diarize Technical Report

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-09T02:19:49.762568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-28T13:05:29.813707Z digest=sha256:339810d63f7d1a7cfbf95960ad4d247775a131fde89cc7b5b0c8d9a99ce1bab8

Observation 83791a05-7fb8-4424-8bbf-857493c59606 · inbound

Balancing ASR and diarization in end-to-end LLMs for multi-talker speech recognition cites this paper.

Balancing ASR and diarization in end-to-end LLMs for multi-talker speech recognition MOSS Transcribe Diarize Technical Report

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-09T02:19:49.762568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T06:04:34.129424Z digest=sha256:cba671a11804ab443d48f00dc5f387e770cb89abb4e6b755c77f1fa6b2d7a0a4

Observation 3a7f2bc8-d5bd-4b2c-8836-e560af80486b · inbound

Grounding Spoken LLMs in Multi-Speaker Audio via Diarization Conditioning cites this paper.

Grounding Spoken LLMs in Multi-Speaker Audio via Diarization Conditioning MOSS Transcribe Diarize Technical Report

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-09T02:19:49.762568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-26T22:39:09.967750Z digest=sha256:cfb7de6c9ace591c27c36aa20f925b377d64ae4514d4dc2c543ae683a390af31

Observation 956adcb4-f2cb-4e1f-991c-50c6418120c3 · inbound

GigaChat Audio: Time-aware Large Audio Language Model cites this paper.

GigaChat Audio: Time-aware Large Audio Language Model MOSS Transcribe Diarize Technical Report

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-14T12:07:39.817307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:07:39.817307Z digest=sha256:08cc456e9de60492f41c1a16a07ecc66ddca1fbd4340b6a28b6e77414f9490c9

Observation 3a94ce0e-a34b-4f8c-bd7d-7706b9d80f26 · inbound

The SLT 2026 SmartGlasses Challenge: Benchmarking Egocentric Multi-Talker Speech Recognition and Understanding with Audio-Language Models cites this paper.

The SLT 2026 SmartGlasses Challenge: Benchmarking Egocentric Multi-Talker Speech Recognition and Understanding with Audio-Language Models MOSS Transcribe Diarize Technical Report

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T00:23:07.516057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:23:07.516057Z digest=sha256:b0026de783460b95896d22969c68d041fd18f314c077795043c7df2411be5421