Pith. sign in

Paper Citation Record · LEDGER

Sortformer: A Novel Approach for Permutation-Resolved Speaker Supervision in Speech-to-Text Systems

As of 11 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2409.06656.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2409.06656 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:32:28.219916Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation df57b59b-7deb-4786-a949-d47497047a53 · inbound

Joint ASR and Speaker Role Tagging with Serialized Output Training cites this paper.

Joint ASR and Speaker Role Tagging with Serialized Output Training Sortformer: A Novel Approach for Permutation-Resolved Speaker Supervision in Speech-to-Text Systems

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T04:32:28.219916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:32:28.219916Z digest=sha256:553d999dcbbbe06af2591d40abc415a2a91ca88546e0358a20f7e46ade1a5ed1

Observation d29b06c6-a457-4291-9d69-02ce65b9c23a · inbound

Speaker Targeting via Self-Speaker Adaptation for Multi-talker ASR cites this paper.

Speaker Targeting via Self-Speaker Adaptation for Multi-talker ASR Sortformer: A Novel Approach for Permutation-Resolved Speaker Supervision in Speech-to-Text Systems

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T22:07:47.575403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:07:47.575403Z digest=sha256:db1887b5e62caec8225ca2d7e70a9f36425b3d32f2a24059f4a583f0f6adc23e

Observation ba6a276c-475e-4e45-ac1e-57f0e35844b8 · inbound

A Unified Speech LLM for Diarization and Speech Recognition in Multilingual Conversations cites this paper.

A Unified Speech LLM for Diarization and Speech Recognition in Multilingual Conversations Sortformer: A Novel Approach for Permutation-Resolved Speaker Supervision in Speech-to-Text Systems

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T22:42:22.701392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:42:22.701392Z digest=sha256:8ddfe4916a28aa16484804ec57697b7e5804da9be60c4337111a4d236204225b

Observation f059a1fd-6090-459c-b269-2647fa982092 · inbound

Streaming Sortformer: Speaker Cache-Based Online Speaker Diarization with Arrival-Time Ordering cites this paper.

Streaming Sortformer: Speaker Cache-Based Online Speaker Diarization with Arrival-Time Ordering Sortformer: A Novel Approach for Permutation-Resolved Speaker Supervision in Speech-to-Text Systems

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T14:35:51.725915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:35:51.725915Z digest=sha256:830cf19acde0ea1f4dd4574105a7240ff868f312293e7c584757c5220deceb15

Observation eab73eb0-4f5e-4f6e-807b-950354e6ce6d · inbound

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation cites this paper.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation Sortformer: A Novel Approach for Permutation-Resolved Speaker Supervision in Speech-to-Text Systems

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:18:03.095617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-20T05:17:11.690484Z digest=sha256:5f83b09ad5113ed3332681462bb267728d81509ca95c595a81c98d3be784776d

Observation 14d97904-3a22-40ed-920a-980f82fd1886 · inbound

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation cites this paper.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation Sortformer: A Novel Approach for Permutation-Resolved Speaker Supervision in Speech-to-Text Systems

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:04:58.049750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:031f80a4ad71d47bdb45f16238018e80ff7f2c2325a32b499f555f2f68a70d1e

Observation 6e02d4d5-1fba-4931-8dd9-f269d5b180e6 · inbound

Audio-Mind: An Auditable Agentic Framework for Audio Understanding cites this paper.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding Sortformer: A Novel Approach for Permutation-Resolved Speaker Supervision in Speech-to-Text Systems

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T10:13:17.553901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:a0cede2dda716717976fd26c29103545ae45319b83583d98e03de63ad4bfff79

Observation 3e0f178c-107a-4681-9e80-8a7186c4cc47 · inbound

Afrispeech Semantics: Evaluating Audio Semantic Reasoning in Spoken Language Models Across Domains and Accents cites this paper.

Afrispeech Semantics: Evaluating Audio Semantic Reasoning in Spoken Language Models Across Domains and Accents Sortformer: A Novel Approach for Permutation-Resolved Speaker Supervision in Speech-to-Text Systems

Reference 247

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T22:15:05.666445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-30T22:11:44.891731Z digest=sha256:ef5cca14589248a4591560701f0db74792310e6f2ba73501dd3d1dc0ad9a17bb

Observation 829ccd89-c919-433d-8f19-52085df9cbf3 · inbound

Position-Aware Target Speaker Extraction for Long-Form Multi-Party Conversations: A Diarization-Free Framework for ASR cites this paper.

Position-Aware Target Speaker Extraction for Long-Form Multi-Party Conversations: A Diarization-Free Framework for ASR Sortformer: A Novel Approach for Permutation-Resolved Speaker Supervision in Speech-to-Text Systems

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-06-30T02:54:10.757568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-30T02:07:29.315595Z digest=sha256:48ab6deb267d2c7edf77777109e569a973478e74f4ad26f33b4f0dc732d0a510