Pith. sign in

Paper Citation Record · LEDGER

Sortformer: A Novel Approach for Permutation-Resolved Speaker Supervision in Speech-to-Text Systems

As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2409.06656.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2409.06656 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T14:35:51.725915Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f059a1fd-6090-459c-b269-2647fa982092 · inbound

Streaming Sortformer: Speaker Cache-Based Online Speaker Diarization with Arrival-Time Ordering cites this paper.

Streaming Sortformer: Speaker Cache-Based Online Speaker Diarization with Arrival-Time Ordering Sortformer: A Novel Approach for Permutation-Resolved Speaker Supervision in Speech-to-Text Systems

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T14:35:51.725915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:35:51.725915Z digest=sha256:a9e91c1a786bf4843a92784fedf13b7988f542848af5270e7ba8d2d9e31237aa

Observation eab73eb0-4f5e-4f6e-807b-950354e6ce6d · inbound

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation cites this paper.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation Sortformer: A Novel Approach for Permutation-Resolved Speaker Supervision in Speech-to-Text Systems

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:18:03.095617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T05:17:11.690484Z digest=sha256:122ba740622a35a75c312ec674f7cec33b343d8cb43fd9bf9da76d59d8c0b1f3

Observation 14d97904-3a22-40ed-920a-980f82fd1886 · inbound

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation cites this paper.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation Sortformer: A Novel Approach for Permutation-Resolved Speaker Supervision in Speech-to-Text Systems

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:04:58.049750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:29f56eea51de41f33ebc99de7bb998b192d57b4eed4a903452ecf4265e340ace

Observation 6e02d4d5-1fba-4931-8dd9-f269d5b180e6 · inbound

Audio-Mind: An Auditable Agentic Framework for Audio Understanding cites this paper.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding Sortformer: A Novel Approach for Permutation-Resolved Speaker Supervision in Speech-to-Text Systems

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T10:13:17.553901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:034137d1ff4a9507d86fb0bb7d84bfcc93c175029f15d5980495f99f284ebaa9

Observation 3e0f178c-107a-4681-9e80-8a7186c4cc47 · inbound

Afrispeech Semantics: Evaluating Audio Semantic Reasoning in Spoken Language Models Across Domains and Accents cites this paper.

Afrispeech Semantics: Evaluating Audio Semantic Reasoning in Spoken Language Models Across Domains and Accents Sortformer: A Novel Approach for Permutation-Resolved Speaker Supervision in Speech-to-Text Systems

Reference 247

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T22:15:05.666445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-30T22:11:44.891731Z digest=sha256:81b1b5617ffae2658e0bc178c70faeaf1f05c7183e387aa15329cfd85267c760

Observation 829ccd89-c919-433d-8f19-52085df9cbf3 · inbound

Position-Aware Target Speaker Extraction for Long-Form Multi-Party Conversations: A Diarization-Free Framework for ASR cites this paper.

Position-Aware Target Speaker Extraction for Long-Form Multi-Party Conversations: A Diarization-Free Framework for ASR Sortformer: A Novel Approach for Permutation-Resolved Speaker Supervision in Speech-to-Text Systems

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-06-30T02:54:10.757568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T02:07:29.315595Z digest=sha256:858234ad574e53faefc0025865cc0c6dbda7a47a1f6008583eaf54ff7337d08e