Pith. sign in

Paper Citation Record · LEDGER

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models

As of 8 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2506.11344.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.11344 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:15:57.190957Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact11
  • verified fuzzy0
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 19df71af-a646-453d-bd7f-104a648e1906 · outbound

This paper cites an unresolved cited work.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Unresolved cited work

Reference 1

Resolution
verified exact
doi, observed 2026-08-07T04:15:58.149537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T04:15:54.728975Z digest=sha256:c590c257c5f4bbc17840257d7ccc26785c39266a3d0eaf1117efbfa0c8e2ea2a

Observation 993142f4-f037-4220-971b-ab252f940ef2 · outbound

This paper cites an unresolved cited work.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:16:00.034386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T04:15:54.834097Z digest=sha256:0ab0cd008bd0da4c82cf81d2805171a96ecc15cb34466d85b79a8bfcd123cd6c

Observation 5ed8ae07-b441-4f7b-a5b7-57f8936c76ed · outbound

This paper cites an unresolved cited work.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:54.918846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:15:54.918846Z digest=sha256:a629d63bd057b57896c560cde810178973cefe39191181412e4f74b51d148e70

Observation cb01b61f-fd41-4c8f-b598-28dd8044266e · outbound

This paper cites an unresolved cited work.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Unresolved cited work

Reference 4

Resolution
verified exact
doi, observed 2026-08-07T04:15:57.993991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T04:15:55.020157Z digest=sha256:ab22d820261e6b4f95c0177cf054baa1a4ae2b59a4e064f0c319fe4b95bfae25

Observation c994171b-17d3-42ab-91b2-5f4411e4ea5f · outbound

This paper cites an unresolved cited work.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:55.109864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:15:55.109864Z digest=sha256:2d4354776e753489666bb1796a90cc1e3631dc46a9de53ade5c3b5ab85b3d719

Observation 17be4d74-41e8-4f01-844a-18642bdeb42d · outbound

This paper cites an unresolved cited work.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Unresolved cited work

Reference 6

Resolution
verified exact
doi, observed 2026-08-07T04:15:57.861434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T04:15:55.177313Z digest=sha256:7873aa2111410f2f8ea26c542157384586c3120661e190e630cb30bb7981d085

Observation a1ed1b1b-f3e6-4de4-bb21-50ba6b6d63de · outbound

This paper cites an unresolved cited work.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:15:59.800209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T04:15:55.283442Z digest=sha256:7cea06c81a483160bbc2ba27313f3f0ecc7eb542c61316b3aca09589516e57db

Observation ace6baef-267f-4a29-ace1-3d71ad5e36f5 · outbound

This paper cites Chafe, Charles Meyer, and Sandra A.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Chafe, Charles Meyer, and Sandra A

Reference 8

Resolution
verified exact
doi, observed 2026-08-07T04:15:57.690637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T04:15:55.365852Z digest=sha256:ea0dc7d5812ece14be66d471f8922ba89fbc02f5cf7e009fed1fc97cd866c9a4

Observation 2e03906e-358a-43e6-a2ed-311dbcc09d67 · outbound

This paper cites an unresolved cited work.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Unresolved cited work

Reference 9

Resolution
verified exact
doi, observed 2026-08-07T04:15:57.520517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T04:15:55.466266Z digest=sha256:45190283e616333f776aedf255606870310b8bfb3ab84bfd5cc36184c6990200

Observation 1e835360-9bfd-41d2-a04c-660acea1262a · outbound

This paper cites an unresolved cited work.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:55.551726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:15:55.551726Z digest=sha256:0cb7ebdebaeb7807ae8e135f897e1dd91243883e5fefb68692fd50131aa668b3

Observation 79edbcff-71b0-45fd-9ae9-ec794a32c25b · outbound

This paper cites an unresolved cited work.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:55.634299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:15:55.634299Z digest=sha256:9b6bcf6c88236cfb00871f6323a43329c6a861fb117c4b40fcbf2694f7da012e

Observation 56f28ec8-925f-4801-b71f-28a3d2e7124b · outbound

This paper cites Partially Observed Discrete-Time Risk-Sensitive Mean Field Games.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Partially Observed Discrete-Time Risk-Sensitive Mean Field Games

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:55.779217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:15:55.779217Z digest=sha256:4a1c15df510f30f9c0e0a077ecd91e2ebfe49e828de9f8f0d733fdc750e089c9

Observation 5bf6fa3a-dd23-4e8f-b5a0-d3ce690a832f · outbound

This paper cites Transcribe-to-Diarize: Neural Speaker Diarization for Unlimited Number of Speakers using End-to-End Speaker-Attributed ASR.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Transcribe-to-Diarize: Neural Speaker Diarization for Unlimited Number of Speakers using End-to-End Speaker-Attributed ASR

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:15:59.267143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T04:15:55.873688Z digest=sha256:353fb003e20e5285daf15a0c90161477efd84773cfc93e5eb496fd5b81e9373b

Observation 8a449894-0fcb-4aea-80f3-7b42929f58de · outbound

This paper cites TitaNet: Neural Model for speaker representation with 1D Depth-wise separable convolutions and global context.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models TitaNet: Neural Model for speaker representation with 1D Depth-wise separable convolutions and global context

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:55.984530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:15:55.984530Z digest=sha256:2453b282018d3bdf55821d3218d96e88d71a569202bea891afbd9245617430fc

Observation 4de9f800-773f-4349-84b4-e4badef5a1ff · outbound

This paper cites an unresolved cited work.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:56.124519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:15:56.124519Z digest=sha256:bd6e26cd3501377d738768a3e7a6aba53057c178bb43acb1d66520b7bcf8ae6a

Observation 080148ad-84de-4e4a-a935-70bdf49bc317 · outbound

This paper cites Enhancing Speaker Diarization with Large Language Models: A Contextual Beam Search Approach.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Enhancing Speaker Diarization with Large Language Models: A Contextual Beam Search Approach

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:15:59.038611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T04:15:56.218052Z digest=sha256:2036d750f7131f8ff8c94afd13160dc474238b91ed8e4d99b5e750e4bc315d6a

Observation 3942f902-e29f-430b-82eb-ffb37e5f278b · outbound

This paper cites Multi-scale Speaker Diarization with Dynamic Scale Weighting.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Multi-scale Speaker Diarization with Dynamic Scale Weighting

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:15:58.855167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T04:15:56.335846Z digest=sha256:fd1b031a478a8f7df446329180993325cd4ad5546d85a73b39fd7e964e0f9d15

Observation af58f791-1946-4fc6-ae89-d2669e4cb1e4 · outbound

This paper cites an unresolved cited work.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Unresolved cited work

Reference 18

Resolution
verified exact
doi, observed 2026-08-07T04:15:57.370849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T04:15:56.441327Z digest=sha256:f97ffd4cf94140a9a6f27bdb08282ced81df570980895ca2aec51d904205cef2

Observation 301c54e8-c851-4259-9ddd-bfd8ba67088e · outbound

This paper cites Pedregosa, G.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Pedregosa, G

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:56.524243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:15:56.524243Z digest=sha256:2e79d1913603ac7053634959c06601776286847af386657050f7fccf57e2e046

Observation f58fe505-57b4-4c36-866e-44a38872b049 · outbound

This paper cites an unresolved cited work.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:15:59.558334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T04:15:56.604502Z digest=sha256:caf7214358f55332be887d3363e135b894093d562ddae86a73088d90616d7221

Observation 5092846e-ce14-438f-b405-42cea1176d02 · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Robust Speech Recognition via Large-Scale Weak Supervision

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:56.689663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:15:56.689663Z digest=sha256:701b0f394d2b98ef4abe6abf423a4fd8406d45c0d7f55bb7b7756f627b63d055

Observation b064a4f8-1330-442d-a145-63034f2a7c02 · outbound

This paper cites an unresolved cited work.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:56.769223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:15:56.769223Z digest=sha256:39342b386596c0513eec258ecf1145cfe96c6e4b3a09c6b2a8a561fc861b9a59

Observation 80d5c565-14c0-4f84-9a26-578fc0102d5e · outbound

This paper cites Joint Speech Recognition and Speaker Diarization via Sequence Transduction.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Joint Speech Recognition and Speaker Diarization via Sequence Transduction

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:56.818718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:15:56.818718Z digest=sha256:6a1439e2c16c3943834c7f7c041a1ba8dcb435335a7b46d531cfb026ab32e27e

Observation 5c3e1fbd-692c-4af9-b0eb-b9c200824229 · outbound

This paper cites an unresolved cited work.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:56.910787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:15:56.910787Z digest=sha256:5d5cceb6dd6aff510d5ac7f392b0a0c2e67d469bbd9bed44edbde14dfec8433e

Observation bb8a8c06-2f04-4733-acc7-4a8a7c0ad7f4 · outbound

This paper cites TOLD: A Novel Two-Stage Overlap-Aware Framework for Speaker Diarization.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models TOLD: A Novel Two-Stage Overlap-Aware Framework for Speaker Diarization

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:15:58.644634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T04:15:57.027068Z digest=sha256:5a5b59b0ae1e92e5eb56b3248b1958165fc2a10c9b84dd5b912ee9d46f19284f

Observation ad09241f-898f-46c1-984a-091958cd9c27 · outbound

This paper cites Speaker Diarization with LSTM.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Speaker Diarization with LSTM

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:15:58.310770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T04:15:57.099168Z digest=sha256:26df54506c640551f7e8d847f803107f09c7925d5fb1966edf2c26832c4af991

Observation a53b12da-abca-4030-807c-aecd224d6911 · outbound

This paper cites DiarizationLM: Speaker Diarization Post-Processing with Large Language Models.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models DiarizationLM: Speaker Diarization Post-Processing with Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:57.190957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:15:57.190957Z digest=sha256:19dfaa5cccf34d676d4aa482de3981297c2a8b1ed98d2f4593ed866053f09322

Pith citing papers

No inbound Pith citation observations are available.