Pith. sign in

Paper Citation Record · LEDGER

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models

As of 18 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2506.11344.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.11344 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:15:57.190957Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact11
  • verified fuzzy0
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 19df71af-a646-453d-bd7f-104a648e1906 · outbound

This paper cites an unresolved cited work.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Unresolved cited work

Reference 1

Resolution
verified exact
doi, observed 2026-08-07T04:15:58.149537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:15:54.728975Z digest=sha256:335e0326190524069edd8a5b8f84dabe2c332e0abca27442a1ab98c5130932e8

Observation 993142f4-f037-4220-971b-ab252f940ef2 · outbound

This paper cites an unresolved cited work.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:16:00.034386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:15:54.834097Z digest=sha256:65454de261b244c0f121edd1b3eef1b7024b16cd2a664bf2a8f372481467cecc

Observation 5ed8ae07-b441-4f7b-a5b7-57f8936c76ed · outbound

This paper cites an unresolved cited work.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:54.918846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:15:54.918846Z digest=sha256:d35fbf89089be8d95ebf6c848e49181f4314f082d3d61b1af5fe5ce31e704a1c

Observation cb01b61f-fd41-4c8f-b598-28dd8044266e · outbound

This paper cites an unresolved cited work.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Unresolved cited work

Reference 4

Resolution
verified exact
doi, observed 2026-08-07T04:15:57.993991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:15:55.020157Z digest=sha256:ad305a84af0a80c86c460eb98d775b759c903de95a3db9a38836e169d3204716

Observation c994171b-17d3-42ab-91b2-5f4411e4ea5f · outbound

This paper cites an unresolved cited work.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:55.109864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:15:55.109864Z digest=sha256:7ca54d6dc5616cfbe5ba4f8ee2b408d7f8b6b5ed57b673766b16be699b23b146

Observation 17be4d74-41e8-4f01-844a-18642bdeb42d · outbound

This paper cites an unresolved cited work.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Unresolved cited work

Reference 6

Resolution
verified exact
doi, observed 2026-08-07T04:15:57.861434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:15:55.177313Z digest=sha256:e5112aba291102cc83882a4c9c4df9b098a410cd5b11330d10baf3569d4391ac

Observation a1ed1b1b-f3e6-4de4-bb21-50ba6b6d63de · outbound

This paper cites an unresolved cited work.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:15:59.800209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:15:55.283442Z digest=sha256:d5413f5b8b6cb5da36c6b757b4a2cedef6888293b0fa29b868cf2bc4665dd77b

Observation ace6baef-267f-4a29-ace1-3d71ad5e36f5 · outbound

This paper cites Chafe, Charles Meyer, and Sandra A.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Chafe, Charles Meyer, and Sandra A

Reference 8

Resolution
verified exact
doi, observed 2026-08-07T04:15:57.690637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:15:55.365852Z digest=sha256:bf382307c262048985805c2e9ad530e1f79dc9c2f5c88c3ccde8ebde6d4f551a

Observation 2e03906e-358a-43e6-a2ed-311dbcc09d67 · outbound

This paper cites an unresolved cited work.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Unresolved cited work

Reference 9

Resolution
verified exact
doi, observed 2026-08-07T04:15:57.520517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:15:55.466266Z digest=sha256:dd0bd28e6354d9652d8e082254f5cfd6e48f280affcd6328c760d158b78d2d69

Observation 1e835360-9bfd-41d2-a04c-660acea1262a · outbound

This paper cites an unresolved cited work.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:55.551726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:15:55.551726Z digest=sha256:72ae29561492408beb669a47ccd6beb05e956c3693244b17b35cf89c05e855f7

Observation 79edbcff-71b0-45fd-9ae9-ec794a32c25b · outbound

This paper cites an unresolved cited work.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:55.634299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:15:55.634299Z digest=sha256:2c9e764f17ca303a4b3d0e5d1e06b06735f9dc21ffdf578e93bffdce2852e037

Observation 56f28ec8-925f-4801-b71f-28a3d2e7124b · outbound

This paper cites Partially Observed Discrete-Time Risk-Sensitive Mean Field Games.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Partially Observed Discrete-Time Risk-Sensitive Mean Field Games

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:55.779217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:15:55.779217Z digest=sha256:71ade2a8f9a744929bf761845ba8ae476fd61b1f25eb16be2180574c7e4f61c5

Observation 5bf6fa3a-dd23-4e8f-b5a0-d3ce690a832f · outbound

This paper cites Transcribe-to-Diarize: Neural Speaker Diarization for Unlimited Number of Speakers using End-to-End Speaker-Attributed ASR.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Transcribe-to-Diarize: Neural Speaker Diarization for Unlimited Number of Speakers using End-to-End Speaker-Attributed ASR

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:15:59.267143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:15:55.873688Z digest=sha256:1356d2e1832b54d75b5e44380f8862b5f2dcc1f79bb424d685d717c695f56e30

Observation 8a449894-0fcb-4aea-80f3-7b42929f58de · outbound

This paper cites TitaNet: Neural Model for speaker representation with 1D Depth-wise separable convolutions and global context.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models TitaNet: Neural Model for speaker representation with 1D Depth-wise separable convolutions and global context

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:55.984530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:15:55.984530Z digest=sha256:f5ccd4d3946f034b406b2b121c68e4180940052924f1bdec8673b696f797e147

Observation 4de9f800-773f-4349-84b4-e4badef5a1ff · outbound

This paper cites an unresolved cited work.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:56.124519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:15:56.124519Z digest=sha256:c4253827d431f079107f4a901fcd7854cf16d2c582d4971b8c9284d0c371fe63

Observation 080148ad-84de-4e4a-a935-70bdf49bc317 · outbound

This paper cites Enhancing Speaker Diarization with Large Language Models: A Contextual Beam Search Approach.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Enhancing Speaker Diarization with Large Language Models: A Contextual Beam Search Approach

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:15:59.038611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:15:56.218052Z digest=sha256:a9ace6220797309c5d2a15185a0a877350289e338b3df32e3d144f62be32f28a

Observation 3942f902-e29f-430b-82eb-ffb37e5f278b · outbound

This paper cites Multi-scale Speaker Diarization with Dynamic Scale Weighting.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Multi-scale Speaker Diarization with Dynamic Scale Weighting

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:15:58.855167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:15:56.335846Z digest=sha256:4806bcc3eb82aa00a666fd93f6e86b2260eb1ee066684657f57caacc5e0da7d6

Observation af58f791-1946-4fc6-ae89-d2669e4cb1e4 · outbound

This paper cites an unresolved cited work.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Unresolved cited work

Reference 18

Resolution
verified exact
doi, observed 2026-08-07T04:15:57.370849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:15:56.441327Z digest=sha256:ed1afd9b24217ed1ecc98ecf8c281bdcc9f06332190c374e753d0d48926236b4

Observation 301c54e8-c851-4259-9ddd-bfd8ba67088e · outbound

This paper cites Pedregosa, G.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Pedregosa, G

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:56.524243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:15:56.524243Z digest=sha256:17c21d24f0bb6706a730baf34b7b17cb1a8a7c98b54c72e065aa12d750cec7fe

Observation f58fe505-57b4-4c36-866e-44a38872b049 · outbound

This paper cites an unresolved cited work.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:15:59.558334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:15:56.604502Z digest=sha256:0522d482b510d64f13bbcbb988bd1da914172bd188648b60c3270bb9f1c88bb7

Observation 5092846e-ce14-438f-b405-42cea1176d02 · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Robust Speech Recognition via Large-Scale Weak Supervision

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:56.689663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:15:56.689663Z digest=sha256:67ad2e434da02d7ed57136e6c83989278bcfb2c7d0a3010ea8c6d5c177b72165

Observation b064a4f8-1330-442d-a145-63034f2a7c02 · outbound

This paper cites an unresolved cited work.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:56.769223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:15:56.769223Z digest=sha256:10cc0eafb27894735e544f532769fe29acc4bb0c63904681d27f64a7a5f0110b

Observation 80d5c565-14c0-4f84-9a26-578fc0102d5e · outbound

This paper cites Joint Speech Recognition and Speaker Diarization via Sequence Transduction.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Joint Speech Recognition and Speaker Diarization via Sequence Transduction

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:56.818718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:15:56.818718Z digest=sha256:526527e939d5f465e496a401a0a131bc2502ac707444d57430cff82a144119ad

Observation 5c3e1fbd-692c-4af9-b0eb-b9c200824229 · outbound

This paper cites an unresolved cited work.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:56.910787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:15:56.910787Z digest=sha256:7ee362de927bf98442b342528443ae21ea0ecb79df56be18a87c7475c09726ca

Observation bb8a8c06-2f04-4733-acc7-4a8a7c0ad7f4 · outbound

This paper cites TOLD: A Novel Two-Stage Overlap-Aware Framework for Speaker Diarization.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models TOLD: A Novel Two-Stage Overlap-Aware Framework for Speaker Diarization

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:15:58.644634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:15:57.027068Z digest=sha256:cce401fd2b64e672c6ded56e09eebba8aa8bb3b091b4f4bb218f23c23fa367de

Observation ad09241f-898f-46c1-984a-091958cd9c27 · outbound

This paper cites Speaker Diarization with LSTM.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Speaker Diarization with LSTM

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:15:58.310770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:15:57.099168Z digest=sha256:393a7a693b71a39977c40d4dc8a1a21a74f02a80c3948bbcaa997bacff1dab48

Observation a53b12da-abca-4030-807c-aecd224d6911 · outbound

This paper cites DiarizationLM: Speaker Diarization Post-Processing with Large Language Models.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models DiarizationLM: Speaker Diarization Post-Processing with Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:57.190957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:15:57.190957Z digest=sha256:052ebad72771be5e33d12c07418fb680f446238e42cf23e7fd2beed4216589e5

Pith citing papers

No inbound Pith citation observations are available.