Pith. sign in

Paper Citation Record · LEDGER

Adapting Whisper for Streaming Speech Recognition via Two-Pass Decoding

As of 7 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 1 inbound Pith citation observation for arXiv:2506.12154.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.12154 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T01:07:25.099655Z

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T01:07:23.619613Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T01:07:25.289026Z

Reference resolution

21 of 21 outbound references displayed

  • verified exact0
  • verified fuzzy16
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4ac91c61-097d-4b65-bf6b-7fdf5d0c6e73 · outbound

This paper cites Whisper [1], released by OpenAI, exemplifies this trend.

Adapting Whisper for Streaming Speech Recognition via Two-Pass Decoding Whisper [1], released by OpenAI, exemplifies this trend

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:07:27.501454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T01:07:23.553893Z digest=sha256:8907977aa05d92dd309e5269aed5302bbc6fcdb196f01a849a15cb8cb046ea6b

Observation ca1e33d6-9bd4-42c1-93e6-b6fbe6232ffe · outbound

This paper cites Adapting Whisper for Streaming Speech Recognition via Two-Pass Decoding.

Adapting Whisper for Streaming Speech Recognition via Two-Pass Decoding Adapting Whisper for Streaming Speech Recognition via Two-Pass Decoding

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T01:07:25.341384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T01:07:23.619613Z digest=sha256:b4069d49895c0536b9b988e636d62db86c6b6b3f743b639cf82272fc9ed7a76a

Observation cb3484ec-5e6a-46bf-9f9c-13af8f5f6a02 · outbound

This paper cites We also considered Earnings-22 [16], but excluded it due to its lim- ited size, which is insufficient for training required by the ex- periments.

Adapting Whisper for Streaming Speech Recognition via Two-Pass Decoding We also considered Earnings-22 [16], but excluded it due to its lim- ited size, which is insufficient for training required by the ex- periments

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:07:27.485037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T01:07:23.671326Z digest=sha256:0ef992ce03da3bd66efc34e1ab3e151737516be33b196512952c6b2535b1dfb9

Observation 2bc45b24-e59c-4b40-bdbb-af6e3693695f · outbound

This paper cites $1.3 million.

Adapting Whisper for Streaming Speech Recognition via Two-Pass Decoding $1.3 million

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:07:27.467225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T01:07:23.736169Z digest=sha256:ca86ca707674b070817f239f17f29627eb4e3b85001913168349c86d9bc46d3c

Observation 37af1c70-e9d8-4370-92fb-071e1224f2e4 · outbound

This paper cites We intro- duced a hybrid tokenizer to improve generalization, especially when fine-tuning is performed with limited data.

Adapting Whisper for Streaming Speech Recognition via Two-Pass Decoding We intro- duced a hybrid tokenizer to improve generalization, especially when fine-tuning is performed with limited data

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:07:27.450773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T01:07:23.805663Z digest=sha256:f8a93a186207c8e812961f2dcc1ee458a303d180c1aaccd8b07b8f0df75a76a0

Observation c0cfde51-a754-4468-bedb-2bb38f4fc313 · outbound

This paper cites Robust speech recognition via large-scale weak su- pervision,.

Adapting Whisper for Streaming Speech Recognition via Two-Pass Decoding Robust speech recognition via large-scale weak su- pervision,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:07:27.435203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T01:07:23.871889Z digest=sha256:f79f833440a452692d5d11b77fecd6d6f3f12720651a4ef62db05ea1d3e53e9a

Observation ad42da0d-2c0e-4722-bc13-c88c193411c7 · outbound

This paper cites Turning whisper into real- time transcription system,.

Adapting Whisper for Streaming Speech Recognition via Two-Pass Decoding Turning whisper into real- time transcription system,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:07:27.416034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T01:07:23.933963Z digest=sha256:bc282cf4cac6b9f435762f54432980fb7c55846166b2a268f792f8f5b0daa0ae

Observation 5dc41039-d442-4dfc-a5d1-37b9ca75ecb3 · outbound

This paper cites Simul-whisper: Attention-guided streaming whisper with truncation detection,.

Adapting Whisper for Streaming Speech Recognition via Two-Pass Decoding Simul-whisper: Attention-guided streaming whisper with truncation detection,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:07:27.399148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T01:07:23.999076Z digest=sha256:358f444d718a0645334f0ae0fff8009b3613639d81066d9d5d71ca693067acf5

Observation 36614f02-218a-4b57-a810-26350fb5315d · outbound

This paper cites Con- nectionist temporal classification: labelling unsegmented se- quence data with recurrent neural networks,.

Adapting Whisper for Streaming Speech Recognition via Two-Pass Decoding Con- nectionist temporal classification: labelling unsegmented se- quence data with recurrent neural networks,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:07:27.382396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T01:07:24.066277Z digest=sha256:7ede5b4c74a9f021852b6a344f3711a6ed26bc27b9374f29cff4b1b6b38da34f

Observation f7ebb03e-e715-48f1-b575-1b1185d539e9 · outbound

This paper cites Sequence transduction with recurrent neural net- works,.

Adapting Whisper for Streaming Speech Recognition via Two-Pass Decoding Sequence transduction with recurrent neural net- works,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:07:27.362807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T01:07:24.132653Z digest=sha256:a3eee89e6ed4b34b2fccbaaf66d3c3721638c9147d238531061dfa8ffe183037

Observation e798313b-24dc-4d09-bfe5-56d522e7283b · outbound

This paper cites Unified Streaming and Non-streaming Two-pass End-to-end Model for Speech Recognition.

Adapting Whisper for Streaming Speech Recognition via Two-Pass Decoding Unified Streaming and Non-streaming Two-pass End-to-end Model for Speech Recognition

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T01:07:24.172626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:07:24.172626Z digest=sha256:2fb257e7bd46c04b1c18c7a7f951316e572463a1cf0e2de3bcc29c1b8863a208

Observation 1cbdc094-966d-42ce-8ce1-9ddf74d3d91d · outbound

This paper cites Wenet: Production oriented streaming and non-streaming end-to-end speech recognition toolkit,.

Adapting Whisper for Streaming Speech Recognition via Two-Pass Decoding Wenet: Production oriented streaming and non-streaming end-to-end speech recognition toolkit,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:07:27.041215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T01:07:24.242363Z digest=sha256:a9626117d46477bdfa86f9a3e1d1346f3c2559d3a1a76532a20bc59c46060fe6

Observation 0972ce04-7659-4248-9178-9395c31c8d08 · outbound

This paper cites Wenet 2.0: More productive end- to-end speech recognition toolkit,.

Adapting Whisper for Streaming Speech Recognition via Two-Pass Decoding Wenet 2.0: More productive end- to-end speech recognition toolkit,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:07:26.767482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T01:07:24.324384Z digest=sha256:6c7f16d79f56976628f7273b32e56281ebd09357d1ed85793cb3c8abd8a9f429

Observation a0ca50c7-6981-499a-a693-009c1d3c72ad · outbound

This paper cites Neural machine translation of rare words with subword units,.

Adapting Whisper for Streaming Speech Recognition via Two-Pass Decoding Neural machine translation of rare words with subword units,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:07:26.571849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T01:07:24.474118Z digest=sha256:ba40495b4b88e0d937c2036dd8fc16911e5105379ac09c97370a77e8fe7dbfed

Observation 34ed571d-a5d4-425b-a286-00a1a555a31b · outbound

This paper cites Hy- brid ctc/attention architecture for end-to-end speech recognition,.

Adapting Whisper for Streaming Speech Recognition via Two-Pass Decoding Hy- brid ctc/attention architecture for end-to-end speech recognition,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:07:26.325589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T01:07:24.560953Z digest=sha256:782fcba4eace4a7d259b7acbc75238ab9ffa183275efee813f01297dbfba1899

Observation bb5c193f-16bd-449b-906e-f8e59c2449f7 · outbound

This paper cites A new algorithm for data compression,.

Adapting Whisper for Streaming Speech Recognition via Two-Pass Decoding A new algorithm for data compression,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T01:07:24.687921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:07:24.687921Z digest=sha256:6f9bd9fd6dea50ad27e90f9387375337439618cbd8e3d34f54eb41e5cc8fd5bd

Observation 0cbb3f2d-2aa3-49e8-acaa-8405047703a8 · outbound

This paper cites Language models are unsupervised multitask learn- ers,.

Adapting Whisper for Streaming Speech Recognition via Two-Pass Decoding Language models are unsupervised multitask learn- ers,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:07:26.070765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T01:07:24.695228Z digest=sha256:9e5699004a6b749c8498fe0d31f8f80838de3107d77d943dbbb84b7f8e1949c8

Observation cc6d7787-00e2-43c1-a57d-5325981190f2 · outbound

This paper cites SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing.

Adapting Whisper for Streaming Speech Recognition via Two-Pass Decoding SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T01:07:24.815381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:07:24.815381Z digest=sha256:7a82609590d06617e55ae80690a97163ec07e5ca685f89248bb1a160b3d10d53

Observation 1e1a8f9a-19c2-4b18-85ca-ec57da7d44d8 · outbound

This paper cites PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode Transformation and Graph Compilation,.

Adapting Whisper for Streaming Speech Recognition via Two-Pass Decoding PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode Transformation and Graph Compilation,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:07:25.891366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T01:07:24.919589Z digest=sha256:e1f340b185feb29f3926f38816fa5bbaf1f6dcceef115356bdd08d8aa449dac1

Observation aba90d9b-2c1d-438c-bc3b-32d36ca27b37 · outbound

This paper cites Lib- rispeech: an asr corpus based on public domain audio books,.

Adapting Whisper for Streaming Speech Recognition via Two-Pass Decoding Lib- rispeech: an asr corpus based on public domain audio books,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:07:25.645845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T01:07:24.997475Z digest=sha256:07d7ba391e11daeb1fa3d36bd4ae5b1585cf91d7924724e548874b4395fdbcfc

Observation 93dc0f56-3c94-4a0f-abbd-4311cbcc27ba · outbound

This paper cites Earnings-22: A Practical Benchmark for Accents in the Wild.

Adapting Whisper for Streaming Speech Recognition via Two-Pass Decoding Earnings-22: A Practical Benchmark for Accents in the Wild

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T01:07:25.099655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:07:25.099655Z digest=sha256:c69ea65f599100d032f3b5a55b794773c639b6bc63557e40f435bb85f1b254c1

Pith citing papers

Observation ca1e33d6-9bd4-42c1-93e6-b6fbe6232ffe · inbound

Adapting Whisper for Streaming Speech Recognition via Two-Pass Decoding cites this paper.

Adapting Whisper for Streaming Speech Recognition via Two-Pass Decoding Adapting Whisper for Streaming Speech Recognition via Two-Pass Decoding

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T01:07:25.341384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T01:07:23.619613Z digest=sha256:b4069d49895c0536b9b988e636d62db86c6b6b3f743b639cf82272fc9ed7a76a