Pith. sign in

Paper Citation Record · LEDGER

WhisperFlow: speech foundation models in real time

As of 14 August 2026, this Paper Citation Record lists 76 of 76 outbound references and 2 inbound Pith citation observations for arXiv:2412.11272.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.11272 v2

Coverage vector

measured 76 of 76 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T15:11:57.955428Z

measured 78 of 78 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-18T22:16:51.917336Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T22:21:53.649506Z

Reference resolution

76 of 76 outbound references displayed

  • verified exact1
  • verified fuzzy35
  • unresolved37
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e1fee2c3-e370-41a2-9439-93222cf0253f · outbound

This paper cites Accessed: 2024-11-3.

WhisperFlow: speech foundation models in real time Accessed: 2024-11-3

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.936833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T15:11:57.731021Z digest=sha256:58813741f5b1c9a01adb5069425ea3d304daa428503249959e93d8e8d4a91fd1

Observation dc2052ab-9b74-4902-b738-bf19f9a44b52 · outbound

This paper cites GPT-4 Technical Report.

WhisperFlow: speech foundation models in real time GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.735305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.735305Z digest=sha256:a6198446a25afef2c7ea14b27666f376278102067547bd27e0573fc8a138418a

Observation a17cf4bb-74a5-464f-ba1c-3fa280105781 · outbound

This paper cites Did you hear that? Adversarial Examples Against Automatic Speech Recognition.

WhisperFlow: speech foundation models in real time Did you hear that? Adversarial Examples Against Automatic Speech Recognition

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.738385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.738385Z digest=sha256:d694cd7b9b7198f95984a6a98f022f62521c566635f37363c6289581223b7dc3

Observation 30577563-e2bb-4df3-8748-a55b7caaaadf · outbound

This paper cites Apple macbook air tech specs, 2024.

WhisperFlow: speech foundation models in real time Apple macbook air tech specs, 2024

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.926431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T15:11:57.742310Z digest=sha256:0ba23d31045e61b02ee0d2982389c132274b233cd6e28ee239b0328d8a029299

Observation 7ca7c6ef-9f0b-4613-9aee-f4876768af59 · outbound

This paper cites Neural Machine Translation by Jointly Learning to Align and Translate.

WhisperFlow: speech foundation models in real time Neural Machine Translation by Jointly Learning to Align and Translate

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.745362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.745362Z digest=sha256:86e16d39099ce68bb0df7e01a1d0fb3a26973accd4f701aed5af41c9a2c6c1bc

Observation 246f45ca-6e4b-40d1-b74b-08e9cbebf41d · outbound

This paper cites an unresolved cited work.

WhisperFlow: speech foundation models in real time Unresolved cited work

Reference 6

Resolution
metadata mismatch
raw_fallback, observed 2026-08-11T15:11:58.596857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T15:11:57.748694Z digest=sha256:f1d6f8a4c03a3fd250837c0a75225d8bf414296e77cf4f1c4af91ca0493735f0

Observation 62593c14-0bf6-46d0-b9da-25d63e01403c · outbound

This paper cites A mathematical theory of adaptive control processes.

WhisperFlow: speech foundation models in real time A mathematical theory of adaptive control processes

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.751674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.751674Z digest=sha256:eeb07dd850bd2d9053409f6f11658158bdb9cc4f14389ab6ff2fa686b0e87a78

Observation 090f84b0-c612-4419-84ef-3452f7ae6132 · outbound

This paper cites Speech recognition for clinical documentation from 1990 to 2018: a systematic review.

WhisperFlow: speech foundation models in real time Speech recognition for clinical documentation from 1990 to 2018: a systematic review

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.913193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T15:11:57.755262Z digest=sha256:53776757127bae3e2f6160da6e17fd476a60a6d42675d2c6303b39eebcbf813e

Observation 494cbbb8-5a55-4348-8ca2-80fa349759ec · outbound

This paper cites Language models are few-shot learners.

WhisperFlow: speech foundation models in real time Language models are few-shot learners

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.905237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T15:11:57.757951Z digest=sha256:0704bd2acf5417bcdaa604286f6ebc9a2abd1f68a7c4eca186a06380bf27c3b2

Observation 094d6cd5-e6f3-48a0-adb3-52eb0e1797be · outbound

This paper cites Audio adversarial examples: Targeted attacks on speech-to-text.

WhisperFlow: speech foundation models in real time Audio adversarial examples: Targeted attacks on speech-to-text

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.896876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T15:11:57.760645Z digest=sha256:cd3a2efaf046849397fbac6a536c78c9d4296eafed8016525c1581780631c606

Observation 725460a3-7847-4591-949e-10a6cb67273c · outbound

This paper cites https://github.com/corsix/amx, 2022.

WhisperFlow: speech foundation models in real time https://github.com/corsix/amx, 2022

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.888962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T15:11:57.763351Z digest=sha256:246b0bd3f916f65c067b634a21202930599b08e9f7d6a3b1ab010dc1b2e1876d

Observation 3877e876-ef45-4874-afa7-83e9d4a5c606 · outbound

This paper cites In 29th USENIX Security Symposium (USENIX Security 20) , pages 2667–2684, 2020.

WhisperFlow: speech foundation models in real time In 29th USENIX Security Symposium (USENIX Security 20) , pages 2667–2684, 2020

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.881328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T15:11:57.766015Z digest=sha256:c338bf4b34f0a6a6cf20e56155c041275db92c945df7da79f31faa95d2cd8497

Observation 1e62c57e-2122-4f3f-a686-70b5754b17a6 · outbound

This paper cites Fleurs: Few-shot learning evaluation of universal representations of speech.

WhisperFlow: speech foundation models in real time Fleurs: Few-shot learning evaluation of universal representations of speech

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.768592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.768592Z digest=sha256:8ebad7c0c6204461972f31156370e056e69d8800603074982e7772e140fb5127

Observation 50cbdccf-acf2-4b36-9cee-f03a241e06f0 · outbound

This paper cites Automatic recognition of spoken digits.

WhisperFlow: speech foundation models in real time Automatic recognition of spoken digits

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.869693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T15:11:57.771062Z digest=sha256:1f9b87c9b95f1bf00daa0d18bf9e9215c526954d610d0db9de16a1f6ff56ae9d

Observation e23ba16d-0950-4668-a35b-636f32acc0eb · outbound

This paper cites Bert: Pre- training of deep bidirectional transformers for language understanding.

WhisperFlow: speech foundation models in real time Bert: Pre- training of deep bidirectional transformers for language understanding

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.862092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T15:11:57.774604Z digest=sha256:4557cf310f7f1e0099f9a2038f837f9932e5384dc832506f0147e8ba71376dd0

Observation 2040508e-5056-4dd7-9985-87d990395ede · outbound

This paper cites Speculative decoding for 2x faster whisper inference, 2023.

WhisperFlow: speech foundation models in real time Speculative decoding for 2x faster whisper inference, 2023

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.854362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T15:11:57.777321Z digest=sha256:d1be7f7a11b150103d23f61a8b10968c844dc0849e97c742f42f22d839357755

Observation 1ac83dde-2c25-45d3-b37c-06b8261f611e · outbound

This paper cites ggerganov/llama.cpp, 2022.

WhisperFlow: speech foundation models in real time ggerganov/llama.cpp, 2022

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.845564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T15:11:57.780439Z digest=sha256:aa9dad6272d840132c57983d930cafb0dbef221e463a7d4f14b001febffdb233

Observation 26fb2ca9-a17f-4164-a9f1-041b2550ceb3 · outbound

This paper cites ggerganov/whisper.cpp, 2022.

WhisperFlow: speech foundation models in real time ggerganov/whisper.cpp, 2022

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.838523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T15:11:57.783040Z digest=sha256:65f00ea81e7efb133c0498e90cd7f6cba0548ac75cf81cbcd2251ddb3197b889

Observation e82e09b5-dc9e-45b2-9b86-ac3f1a5d35df · outbound

This paper cites Sequence Transduction with Recurrent Neural Networks.

WhisperFlow: speech foundation models in real time Sequence Transduction with Recurrent Neural Networks

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.785464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.785464Z digest=sha256:79d31565b959c8deb5fa1da0982d1a1962077217e876f133736b362cbe8bcb45

Observation 468ba433-b7d5-4625-9837-42621a8a68e5 · outbound

This paper cites Conformer: Convolution-augmented Transformer for Speech Recognition.

WhisperFlow: speech foundation models in real time Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.788464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.788464Z digest=sha256:6b30e1c7ec1e77856b0c106d03059a60c932bd63398f3cfaa96731914041fc4b

Observation bef0cb85-d03c-4b74-af88-aaaacca45f7f · outbound

This paper cites MLX: Efficient and flexible machine learning on apple silicon, 2023.

WhisperFlow: speech foundation models in real time MLX: Efficient and flexible machine learning on apple silicon, 2023

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.829653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T15:11:57.791179Z digest=sha256:f68988fac7317889854eb5603c8178bf07de5d88114ae516aa3cb493e906cf9a

Observation 645054e9-16f2-4b42-9754-79d1b3eac5a9 · outbound

This paper cites Speech understanding systems: Summary of results of the five-year research effort, 1976.

WhisperFlow: speech foundation models in real time Speech understanding systems: Summary of results of the five-year research effort, 1976

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.821807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T15:11:57.793723Z digest=sha256:4b8664a2bae66c4b56d1c0cd559df4e41477588166799687130a38fd18e3f692

Observation 20530a75-8258-4a5c-9c6a-e56172710f77 · outbound

This paper cites Streaming end-to-end speech recognition for mobile devices.

WhisperFlow: speech foundation models in real time Streaming end-to-end speech recognition for mobile devices

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.814296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T15:11:57.796197Z digest=sha256:1ae02d80cd47c3f136b41462f96b6ae4b74152e27f01e40c7d43764402c16bc1

Observation 2344ecb9-9d73-415c-a5be-0642d983e51b · outbound

This paper cites Ted-lium 3: Twice as much data and corpus repartition for experiments on speaker adaptation.

WhisperFlow: speech foundation models in real time Ted-lium 3: Twice as much data and corpus repartition for experiments on speaker adaptation

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.805651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T15:11:57.798672Z digest=sha256:6b1c1d108e130199eb91afc360a0e78cfbfd989f9a9b9ec337007548ff81c77e

Observation c88a3b4b-3859-48d0-9a3c-40afca78a3ff · outbound

This paper cites Advances in Joint CTC-Attention based End-to-End Speech Recognition with a Deep CNN Encoder and RNN-LM.

WhisperFlow: speech foundation models in real time Advances in Joint CTC-Attention based End-to-End Speech Recognition with a Deep CNN Encoder and RNN-LM

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.801247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.801247Z digest=sha256:54604521ef20634cc89f1212f7df7ff4d8d412bdc24f0c08b378d490237c617d

Observation e55ac934-d788-4dcf-988d-6884d4947301 · outbound

This paper cites Swapadvisor: Pushing deep learning beyond the gpu memory limit via smart swapping.

WhisperFlow: speech foundation models in real time Swapadvisor: Pushing deep learning beyond the gpu memory limit via smart swapping

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.804177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.804177Z digest=sha256:94f010a89285afb345a283415e7f03f93db8b05dd476fda13c5550f464257e97

Observation 9401c046-5b23-430e-af94-ab6774ffb69f · outbound

This paper cites Deepum: Tensor migration and prefetching in unified memory.

WhisperFlow: speech foundation models in real time Deepum: Tensor migration and prefetching in unified memory

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.806689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.806689Z digest=sha256:d9fb1f2ddf116884b495e3e1387ccf5982c789bce21cb73e06f9a80dcee92f0c

Observation 091c9561-37b9-4a0e-90de-d888f886085b · outbound

This paper cites Speech and language processing, 2000.

WhisperFlow: speech foundation models in real time Speech and language processing, 2000

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.797821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T15:11:57.809268Z digest=sha256:61cd10dc4d61db4a432eb51492cb7b0a75c340ed66ea07fbeb5a778c1c173979

Observation 2a4a9b6a-28c3-4fe8-bff0-2a03ec80140a · outbound

This paper cites Large-Scale Multilingual Speech Recognition with a Streaming End-to-End Model.

WhisperFlow: speech foundation models in real time Large-Scale Multilingual Speech Recognition with a Streaming End-to-End Model

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.811984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.811984Z digest=sha256:e218af8e947634129ad986a6b5f8c7c1af331e3b476c24c56832d8a5a45b966b

Observation af25c6a9-e662-4d0c-b7ab-f308a048cb80 · outbound

This paper cites Scaling Laws for Neural Language Models.

WhisperFlow: speech foundation models in real time Scaling Laws for Neural Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.815931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.815931Z digest=sha256:e7789c104ede6a73538dfe45f449fed4494db8093f9e4c14a4d34241b61a8d57

Observation a8b0faf0-90df-446a-9a85-f6e0d22b16b4 · outbound

This paper cites Convolution-augmented parameter-efficient fine-tuning for speech recognition.

WhisperFlow: speech foundation models in real time Convolution-augmented parameter-efficient fine-tuning for speech recognition

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.789991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T15:11:57.819593Z digest=sha256:2eff9c867f67c3f88623056bf1b32b4c813f0cd9d65c80b895785779d6cc9655

Observation 8133820a-9a97-48b7-a7b1-341eb3ebde10 · outbound

This paper cites Speculative Decoding with Big Little Decoder.

WhisperFlow: speech foundation models in real time Speculative Decoding with Big Little Decoder

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.822034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.822034Z digest=sha256:ae9851c26f3a689d72f1312b83d1309ac6524af95b207137ead45ea647e2ab74

Observation 3cbfb0bb-9383-4450-a179-4372dfc16d32 · outbound

This paper cites Low-latency sequence-to- sequence speech recognition and translation by partial hypothesis selection.

WhisperFlow: speech foundation models in real time Low-latency sequence-to- sequence speech recognition and translation by partial hypothesis selection

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.781953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T15:11:57.825115Z digest=sha256:318313ffde8ea39c827a6c25a71969a7bc4a76377ff6606e99ea119cecdc3d5b

Observation c0e2a732-22d5-4cbc-95cf-eebe0733d498 · outbound

This paper cites Turning Whisper into Real-Time Transcription System.

WhisperFlow: speech foundation models in real time Turning Whisper into Real-Time Transcription System

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.828381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.828381Z digest=sha256:ab15efe5497ee8976e53cd3b09db183ae429b2fa0bdfb685da0afe7435e71efa

Observation e04c2db0-c5c3-4d73-96f9-8ca1ec3b96b6 · outbound

This paper cites Streaming automatic speech recog- nition with the transformer model.

WhisperFlow: speech foundation models in real time Streaming automatic speech recog- nition with the transformer model

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.772640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T15:11:57.831725Z digest=sha256:9eb123d16e9dece0590aa239f7ce26438ccb1793b6290541e6d2877ea1581025

Observation 76335910-b08d-4256-9697-772306a91042 · outbound

This paper cites Universal Adversarial Perturbations for Speech Recognition Systems.

WhisperFlow: speech foundation models in real time Universal Adversarial Perturbations for Speech Recognition Systems

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.834442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.834442Z digest=sha256:fe3ce7c496be05d03411a2b4a668aa18aab31511ae6d1c7e795d4cd550083127

Observation 00224c54-9fb2-4696-b3fd-d2add7375363 · outbound

This paper cites There is more than one kind of robustness: Fooling Whisper with adversarial examples.

WhisperFlow: speech foundation models in real time There is more than one kind of robustness: Fooling Whisper with adversarial examples

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.837979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.837979Z digest=sha256:b5e5a3dc74deadb5273ceeb358f2feec9d87af2eb4655de27782f197de29434a

Observation b053425f-17b8-482c-aebc-27494d4481a5 · outbound

This paper cites Train- ing language models to follow instructions with human feedback.

WhisperFlow: speech foundation models in real time Train- ing language models to follow instructions with human feedback

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.764019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T15:11:57.841017Z digest=sha256:7133c98bdef77aa12764e9f5f81738a76ceb6e7dae5ae7af72dd14dc5c804826

Observation f28c8d69-25ac-4a81-a9ed-dec47100c045 · outbound

This paper cites Lib- rispeech: an asr corpus based on public domain audio books.

WhisperFlow: speech foundation models in real time Lib- rispeech: an asr corpus based on public domain audio books

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.756208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T15:11:57.843498Z digest=sha256:59a58e240e8e278ce5b11408a543ccb99337255303de6d42e48a2799a0d1225a

Observation 6b295caa-a991-45bb-a78a-0597aaed94e0 · outbound

This paper cites Splitwise: Efficient generative llm inference using phase splitting.

WhisperFlow: speech foundation models in real time Splitwise: Efficient generative llm inference using phase splitting

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.846091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.846091Z digest=sha256:80380d710391ca93714e6b597714f5cbad9d3e8102e6fe4f435a8edf03ada3f0

Observation 698b7f08-3809-4cb9-aba8-0c5b0a5ee70a · outbound

This paper cites Branchformer: Parallel mlp-attention architectures to capture local and global context for speech recognition and understanding.

WhisperFlow: speech foundation models in real time Branchformer: Parallel mlp-attention architectures to capture local and global context for speech recognition and understanding

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.742539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T15:11:57.848604Z digest=sha256:8c858be00ac1222e35d27cfd323cb74d7e9209c6e406a4123f118394965e1be3

Observation b75de65d-c04d-42fb-9686-e1c49c643009 · outbound

This paper cites OWSM-CTC: An Open Encoder-Only Speech Foundation Model for Speech Recognition, Translation, and Language Identification.

WhisperFlow: speech foundation models in real time OWSM-CTC: An Open Encoder-Only Speech Foundation Model for Speech Recognition, Translation, and Language Identification

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.851647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.851647Z digest=sha256:8db4f1466c8d5c7bc83b0e9f9ef37c7a8fc28a0b43a76c4654ecbe92b95779a3

Observation 138f80b3-879a-439e-b263-28b3ba8a0e7f · outbound

This paper cites OWSM v3.1: Better and Faster Open Whisper-Style Speech Models based on E-Branchformer.

WhisperFlow: speech foundation models in real time OWSM v3.1: Better and Faster Open Whisper-Style Speech Models based on E-Branchformer

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.855407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.855407Z digest=sha256:88d48560c4a5a3e024d71e972bc36e4ebc55f352f64eaab227dfa93a62f4704d

Observation 98f7d38c-4830-41c1-913c-6c03fff58770 · outbound

This paper cites Reproducing whisper-style training using an open-source toolkit and publicly available data.

WhisperFlow: speech foundation models in real time Reproducing whisper-style training using an open-source toolkit and publicly available data

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.733664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T15:11:57.858375Z digest=sha256:b16388cb8e03aa003d43abd9be3cb60cf4b82199e5e62595fb28e527f4c80b0e

Observation d94484bb-c132-4fd9-b332-6f766f875476 · outbound

This paper cites Speech percep- tion at the interface of neurobiology and linguistics.

WhisperFlow: speech foundation models in real time Speech percep- tion at the interface of neurobiology and linguistics

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.724628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T15:11:57.860881Z digest=sha256:a95a4d144ddd3010e8f673e2a4da0930391d708c5a3546e330e35d6783eee71d

Observation 251e98f7-4a70-411e-b3c8-3de52bddfbf9 · outbound

This paper cites Improving language understanding by generative pre-training.

WhisperFlow: speech foundation models in real time Improving language understanding by generative pre-training

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.863590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.863590Z digest=sha256:a755b6e0efef9b9d394841a4cf97223134b7cc4a34be50a3f414dbf0796966a2

Observation 6cb8518c-9b3a-4ebd-81fb-8dd3c92faf3f · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

WhisperFlow: speech foundation models in real time Robust speech recognition via large-scale weak supervision

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.866274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.866274Z digest=sha256:541597458270741e3e7ca0dbb5d624a7bc83f1a6872a5646d5b6c4ccd2930863

Observation af9b2efb-3254-4f00-91fc-7315dbf3a9db · outbound

This paper cites Language models are unsupervised multitask learners.

WhisperFlow: speech foundation models in real time Language models are unsupervised multitask learners

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.868942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.868942Z digest=sha256:74d2ed600cb4f9f3d3298d68892c5b62b6d999baddc4246c022ae89869cdfd7b

Observation 07330927-76d8-4dae-a04f-5db36bfcee3a · outbound

This paper cites Controlling Whisper: Universal Acoustic Adversarial Attacks to Control Speech Foundation Models.

WhisperFlow: speech foundation models in real time Controlling Whisper: Universal Acoustic Adversarial Attacks to Control Speech Foundation Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.874444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.874444Z digest=sha256:4d63d51d300d3e914e658ea99fe5eb59e4f65faa200651c3a4b37aff77bec7f8

Observation 392bc3de-f999-4476-b06d-b3ef08c5a579 · outbound

This paper cites Muting whisper: A universal acoustic adversarial attack on speech foundation models.

WhisperFlow: speech foundation models in real time Muting whisper: A universal acoustic adversarial attack on speech foundation models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.877075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.877075Z digest=sha256:9414e88df4923811c57e9e9af809b4728aad2156efc1309cb62af451999f0e67

Observation 006ed930-ce33-42b2-9c7b-ad34ce27e3dc · outbound

This paper cites Paul Robinson, and Bradley S.

WhisperFlow: speech foundation models in real time Paul Robinson, and Bradley S

Reference 52

Resolution
verified exact
doi, observed 2026-08-11T15:11:57.982976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T15:11:57.879758Z digest=sha256:9b4eedd282ebe279d3b4032ba4789727c3fea026c12f8d8c90ace7e412f256e9

Observation c359b220-7008-4c8f-be90-4b7f55c4234c · outbound

This paper cites Exploring architectures, data and units for streaming end-to-end speech recognition with rnn-transducer.

WhisperFlow: speech foundation models in real time Exploring architectures, data and units for streaming end-to-end speech recognition with rnn-transducer

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.704061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T15:11:57.882478Z digest=sha256:1fe5522530a0e224a06b0a41ede277ded2c219369a861ea9ff811487a081f2e9

Observation 93db44f8-3349-4a37-86e6-d80025cece7c · outbound

This paper cites ZeRO-Offload: De- mocratizing Billion-Scale model training.

WhisperFlow: speech foundation models in real time ZeRO-Offload: De- mocratizing Billion-Scale model training

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.695451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T15:11:57.885034Z digest=sha256:262665803776daec1fd385a605a2ad0f13f988b9bb9150bf0b0b6603c8419d66

Observation 1a6f2e7e-3144-4ebd-98a7-85a1084640b2 · outbound

This paper cites Enhancing the ted-lium corpus with selected data for language modeling and more ted talks.

WhisperFlow: speech foundation models in real time Enhancing the ted-lium corpus with selected data for language modeling and more ted talks

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.687304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T15:11:57.887689Z digest=sha256:5f82bd1247db45bbe329281ffb3247cd7d19759ea4f2b022fa76af08a95704d9

Observation 34169d80-0d54-47e2-a1af-2b5148244e3a · outbound

This paper cites Adversarial Attacks Against Automatic Speech Recognition Systems via Psychoacoustic Hiding.

WhisperFlow: speech foundation models in real time Adversarial Attacks Against Automatic Speech Recognition Systems via Psychoacoustic Hiding

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.890946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.890946Z digest=sha256:d0bf4a154317113041cadfcbf347805fafc139ffb749b0230b6e8806f9365440

Observation 7eb371e8-3569-4b44-93e6-ce397c7d49fe · outbound

This paper cites Review of speech-to-text recognition technology for enhancing learning.

WhisperFlow: speech foundation models in real time Review of speech-to-text recognition technology for enhancing learning

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.679625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T15:11:57.893794Z digest=sha256:698cf6d0a659ba910c453b4764bdd3e24f0caae69db5ab859a1da736d574e7f1

Observation ff1dba35-b561-490c-8c76-5ade1e72c3a1 · outbound

This paper cites Dissecting User-Perceived Latency of On-Device E2E Speech Recognition.

WhisperFlow: speech foundation models in real time Dissecting User-Perceived Latency of On-Device E2E Speech Recognition

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.896422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.896422Z digest=sha256:3397702a15fc3b33166806848879a1d95d3bc0a1938d1c9a3dc682e06fb082be

Observation ada30cfe-0565-4506-b4a1-d835bfc7a6b6 · outbound

This paper cites PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU.

WhisperFlow: speech foundation models in real time PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.899379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.899379Z digest=sha256:6d7a52a87e731e858f8fb32541d95d68bd138a07e0c5e849da7c1be83465e56f

Observation b63a376d-5d78-4531-94ae-2497a4566db1 · outbound

This paper cites Instantaneous Grammatical Error Correction with Shallow Aggressive Decoding.

WhisperFlow: speech foundation models in real time Instantaneous Grammatical Error Correction with Shallow Aggressive Decoding

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.904845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.904845Z digest=sha256:b196cc277f544e53e71129f69e859c14754d24e81948fd89dd332e9c38df46a3

Observation 73da8f7c-9092-44f6-8757-4afae33bd75c · outbound

This paper cites Intriguing properties of neural networks.

WhisperFlow: speech foundation models in real time Intriguing properties of neural networks

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.907702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.907702Z digest=sha256:6e527951d659ceec1dda94883110ceccee4efa14783a7d558f24338249c71db2

Observation 4fff794c-ac5e-42d5-b128-04ba5146e1a8 · outbound

This paper cites Streaming trans- former asr with blockwise synchronous beam search.

WhisperFlow: speech foundation models in real time Streaming trans- former asr with blockwise synchronous beam search

Reference 63

Resolution
malformed identifier
no resolver link, observed 2026-08-11T15:11:57.910917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.910917Z digest=sha256:58294369e34e066bb004fc6a2b35ccd5cb4761e223c4618a84a8bd0c3b87bada

Observation 991113f4-2e3a-40c0-b0c3-38b0e68e9f90 · outbound

This paper cites The calo meeting speech recognition and understanding system.

WhisperFlow: speech foundation models in real time The calo meeting speech recognition and understanding system

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.671050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T15:11:57.913485Z digest=sha256:99effd3ec0734a63065cc52f538548c25b41f734245010716c13db1f941af15d

Observation 3dc58165-3aab-40f6-a3b2-0a83fc81a30f · outbound

This paper cites Attention is all you need.

WhisperFlow: speech foundation models in real time Attention is all you need

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.916319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.916319Z digest=sha256:3573c55591567304e662947c1f135f54ad5e9d0297870db5194381df92509047

Observation d6fff2f0-5bbf-43a4-b963-9db803f49ac5 · outbound

This paper cites Simul-Whisper: Attention-Guided Streaming Whisper with Truncation Detection.

WhisperFlow: speech foundation models in real time Simul-Whisper: Attention-Guided Streaming Whisper with Truncation Detection

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.919068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.919068Z digest=sha256:f28f4be8473cdddba99cca689b15e2f070dbf7f7a045b53a59a041a558390375

Observation 03387295-7265-473b-821a-24b96bc67d39 · outbound

This paper cites Turbocharge speech understanding with pilot inference.

WhisperFlow: speech foundation models in real time Turbocharge speech understanding with pilot inference

Reference 67

Resolution
malformed identifier
no resolver link, observed 2026-08-11T15:11:57.921997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.921997Z digest=sha256:5aa75d17945db342444a9ff18a9de38ce224f2046e9a6029f27276ef2e7bf9f8

Observation 03455b10-86cd-4b1a-ad0f-9a48a940dd5f · outbound

This paper cites ESPnet: End-to- end speech processing toolkit.

WhisperFlow: speech foundation models in real time ESPnet: End-to- end speech processing toolkit

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.658540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T15:11:57.924604Z digest=sha256:76bea65b71e200ac67b17f0da7d6ffcbf9dd94ca9bda486a71855a0d6449ceed

Observation 42c84d03-9e8b-4e7b-a97f-b066a9915e7a · outbound

This paper cites Loongserve: Efficiently serving long-context large language models with elastic sequence parallelism.

WhisperFlow: speech foundation models in real time Loongserve: Efficiently serving long-context large language models with elastic sequence parallelism

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.649972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T15:11:57.930867Z digest=sha256:4a7cc6301bc8505fc14359dd5c66563dd6a448de384e38608c03d9daab0a7807

Observation f28cfe96-0ce5-4b57-be1b-aa55c22217d2 · outbound

This paper cites Speculative Decoding: Exploiting Speculative Execution for Accelerating Seq2seq Generation.

WhisperFlow: speech foundation models in real time Speculative Decoding: Exploiting Speculative Execution for Accelerating Seq2seq Generation

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.936653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.936653Z digest=sha256:e8b8a1b1dfd4a12e19224e1feae1ff7d51bfb0f6728661b07d88d60ccb1f123c

Observation 5a5b0c0d-2af8-4d98-8484-c827f9c37b3f · outbound

This paper cites Toward human parity in con- versational speech recognition.

WhisperFlow: speech foundation models in real time Toward human parity in con- versational speech recognition

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.641768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T15:11:57.940291Z digest=sha256:b148a98dc7e44b3a4a132a749066b412a9d447bef25da3b7fc594da764316e68

Observation 0f39021f-fe3e-452b-af6b-d3d1fd2dd17c · outbound

This paper cites Fast On-device LLM Inference with NPUs.

WhisperFlow: speech foundation models in real time Fast On-device LLM Inference with NPUs

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.943619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.943619Z digest=sha256:570507591e9e874ad0e799bc094964e3f0c12d81bfb3599c211df4b9e1f89891

Observation ad35efc5-74ed-4041-9ed5-32b43527ef68 · outbound

This paper cites Inference with Reference: Lossless Acceleration of Large Language Models.

WhisperFlow: speech foundation models in real time Inference with Reference: Lossless Acceleration of Large Language Models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.946433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.946433Z digest=sha256:0e819bcb4f4267419d30152824336557ec2f4212360d39bebb8972c17d5d9441

Observation ad388ebc-0738-4f51-a6b9-13e26e4f2cbe · outbound

This paper cites Transformer transducer: A streamable speech recog- nition model with transformer encoders and rnn-t loss.

WhisperFlow: speech foundation models in real time Transformer transducer: A streamable speech recog- nition model with transformer encoders and rnn-t loss

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.949322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.949322Z digest=sha256:71140ccb9134d8e95c39336cbdc6f9aec41902695b27f3c6d1df09c9beaf59c1

Observation 25c14384-0ad6-4aa9-bafd-b20256e9b58d · outbound

This paper cites Black-box adversarial attacks on commercial speech platforms with minimal information.

WhisperFlow: speech foundation models in real time Black-box adversarial attacks on commercial speech platforms with minimal information

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.633008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T15:11:57.951868Z digest=sha256:b267efb9f08353a27d32c080a2ee239cecfb321bce9245fe58b25b1d3ba3d5f2

Observation 85c56d45-b835-4286-8d99-be546aa2d969 · outbound

This paper cites In 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24) , pages 193–210, 2024.

WhisperFlow: speech foundation models in real time In 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24) , pages 193–210, 2024

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.624249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T15:11:57.955428Z digest=sha256:1746ca4b29a25172cf4f4f41f1294f49ee7ffd59806e1893064dce5c15cf9814

Observation 6a3efb1f-8404-416b-9c99-cd82ef5f4442 · outbound

This paper cites an unresolved cited work.

WhisperFlow: speech foundation models in real time Unresolved cited work

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.927426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.927426Z digest=sha256:4eb41c620bed7f3cee0414664c0c007714b19939ba52db808e7d04a6811ab1c1

Observation d164851d-0a5a-4698-92eb-9e96d9578ee8 · outbound

This paper cites doi:10.1145/3694715.3695948.

WhisperFlow: speech foundation models in real time doi:10.1145/3694715.3695948

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.933707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.933707Z digest=sha256:6ff53f5549bcbcd2d412612ad572379ba10c2e998ad41f1b4424fd9920703439

Pith citing papers

Observation 818eced5-2f8e-4812-9425-a9adefa861aa · inbound

WhisperRT -- Turning Whisper into a Causal Streaming Model cites this paper.

WhisperRT -- Turning Whisper into a Causal Streaming Model WhisperFlow: speech foundation models in real time

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-18T22:21:53.652242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T22:16:51.917336Z digest=sha256:a7492ac0d7bd7279fa0381bc213a7af54df5fcdb0a3e996ca3a8536b3f24c6b9

Observation 4440a6c6-e16d-4a46-98e2-82eba009ed46 · inbound

Sink or SWIM: Tackling Real-Time ASR at Scale cites this paper.

Sink or SWIM: Tackling Real-Time ASR at Scale WhisperFlow: speech foundation models in real time

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:10:53.532913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-16T12:09:17.087477Z digest=sha256:fe7d36a342bc8f3ccb8c7b150f77284cf64c693aa99195816441f06700ac01c1