Pith. sign in

Paper Citation Record · LEDGER

Towards Maximum Likelihood Training for Transducer-based Streaming Speech Recognition

As of 13 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2411.17537.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.17537 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T12:13:11.939403Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact1
  • verified fuzzy21
  • unresolved5
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d1fb53bf-ccba-4b84-8587-837ac432bed3 · outbound

This paper cites Improving proper noun recog- nition in end-to-end asr by customization of the mwer loss criterion,.

Towards Maximum Likelihood Training for Transducer-based Streaming Speech Recognition Improving proper noun recog- nition in end-to-end asr by customization of the mwer loss criterion,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:13:12.383662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:13:11.816274Z digest=sha256:3e40a194d8412e3e8f84ff604a9f36bf00fab21d362e68daabb9a483d34819d0

Observation 171e2db0-00b3-48bc-b65f-c505cbef4ef0 · outbound

This paper cites Personalization of end-to-end speech recognition on mobile devices for named entities,.

Towards Maximum Likelihood Training for Transducer-based Streaming Speech Recognition Personalization of end-to-end speech recognition on mobile devices for named entities,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:13:12.367127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:13:11.821661Z digest=sha256:2dbf59d5a66f09fe3324afec658b5825ea80207ad5a7f89e74ae525b51c3877c

Observation b12b8d2b-9278-414c-878c-72d4e96a80d5 · outbound

This paper cites Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages.

Towards Maximum Likelihood Training for Transducer-based Streaming Speech Recognition Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T12:13:11.826465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:13:11.826465Z digest=sha256:10256297780cbfb936f6285efb844ffcb7809dde39497b312511ad8d33eb996f

Observation a600b9bd-bfdb-40d5-8f86-e50a9f9d4cc9 · outbound

This paper cites Scaling Speech Technology to 1,000+ Languages.

Towards Maximum Likelihood Training for Transducer-based Streaming Speech Recognition Scaling Speech Technology to 1,000+ Languages

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T12:13:11.831434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:13:11.831434Z digest=sha256:725c1e6b743221f80b928ede31de55dd31fbc6968f6bf310dd6ce8753d562522

Observation 7cad08d1-e2ab-45dd-aa9a-37f66686eb3a · outbound

This paper cites A better and faster end-to- end model for streaming asr,.

Towards Maximum Likelihood Training for Transducer-based Streaming Speech Recognition A better and faster end-to- end model for streaming asr,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:13:12.350590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:13:11.836718Z digest=sha256:e5be1b75b716ea3acbe55de436851ebf52021b09adfe754ba51e24191bce98d8

Observation 8d1623b1-1960-4911-9663-a6a033a65439 · outbound

This paper cites Cascaded encoders for unifying streaming and non-streaming asr,.

Towards Maximum Likelihood Training for Transducer-based Streaming Speech Recognition Cascaded encoders for unifying streaming and non-streaming asr,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:13:12.335630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:13:11.841431Z digest=sha256:0d32543bebb537d134ac4132af6c592f39706637b97482cad81f5517c5f9a1ef

Observation 6f2010c9-a873-4e92-b3ba-2c06ffbcbeff · outbound

This paper cites Dual-mode asr: Unify and improve streaming asr with full-context modeling,.

Towards Maximum Likelihood Training for Transducer-based Streaming Speech Recognition Dual-mode asr: Unify and improve streaming asr with full-context modeling,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:13:12.320471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:13:11.846582Z digest=sha256:b955ea1e2313ce78bcf90f7a929705d17c273fd10c89f73caccf519eef1f207a

Observation 5a344ed2-2b68-45de-9f4d-a90fc9b05648 · outbound

This paper cites Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups,.

Towards Maximum Likelihood Training for Transducer-based Streaming Speech Recognition Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T12:13:11.851129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:13:11.851129Z digest=sha256:9c1af09fc8d6b6acb5f0f2a1617b1ed9b83a7132e9b3b51172ee5797af00edfc

Observation b46cab54-a763-478d-a4e9-e98026def95a · outbound

This paper cites Connectionist temporal classification: labelling unsegmented sequence data with recur- rent neural networks,.

Towards Maximum Likelihood Training for Transducer-based Streaming Speech Recognition Connectionist temporal classification: labelling unsegmented sequence data with recur- rent neural networks,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:13:12.294940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:13:11.855609Z digest=sha256:c7b317254b0652a529611a7b9d66904bc849301dba6a5f6a08596e6dbba1fac7

Observation d4108838-eac7-4adb-9417-b30939bcd4d3 · outbound

This paper cites Sequence transduction with recurrent neural networks,.

Towards Maximum Likelihood Training for Transducer-based Streaming Speech Recognition Sequence transduction with recurrent neural networks,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:13:12.279887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:13:11.860027Z digest=sha256:05f0e1e38aa0599a3cbe670e698e0d13738577ba4d6bc435d1c1ce93ca901b93

Observation 0b875d82-d58c-4c7a-9d4b-53ce49212ab7 · outbound

This paper cites End-to-end attention-based large vocabulary speech recognition,.

Towards Maximum Likelihood Training for Transducer-based Streaming Speech Recognition End-to-end attention-based large vocabulary speech recognition,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:13:12.263759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:13:11.864442Z digest=sha256:af107e18371163b01a46cba79774debb5bd61c025d4778248e3e69cb9b09eba1

Observation 3e4a5c99-a609-41d8-b611-ac70dfbab78f · outbound

This paper cites Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,.

Towards Maximum Likelihood Training for Transducer-based Streaming Speech Recognition Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:13:12.249338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:13:11.870010Z digest=sha256:f79f7e881124e8af33eb53554b7fe04c9d25cead90e323f805d7202be355aea1

Observation 4d9306e7-1ed0-4a94-89ec-db00dc4286a2 · outbound

This paper cites CIF: Continuous Integrate-and-Fire for End-to-End Speech Recognition.

Towards Maximum Likelihood Training for Transducer-based Streaming Speech Recognition CIF: Continuous Integrate-and-Fire for End-to-End Speech Recognition

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-12T12:13:12.016770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:13:11.875107Z digest=sha256:ff187b7d89ea303eb0a3156c9ae242e50ece42cdf62d3ef86d7486f20eef5a9f

Observation 10e80c2f-73f4-4c60-b04a-aa19bcfc006b · outbound

This paper cites Crf-based single-stage acoustic modeling with ctc topology,.

Towards Maximum Likelihood Training for Transducer-based Streaming Speech Recognition Crf-based single-stage acoustic modeling with ctc topology,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:13:12.234353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:13:11.879782Z digest=sha256:12a0f89bf1cb74d4886f2596f7d3f1b4ccd36ce9db0ce5db4041581ccaca640e

Observation 5903ca51-3a82-4b88-a792-a8a449583128 · outbound

This paper cites Imputer: Sequence modelling via imputation and dynamic programming,.

Towards Maximum Likelihood Training for Transducer-based Streaming Speech Recognition Imputer: Sequence modelling via imputation and dynamic programming,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:13:12.219564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:13:11.884138Z digest=sha256:f03cd42efb806452975699b7ec7e18089786f0b4de738f0a9e86ba293589a199

Observation e9b2fdcd-caf4-44c9-880f-b99446e699a6 · outbound

This paper cites Global normalization for streaming speech recognition in a modular framework,.

Towards Maximum Likelihood Training for Transducer-based Streaming Speech Recognition Global normalization for streaming speech recognition in a modular framework,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:13:12.204049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:13:11.888434Z digest=sha256:9576541b8adc35bf3f731074800d934aa264290a22cc0eb84a6bf2e2f3f746b5

Observation c55bfd57-f5ec-470f-bc88-34313b235760 · outbound

This paper cites An unsupervised autoregressive model for speech representation learning,.

Towards Maximum Likelihood Training for Transducer-based Streaming Speech Recognition An unsupervised autoregressive model for speech representation learning,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:13:12.189145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:13:11.893079Z digest=sha256:e145ec4c44b94912fb672241b2d5de0f99e9ec5b62a5fd8dc5a98e8d7249ee96

Observation 7d281fd7-cbe5-4a1f-9e41-fb57c8f17546 · outbound

This paper cites Variational inference with normalizing flows,.

Towards Maximum Likelihood Training for Transducer-based Streaming Speech Recognition Variational inference with normalizing flows,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:13:12.174075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:13:11.897533Z digest=sha256:d39849411885b774e1174681c4d10c45703d6b2c9cfe5e9066e9d676270e3d1d

Observation 0d4597f1-bc0a-447f-9fe8-b7d039341609 · outbound

This paper cites https://github.com/k2-fsa/icefall.

Towards Maximum Likelihood Training for Transducer-based Streaming Speech Recognition https://github.com/k2-fsa/icefall

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:13:12.158985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:13:11.901971Z digest=sha256:c3cca8452f9e893ed7b8a992341aafc38f95e64ff0ccb4d672b9272b66e6afca

Observation 7eef7b32-dfc2-4500-b436-de80460cd04d · outbound

This paper cites Librispeech: an asr corpus based on public domain audio books,.

Towards Maximum Likelihood Training for Transducer-based Streaming Speech Recognition Librispeech: an asr corpus based on public domain audio books,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:13:12.143688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:13:11.906742Z digest=sha256:f26c23a52d9059399250f9382fdc3cb214345c6009e4b690d17afff7c97264ee

Observation 343bbec1-16af-49d8-8430-fb1ee3a486aa · outbound

This paper cites Ted-lium 3: Twice as much data and corpus repartition for experiments on speaker adaptation,.

Towards Maximum Likelihood Training for Transducer-based Streaming Speech Recognition Ted-lium 3: Twice as much data and corpus repartition for experiments on speaker adaptation,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:13:12.127432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:13:11.911102Z digest=sha256:8c3702c0854b9408c41aba7e353a5a00f04cfa79bf14850363f152b109c894c9

Observation 74753a00-dfa4-4b74-83f1-23b4aadd3664 · outbound

This paper cites A new algorithm for data compression,.

Towards Maximum Likelihood Training for Transducer-based Streaming Speech Recognition A new algorithm for data compression,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:13:12.112275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:13:11.915698Z digest=sha256:2a94656a27340df50cc7785a07dc475d18cd926338c3da1003f8ffa40032b0e5

Observation 1f671f66-86ad-4e7b-84ce-7b688cad4f50 · outbound

This paper cites Neural Machine Translation of Rare Words with Subword Units.

Towards Maximum Likelihood Training for Transducer-based Streaming Speech Recognition Neural Machine Translation of Rare Words with Subword Units

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T12:13:11.920124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:13:11.920124Z digest=sha256:5bff8534be49f874821c6cfcd7d871debd397a5eeda5215a6e01141f8a2c1937

Observation 28b369b5-17ba-4972-859f-618ccaef638b · outbound

This paper cites Zipformer: A faster and better encoder for automatic speech recognition.

Towards Maximum Likelihood Training for Transducer-based Streaming Speech Recognition Zipformer: A faster and better encoder for automatic speech recognition

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T12:13:11.925137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:13:11.925137Z digest=sha256:049d7df51ba866fc95a02042fe9e433389f739e4ab2fed54eb282ebd76e17e55

Observation 250cd683-1af0-4a57-91ea-0b7d501eabd6 · outbound

This paper cites Rnn- transducer with stateless prediction network,.

Towards Maximum Likelihood Training for Transducer-based Streaming Speech Recognition Rnn- transducer with stateless prediction network,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:13:12.097217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:13:11.930273Z digest=sha256:235fbaeb2ed73a5e7ef3a3724a317842d91b631ebebb0c8f8ec06b3ce189abf4

Observation 011b3b63-fd94-4e61-8ad7-803d89938d9e · outbound

This paper cites Made: Masked autoencoder for distribution estimation,.

Towards Maximum Likelihood Training for Transducer-based Streaming Speech Recognition Made: Masked autoencoder for distribution estimation,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:13:12.081969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:13:11.934933Z digest=sha256:f4e85299c76c294a7376e7c3591027a17e44a7cdab22435f736af7bcf9a9aa1a

Observation 42beed4c-1604-49a0-8c7b-e468d514b2cb · outbound

This paper cites Specaugment: A simple data augmentation method for automatic speech recognition,.

Towards Maximum Likelihood Training for Transducer-based Streaming Speech Recognition Specaugment: A simple data augmentation method for automatic speech recognition,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:13:12.066442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:13:11.939403Z digest=sha256:6abd66c6bb4a0d44731053f62afc2d37e1f8b9a0728b3a51cc2e6d5cea82ea08

Pith citing papers

No inbound Pith citation observations are available.