Pith. sign in

Paper Citation Record · LEDGER

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs

As of 18 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 0 inbound Pith citation observations for arXiv:2607.06831.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.06831 v1

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-10T20:25:33.127942Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

50 of 50 outbound references displayed

  • verified exact11
  • verified fuzzy39
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d966a59c-cd8c-476f-a0b9-3b35211818d5 · outbound

This paper cites The application of hidden Markov models in speech recognition,.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs The application of hidden Markov models in speech recognition,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:27:36.749417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:cd48c48b695a7b1792e0010e6d1799c6ddc887f2ed21197bae014631e7ebd8ab

Observation d45345d8-d7fe-4b76-9a97-2b18a0f7cf52 · outbound

This paper cites A tutorial on hidden Markov models and selected applications in speech recognition,.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs A tutorial on hidden Markov models and selected applications in speech recognition,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:27:36.732459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:af1e9e713844e9bb1bb5c886d70ea139139c223d7aff7876ed61d19a311d4e37

Observation aeb462a3-2d33-4f03-b2df-157037d21e14 · outbound

This paper cites Montreal Forced Aligner: Trainable text-speech alignment using Kaldi,.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs Montreal Forced Aligner: Trainable text-speech alignment using Kaldi,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:27:36.745973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:2580461720452ae2276d7652cb75bf4f6ed7acadaea8c95246f6056298702261

Observation 6f3e1311-fec4-47c0-93b7-664107fb0ff3 · outbound

This paper cites Less peaky and more accurate CTC forced alignment by label priors,.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs Less peaky and more accurate CTC forced alignment by label priors,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:27:36.747709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:2dd3bc3f6e6638ace0951f1503d146dac26676bb2987f85dd06a873656379b4d

Observation 50d9969e-337f-4be5-8045-927e882fac6d · outbound

This paper cites Tradition or inno- vation: A comparison of modern asr methods for forced alignment,.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs Tradition or inno- vation: A comparison of modern asr methods for forced alignment,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:27:36.726454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:26a347b45b53fa4d77ffb5abf387605dab85ba4c428b65990da8a02ed6173b66

Observation 810712ad-b30a-4540-a290-e964861f9a83 · outbound

This paper cites End-to-end speech recognition: A survey,.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs End-to-end speech recognition: A survey,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:27:36.742389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:53a974c2a1ab43f291549401a1b9a5e275058535fb94f473f2d8e467a91a3060

Observation 857bdb36-5b2c-4174-a6c7-da2b75041ccf · outbound

This paper cites Connection- ist temporal classification: labelling unsegmented sequence data with recurrent neural networks,.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs Connection- ist temporal classification: labelling unsegmented sequence data with recurrent neural networks,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:27:36.734322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:a5a0184a5d0ac16ea14e69859e909d09a976d7886fff3581dd18459b30226637

Observation c54e7385-2112-4562-ae87-af83100b8fae · outbound

This paper cites Sequence Transduction with Recurrent Neural Networks.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs Sequence Transduction with Recurrent Neural Networks

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-10T20:27:36.593477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:106b993cb6c757de50859284449d3e7a00bcae6dfe49c8a0ac533639abc469e0

Observation f47e5bcf-50fd-4cae-9ac1-d048fd931d7a · outbound

This paper cites Attention-based models for speech recognition,.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs Attention-based models for speech recognition,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:27:36.760103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:acc9e23447ba3ed9f185cf0301754a18bb3e286e0e13d4dc0420a59fc9f7fad0

Observation 3a4df90a-b69d-4449-939a-776b5f7ec512 · outbound

This paper cites Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:27:36.758483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:f0dcbc768b80772bce4241aecf47f00422bf1b661f0cab0ca050594c335e85dd

Observation 68e4eecd-db03-4145-b76c-bb713c403e0d · outbound

This paper cites Improved training of end-to- end attention models for speech recognition,.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs Improved training of end-to- end attention models for speech recognition,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:27:36.736486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:e3f344361ebd16006830ce5c3511018f43b6fb06ddfc89919e7636a59fd16500

Observation 9c46d745-96f8-469c-91f5-b5f7f68023dc · outbound

This paper cites SpeechGPT: Empowering large language models with intrinsic cross-modal conversational abilities,.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs SpeechGPT: Empowering large language models with intrinsic cross-modal conversational abilities,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:27:36.728620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:a112b83bece58f0f547fb1c0c32b06a8e44f627bffa69f2802b3f815a7127bd7

Observation 7ce561c0-0674-4ab2-b590-4b95fa456615 · outbound

This paper cites Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-10T20:27:36.581542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:1e1a569fe422b02db60f33a21ba3e61ce5e630dac8a7c791dd3cd3fcee415631

Observation 399b1390-7481-4696-a009-023215cc1122 · outbound

This paper cites LLMs and Speech: Integration vs. Combination.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs LLMs and Speech: Integration vs. Combination

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-10T20:27:36.588062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:a19636ad4f0b31a8b313158ac1100beb91dd0d74387b3433a9f3fcf5852b6de8

Observation ae15985d-1df6-4eb3-bf05-092468164318 · outbound

This paper cites Robust speech recognition via large-scale weak supervision,.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs Robust speech recognition via large-scale weak supervision,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:27:36.781600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:1ff6297b0f064a1d128a59f8e28e6718d54bf40d4f2e2ef7a620d2ff5aa51cda

Observation 60c69c61-7a2a-4850-a787-41f8ad045c9d · outbound

This paper cites WhisperX: Time-accurate speech transcription of long-form audio,.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs WhisperX: Time-accurate speech transcription of long-form audio,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:27:36.754796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:c7a431f5f7478c0b0dfdcf3d019414a606b31cc1f7601d8743f04778dfb32250

Observation b199a002-115e-49f6-bdbb-7ae779d26b93 · outbound

This paper cites CrisperWhisper: Accurate timestamps on verbatim speech transcriptions,.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs CrisperWhisper: Accurate timestamps on verbatim speech transcriptions,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:27:36.792339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:e4e621a1f4574be8365ed47511f14b3a31197752d06787baccb484f4d474d82b

Observation f96fb2a6-db40-4145-a810-8a5fed11c603 · outbound

This paper cites Whisper has an internal word aligner,.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs Whisper has an internal word aligner,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:27:36.788911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:b6d534785f74e61aee86ad0c9a6cb06c4bf1b92f8beb7919a0b777ff1e0bac84

Observation fb35d85a-b9c4-4f0e-98b6-f83dc377ac08 · outbound

This paper cites Attention-Constrained Inference for Robust Decoder-Only Text-to-Speech.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs Attention-Constrained Inference for Robust Decoder-Only Text-to-Speech

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-07-10T20:27:36.605092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:6da4ed37ebac018bf5e88f7e59dca93ad51531c96a452fb9f6c80fb54ea64527

Observation db745102-9d9f-4963-8f66-cf92f78b966d · outbound

This paper cites DOA: Training-Free Decoder-Only Attention Policy for Long-Form Simultaneous Translation with SpeechLLMs.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs DOA: Training-Free Decoder-Only Attention Policy for Long-Form Simultaneous Translation with SpeechLLMs

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-10T20:27:36.584712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:c117d4d56385bbe7d248c40aced2398f1a450c8aa3b98051a618effb65f9dbe0

Observation 8fa22560-df82-4b40-b451-dfd3f57a013f · outbound

This paper cites Available: https://arxiv.org/abs/2601.18220.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs Available: https://arxiv.org/abs/2601.18220

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-10T20:27:36.599158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:7f45514227a1fe9293f2bc8bb9a6b1e62a69d698ce0c6e1d629a5cde5fa4d160

Observation 46faea88-3537-461b-89de-0bdc9534c20e · outbound

This paper cites Deep Inside Convolutional Networks: Visualising Image Classification Models and Saliency Maps.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs Deep Inside Convolutional Networks: Visualising Image Classification Models and Saliency Maps

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-07-10T20:27:36.610300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:eae476a2e1eacdf20a4bd9a01df4ce5117d815d974a64ce2040d8c95a3e2da7c

Observation b5413ae3-61ac-42fb-8423-5534e301fb5b · outbound

This paper cites Right Label Context in End-to-End Training of Time-Synchronous ASR Models.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs Right Label Context in End-to-End Training of Time-Synchronous ASR Models

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-07-10T20:27:36.590845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:097e05b4603e8e61f03ff9a37c7bff9f0f067e140b39b19c30ac34fd427d77b3

Observation 14c126d0-ee23-44f9-996c-8862d3fdc5e3 · outbound

This paper cites Saliency-driven word alignment interpretation for neural machine translation,.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs Saliency-driven word alignment interpretation for neural machine translation,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:27:36.756579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:c15309ed2ca063d3da706fdf1bd5b9da6ed25dfe0aa79886e9c60d8acf047706

Observation 80b417cf-69a4-4c7a-8385-efd91f5a14d2 · outbound

This paper cites Joint CTC/attention decoding for end-to-end speech recognition,.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs Joint CTC/attention decoding for end-to-end speech recognition,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:27:36.786985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:a07dd07d606b597059232cdf9a3de1fabf2dd5749acc454bfa1b1c9fd9c6119b

Observation 2b14ce76-a4ce-442e-a644-f717c54a511d · outbound

This paper cites Scaling speech tech- nology to 1,000+ languages,.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs Scaling speech tech- nology to 1,000+ languages,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:27:36.785158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:7f4d78ba008284a35d9577ec750f530bd99a08d5ffe2d7bbf46a443199a3684c

Observation 141c56c5-d111-482b-ad7d-12872cc90159 · outbound

This paper cites TorchAudio 2.1: Advancing speech recognition, self-supervised learning, and audio processing components for PyTorch,.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs TorchAudio 2.1: Advancing speech recognition, self-supervised learning, and audio processing components for PyTorch,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:27:36.724159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:622aae7d195ba91fd887d67a4d969ca81a62c986a27a100c499d8ca6df7598c1

Observation 7e474b36-d535-4835-af3f-56cce3229774 · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech representations,.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs wav2vec 2.0: A framework for self-supervised learning of speech representations,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:27:36.765630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:260aa9bfe03d932170d287b11aad900bc812742e9e193f2187130d412ef77e50

Observation 909c4e4a-76a3-4b23-afcd-9256897df7fd · outbound

This paper cites Automatic phoneme recognition on TIMIT dataset with Wav2Vec 2.0,.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs Automatic phoneme recognition on TIMIT dataset with Wav2Vec 2.0,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:27:36.769218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:364ea48f3dc8111902457d84087f1a1ef9bc63631d987c703141158e0a4471bf

Observation faa55225-00b2-46c4-bc19-88ece0e60d4c · outbound

This paper cites XLS-R: Self-supervised cross-lingual speech representation learning at scale,.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs XLS-R: Self-supervised cross-lingual speech representation learning at scale,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:27:36.776543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:2c237aac32e814ca809e1f451d38ee8a9f7f41df18ce32ac43c44230e509b999

Observation f8a77ee5-878e-41e4-9de1-0f4b186747bb · outbound

This paper cites Fast conformer with linearly scalable attention for efficient speech recognition,.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs Fast conformer with linearly scalable attention for efficient speech recognition,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:27:36.783272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:cfbd3bdf8a74d19d4eac04d299e7ad03fe5187d9a9027c27cd2fbb32104fc51d

Observation e61fd5aa-ba69-4030-87ff-2ca320b7a803 · outbound

This paper cites OWSM v4: Improving open whisper-style speech models via data scaling and cleaning,.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs OWSM v4: Improving open whisper-style speech models via data scaling and cleaning,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:27:36.790646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:21f47e2123ca4b5d21e223c3ca3f9262ffad81cafcb0c749c647706057020f1b

Observation 192f1e8d-8680-45ed-9418-f3228e4196ea · outbound

This paper cites OWSM-CTC: An open encoder-only speech foundation model for speech recognition, translation, and language identification,.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs OWSM-CTC: An open encoder-only speech foundation model for speech recognition, translation, and language identification,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:27:36.767507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:2fd67eed0ea9fb1b719c414a363d46deb3d29966d5a8b3ff93c8256d677240f1

Observation a9eaa9a3-1064-4281-b3d4-2894d73a38a3 · outbound

This paper cites Stateful conformer with cache-based inference for streaming automatic speech recognition,.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs Stateful conformer with cache-based inference for streaming automatic speech recognition,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:27:36.770999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:7f5cb5b7a45f4b3ac8ec8a78c8afa74b667f2afbb33e247d68cbbd8b0faa658e

Observation 6c777291-1348-41e2-85d0-09d8e82d52cf · outbound

This paper cites Efficient sequence transduction by jointly predicting tokens and durations,.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs Efficient sequence transduction by jointly predicting tokens and durations,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:27:36.763726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:0be6f2af58273b0cf3f8b0235f041aa710a6cfd60b8b9f532380a1e317c1f427

Observation 2b966bde-edd6-4365-83b5-d13f50453e29 · outbound

This paper cites Emformer: Efficient memory transformer based acoustic model for low latency streaming speech recognition,.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs Emformer: Efficient memory transformer based acoustic model for low latency streaming speech recognition,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:27:36.753093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:4f7830cc83834641fe593d187434fcbe3c9a548456601d3a7f655a87b66839ef

Observation bfa1780e-6102-4e67-b507-77b61a9746c3 · outbound

This paper cites TorchAudio: Building blocks for audio and speech processing,.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs TorchAudio: Building blocks for audio and speech processing,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:27:36.778265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:7c3f76f32504bf79afbc8a181af516dd50f8c5f052af4abf9b8f742be55ea3bf

Observation 8d6223ec-4015-456d-ad38-511db9c02f77 · outbound

This paper cites OWLS: Scaling laws for multilingual speech recognition and translation models,.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs OWLS: Scaling laws for multilingual speech recognition and translation models,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:27:36.762010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:a0c64ef2f6c88c3a9102ba5a379650b9ae94c624874775b03531f73f54eb27c0

Observation 3db9f10c-f72e-4a63-802f-e476e6864eec · outbound

This paper cites Voxtral.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs Voxtral

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-07-10T20:27:36.607654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:78ec8f04a36f479142917489b6ea3627801e227ca17680358c25e60ab23320a0

Observation 0f200d49-b50b-44bd-b3c5-d4edadfefd4d · outbound

This paper cites Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-07-10T20:27:36.602301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:83f888c160abaf5c0064166f133d8b2c40b2f2d4b8da2b7dc7660c02c4715ee3

Observation 2d604f11-08f7-4c62-86ce-c490b9e6441f · outbound

This paper cites Less is more: Accurate speech recognition & translation without web-scale data,.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs Less is more: Accurate speech recognition & translation without web-scale data,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:27:36.740428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:77a51290e380799d2b4d5c3f611ca8be6ad8090f811da0b54364a08531376d1b

Observation 5f231df0-a75d-4f6e-885a-dee428eb86c7 · outbound

This paper cites TIMIT acoustic-phonetic continuous speech corpus,.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs TIMIT acoustic-phonetic continuous speech corpus,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:27:36.779980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:ea7316365348988bd2955eb41f14232e917e1e61ce09444ac5aa6b6f744e5fc0

Observation 84564faf-fd3a-448e-abb3-4dfb050d909e · outbound

This paper cites The Buckeye corpus of conversational speech: Labeling conventions and a test of transcriber reliability,.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs The Buckeye corpus of conversational speech: Labeling conventions and a test of transcriber reliability,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:27:36.744190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:3784d1daf4f6af9eb7ffc4dce4111f33d23e4ef0d0a5b275158a48fa3ad325c3

Observation cc9ab90d-6d58-4ff2-8ea6-9c477c6d02dd · outbound

This paper cites Learning important features through propagating activation differences,.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs Learning important features through propagating activation differences,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:27:36.738421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:48569a40be66219caad40151615047659745fdaebfaf3687b20faaafced08836

Observation d7e1feb3-8b81-4716-a2eb-1b8bbd2bf388 · outbound

This paper cites Towards better understanding of gradient-based attribution methods for deep neural networks,.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs Towards better understanding of gradient-based attribution methods for deep neural networks,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:27:36.751310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:fcdec730bd228ffc5dfdd8b72c6bfa69e281d84a164193b24dbc8c05fa8848c9

Observation 2290b8b4-2ec2-4619-9cca-56ae8b26d3dc · outbound

This paper cites SmoothGrad: removing noise by adding noise.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs SmoothGrad: removing noise by adding noise

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-07-10T20:27:36.596077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:595eb031755a4b89776b2c854cad9b8967718e630861c54543a51f5ff1a43788

Observation 5ba70317-85aa-43c0-a86f-3db544ec09e8 · outbound

This paper cites Sanity checks for saliency maps,.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs Sanity checks for saliency maps,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:27:36.772994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:898cd92353dfd8a4a67abead553393390d096fc2637b074b2738085d35f1260a

Observation 41457776-1484-4b60-adca-67d6ee3eee18 · outbound

This paper cites Available: https://proceedings.neurips.cc/paper files/ paper/2018/file/294a8ed24b1ad22ec2e7efea049b8737-Paper.pdf.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs Available: https://proceedings.neurips.cc/paper files/ paper/2018/file/294a8ed24b1ad22ec2e7efea049b8737-Paper.pdf

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:27:36.794010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:c01845fce9c22f594ed970439785ccf73661405f46b483d2c3df89cba78d3ec3

Observation c0fb4029-8cc0-4385-9891-18f15ccd52b5 · outbound

This paper cites Axiomatic attribution for deep networks,.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs Axiomatic attribution for deep networks,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:27:36.730491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:aee75f4e215c35835af6109940b578002e981685803e37839bf54e6744a9ffee

Observation bb1834f8-f027-4183-b870-7fb7c94926f4 · outbound

This paper cites Improving performance of deep learning models with axiomatic attribution priors and expected gradients,.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs Improving performance of deep learning models with axiomatic attribution priors and expected gradients,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:27:36.774812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:e7ea3009ee626228a39aa331c351c47b20092f112fa1740f86c2bdc01fb9ca3d

Pith citing papers

No inbound Pith citation observations are available.