Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T20:15:28.530654Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 0 inbound Pith citation observations for arXiv:2501.09104.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T20:15:28.530654Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
38 of 38 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 19382c3b-43d4-4f07-a5f9-f554ba6d7e2f · outbound
A Non-autoregressive Model for Joint STT and TTS Almost unsupervised text to speech and automatic speech recognition,
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1cbe7a3b-e633-4f52-b960-bb91073d34ee · outbound
A Non-autoregressive Model for Joint STT and TTS Speecht5: Unified-modal encoder-decoder pre-training for spoken language processing,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ce8b6928-7725-4f9a-ad72-f4fcab158215 · outbound
A Non-autoregressive Model for Joint STT and TTS LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e745e0e5-a8ee-4a6a-bcf3-0b8162d14bab · outbound
A Non-autoregressive Model for Joint STT and TTS SeamlessM4T: Massively Multilingual & Multimodal Machine Translation
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 986c61e3-f5ce-4887-9e99-81673de043b5 · outbound
A Non-autoregressive Model for Joint STT and TTS SpeechVerse: A Large-scale Generalizable Audio Language Model
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5070c066-6467-47c6-9d22-90b4fb84ef2b · outbound
A Non-autoregressive Model for Joint STT and TTS Viola: Conditional language models for speech recognition, synthesis, and translation,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 210853eb-d742-4853-9626-2bc877e506f7 · outbound
A Non-autoregressive Model for Joint STT and TTS FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e92e1ccd-e614-476f-b7f5-bd0aca3b7af6 · outbound
A Non-autoregressive Model for Joint STT and TTS OWSM-CTC: An Open Encoder-Only Speech Foundation Model for Speech Recognition, Translation, and Language Identification
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e185ac2e-fefd-4400-94fb-57479b7a0087 · outbound
A Non-autoregressive Model for Joint STT and TTS Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d7e56b1-3251-4f64-8795-59b5b151f0e1 · outbound
A Non-autoregressive Model for Joint STT and TTS Fastspeech: Fast, robust and controllable text to speech,
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9135e5b5-2be5-4f6a-9626-61a1d0b3309e · outbound
A Non-autoregressive Model for Joint STT and TTS Joist: A joint speech and text streaming model for asr,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f8c5f476-a811-4f5a-a647-fdfd02816f7a · outbound
A Non-autoregressive Model for Joint STT and TTS Integrating text inputs for training and adapting rnn transducer asr models,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3c595c02-b884-4bc0-8e2e-5249f72347bb · outbound
A Non-autoregressive Model for Joint STT and TTS Semi-autoregressive streaming asr with label context,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c82be303-93f6-488a-b8f6-e55b71a57ae6 · outbound
A Non-autoregressive Model for Joint STT and TTS Align-Refine: Non-Autoregressive Speech Recognition via Iterative Realignment
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6d57ffeb-8354-48c8-b6a0-a6e84ac19d02 · outbound
A Non-autoregressive Model for Joint STT and TTS Mask CTC: Non-Autoregressive End-to-End ASR with CTC and Mask Predict
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a5baa8fd-78f8-46dc-bc2c-7badce02f133 · outbound
A Non-autoregressive Model for Joint STT and TTS BERT Meets CTC: New Formulation of End-to-End Speech Recognition with Pre-trained Masked Language Model
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2e82dfc9-12b7-4d54-9ae0-f774b1688393 · outbound
A Non-autoregressive Model for Joint STT and TTS Bectra: Transducer-based end-to-end asr with bert-enhanced encoder,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 52ffb3d8-360b-4cb8-bd8d-dddbaa6cada6 · outbound
A Non-autoregressive Model for Joint STT and TTS Mask-conformer: Augmenting conformer with mask-predict decoder,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5e40e57e-16cc-46c1-93a4-c962b429327d · outbound
A Non-autoregressive Model for Joint STT and TTS wav2vec 2.0: A framework for self-supervised learning of speech representations,
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation afc07b86-e777-44ab-b611-e14c315d9f89 · outbound
A Non-autoregressive Model for Joint STT and TTS Layer normalization,
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7279e4c5-b81f-4fdb-b014-702de35ec5c3 · outbound
A Non-autoregressive Model for Joint STT and TTS UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d529b1a-5af5-4d0c-a21c-e91333cc6ae9 · outbound
A Non-autoregressive Model for Joint STT and TTS Robust speech recognition via large-scale weak supervision,
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82b3c130-1de6-46c1-aa51-3fb9b5a1f987 · outbound
A Non-autoregressive Model for Joint STT and TTS ESPnet-SPK: full pipeline speaker embedding toolkit with reproducible recipes, self-supervised front-ends, and off-the-shelf models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96cce75b-7029-453f-9f93-4549ae9cf4fe · outbound
A Non-autoregressive Model for Joint STT and TTS The lj speech dataset,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 348305b0-a045-4aff-8f80-e15c448bc940 · outbound
A Non-autoregressive Model for Joint STT and TTS LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 442bc79f-1de7-4d20-913a-e655ee78199b · outbound
A Non-autoregressive Model for Joint STT and TTS Librispeech: an asr corpus based on public domain audio books,
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a0bd6be-8bba-440d-92d7-eeb78616a12b · outbound
A Non-autoregressive Model for Joint STT and TTS LibriTTS-R: A Restored Multi-Speaker Text-to-Speech Corpus
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dba533ba-213b-44b6-b407-465ed4750d52 · outbound
A Non-autoregressive Model for Joint STT and TTS Conformer: Convolution-augmented Transformer for Speech Recognition
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18d15006-6a65-46be-87ac-9ce3b30c07ba · outbound
A Non-autoregressive Model for Joint STT and TTS ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60d7cb1f-ff83-4449-a370-e8cf45f47c44 · outbound
A Non-autoregressive Model for Joint STT and TTS Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7701a5c4-ff0c-4c81-bba2-b716a10662c7 · outbound
A Non-autoregressive Model for Joint STT and TTS SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 580a110f-fe7f-4fd9-9ad7-f31e81b52e7c · outbound
A Non-autoregressive Model for Joint STT and TTS Audio augmentation for speech recognition
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a30e41ff-9fd7-4bbd-a147-5d29be044c01 · outbound
A Non-autoregressive Model for Joint STT and TTS Super-convergence: Very fast training of neural networks using large learning rates,
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0fdf2ab-b8e7-4cb6-b0e6-66a6fa579f6f · outbound
A Non-autoregressive Model for Joint STT and TTS Rethinking the inception architecture for computer vision,
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4c980a6-9f46-484e-95ff-654787756f49 · outbound
A Non-autoregressive Model for Joint STT and TTS Regularization of neural networks using dropconnect,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 004df12f-cf94-43a1-b93f-989747bbc029 · outbound
A Non-autoregressive Model for Joint STT and TTS Sequence noise injected training for end-to-end speech recognition,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9a06238f-8bab-4481-971b-91cb02c8bc78 · outbound
A Non-autoregressive Model for Joint STT and TTS Relaxing the Conditional Independence Assumption of CTC-based ASR by Conditioning on Intermediate Predictions
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25f1d255-6d72-4087-87ff-f4a4ff22e5fa · outbound
A Non-autoregressive Model for Joint STT and TTS FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.