Pith. sign in

Paper Citation Record · LEDGER

A Non-autoregressive Model for Joint STT and TTS

As of 11 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 0 inbound Pith citation observations for arXiv:2501.09104.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.09104 v2

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T20:15:28.530654Z

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

38 of 38 outbound references displayed

  • verified exact3
  • verified fuzzy11
  • unresolved24
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 19382c3b-43d4-4f07-a5f9-f554ba6d7e2f · outbound

This paper cites Almost unsupervised text to speech and automatic speech recognition,.

A Non-autoregressive Model for Joint STT and TTS Almost unsupervised text to speech and automatic speech recognition,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:15:29.158504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:15:28.346059Z digest=sha256:5274b49e9e2dc24469b5f0667b3a89eb2a5d4862f990b537e2b34e57f8322245

Observation 1cbe7a3b-e633-4f52-b960-bb91073d34ee · outbound

This paper cites Speecht5: Unified-modal encoder-decoder pre-training for spoken language processing,.

A Non-autoregressive Model for Joint STT and TTS Speecht5: Unified-modal encoder-decoder pre-training for spoken language processing,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:15:29.143256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:15:28.351495Z digest=sha256:e06dc2426138cc5f6e55f5b9ef050a9b7179caecbb273d40d1ec21f6ae2f9530

Observation ce8b6928-7725-4f9a-ad72-f4fcab158215 · outbound

This paper cites LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT.

A Non-autoregressive Model for Joint STT and TTS LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:28.356241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:28.356241Z digest=sha256:d4d1341567c24044bc2c8563a6e6ddba964553f49f249db1dc7fcdda65e93b07

Observation e745e0e5-a8ee-4a6a-bcf3-0b8162d14bab · outbound

This paper cites SeamlessM4T: Massively Multilingual & Multimodal Machine Translation.

A Non-autoregressive Model for Joint STT and TTS SeamlessM4T: Massively Multilingual & Multimodal Machine Translation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:28.361491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:28.361491Z digest=sha256:1dc138ddd2fd130146233b50ab549109f5dcfb6c49ac69dd14378633c410ec4c

Observation 986c61e3-f5ce-4887-9e99-81673de043b5 · outbound

This paper cites SpeechVerse: A Large-scale Generalizable Audio Language Model.

A Non-autoregressive Model for Joint STT and TTS SpeechVerse: A Large-scale Generalizable Audio Language Model

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:28.366374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:28.366374Z digest=sha256:df57e543294dac079b44096ab09ec6c1e25eb04e3502cfb7e46df8cfb27a9067

Observation 5070c066-6467-47c6-9d22-90b4fb84ef2b · outbound

This paper cites Viola: Conditional language models for speech recognition, synthesis, and translation,.

A Non-autoregressive Model for Joint STT and TTS Viola: Conditional language models for speech recognition, synthesis, and translation,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:15:29.127478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:15:28.371562Z digest=sha256:3d1b7ac734faac9ae387fb309ca55f455b2718de22df84d0a3e2836dcc02ca64

Observation 210853eb-d742-4853-9626-2bc877e506f7 · outbound

This paper cites FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs.

A Non-autoregressive Model for Joint STT and TTS FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:28.376868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:28.376868Z digest=sha256:9a7d6bf0262caf1d654e19434bf332c4d75427b2a2d4b0e85528693fb0b86339

Observation e92e1ccd-e614-476f-b7f5-bd0aca3b7af6 · outbound

This paper cites OWSM-CTC: An Open Encoder-Only Speech Foundation Model for Speech Recognition, Translation, and Language Identification.

A Non-autoregressive Model for Joint STT and TTS OWSM-CTC: An Open Encoder-Only Speech Foundation Model for Speech Recognition, Translation, and Language Identification

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:28.381704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:28.381704Z digest=sha256:fbbd0bd4c555f518a530ce1cf3343cf9f27cfc17747a0869abe5a3d5622741e0

Observation e185ac2e-fefd-4400-94fb-57479b7a0087 · outbound

This paper cites Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,.

A Non-autoregressive Model for Joint STT and TTS Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:28.387144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:28.387144Z digest=sha256:4aa4b8e667de3f95daa83139d051cc8b6dbb563845d307482c93be96e8e0b36f

Observation 9d7e56b1-3251-4f64-8795-59b5b151f0e1 · outbound

This paper cites Fastspeech: Fast, robust and controllable text to speech,.

A Non-autoregressive Model for Joint STT and TTS Fastspeech: Fast, robust and controllable text to speech,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:28.391635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:28.391635Z digest=sha256:269e80f86aafb2c26453a5d41096763611fbf99c0673625fb086555f0b67abc8

Observation 9135e5b5-2be5-4f6a-9626-61a1d0b3309e · outbound

This paper cites Joist: A joint speech and text streaming model for asr,.

A Non-autoregressive Model for Joint STT and TTS Joist: A joint speech and text streaming model for asr,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:15:29.090188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:15:28.396298Z digest=sha256:36e90939bd94d909a680b43d9ab327cb2965f74d006876bef862d4757b0f60b1

Observation f8c5f476-a811-4f5a-a647-fdfd02816f7a · outbound

This paper cites Integrating text inputs for training and adapting rnn transducer asr models,.

A Non-autoregressive Model for Joint STT and TTS Integrating text inputs for training and adapting rnn transducer asr models,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:15:29.074509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:15:28.401786Z digest=sha256:eca160c16f150d5273b9372edaa7248a54c79e478ed5254dad95e2b2364acb42

Observation 3c595c02-b884-4bc0-8e2e-5249f72347bb · outbound

This paper cites Semi-autoregressive streaming asr with label context,.

A Non-autoregressive Model for Joint STT and TTS Semi-autoregressive streaming asr with label context,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:15:29.057689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:15:28.406510Z digest=sha256:f6584259108454cc4d6862e6d29e6074e625c67d7b653e849a44c9e92bf8265e

Observation c82be303-93f6-488a-b8f6-e55b71a57ae6 · outbound

This paper cites Align-Refine: Non-Autoregressive Speech Recognition via Iterative Realignment.

A Non-autoregressive Model for Joint STT and TTS Align-Refine: Non-Autoregressive Speech Recognition via Iterative Realignment

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-08-10T20:15:28.760841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:15:28.411135Z digest=sha256:2d23e434c5d12358c45be0f413edfbbb8f42fa72b72a441ba71a4972cc67b7b5

Observation 6d57ffeb-8354-48c8-b6a0-a6e84ac19d02 · outbound

This paper cites Mask CTC: Non-Autoregressive End-to-End ASR with CTC and Mask Predict.

A Non-autoregressive Model for Joint STT and TTS Mask CTC: Non-Autoregressive End-to-End ASR with CTC and Mask Predict

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-10T20:15:28.738910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:15:28.416162Z digest=sha256:caea363b2d59c1ef704fc74cf46c94e1ed59c9ac0b6ec49af314a0452ff1fe6d

Observation a5baa8fd-78f8-46dc-bc2c-7badce02f133 · outbound

This paper cites BERT Meets CTC: New Formulation of End-to-End Speech Recognition with Pre-trained Masked Language Model.

A Non-autoregressive Model for Joint STT and TTS BERT Meets CTC: New Formulation of End-to-End Speech Recognition with Pre-trained Masked Language Model

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-10T20:15:28.715509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:15:28.421118Z digest=sha256:18f645cf802629081c9951dc801ac70e5615ea7b8590f2abc92d90e49aef7265

Observation 2e82dfc9-12b7-4d54-9ae0-f774b1688393 · outbound

This paper cites Bectra: Transducer-based end-to-end asr with bert-enhanced encoder,.

A Non-autoregressive Model for Joint STT and TTS Bectra: Transducer-based end-to-end asr with bert-enhanced encoder,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:15:29.040118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:15:28.426110Z digest=sha256:c061be3e2c5846b812ceb0a3ada23dccac52a5d274476b7dd223c4b22da8ee87

Observation 52ffb3d8-360b-4cb8-bd8d-dddbaa6cada6 · outbound

This paper cites Mask-conformer: Augmenting conformer with mask-predict decoder,.

A Non-autoregressive Model for Joint STT and TTS Mask-conformer: Augmenting conformer with mask-predict decoder,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:15:29.022217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:15:28.430810Z digest=sha256:57f4378f7db9bef0fc4981476bc0405d32b56b28b4f0be03ec4ee54681b5f119

Observation 5e40e57e-16cc-46c1-93a4-c962b429327d · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech representations,.

A Non-autoregressive Model for Joint STT and TTS wav2vec 2.0: A framework for self-supervised learning of speech representations,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:28.435399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:28.435399Z digest=sha256:e7e535316152fb8ca1822af903a2ac919923b06316e32941bd61792fd39937ba

Observation afc07b86-e777-44ab-b611-e14c315d9f89 · outbound

This paper cites Layer normalization,.

A Non-autoregressive Model for Joint STT and TTS Layer normalization,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:28.440501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:28.440501Z digest=sha256:75f6b16e86bb6e05981fc6c10bd15c076fbf90cf743924791d1f0a575a775ab7

Observation 7279e4c5-b81f-4fdb-b014-702de35ec5c3 · outbound

This paper cites UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022.

A Non-autoregressive Model for Joint STT and TTS UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:28.445462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:28.445462Z digest=sha256:269c1d3dc735e596e3b1fc0d8bafe38fae3e4d78a0ab0349639a744dab5c74a0

Observation 3d529b1a-5af5-4d0c-a21c-e91333cc6ae9 · outbound

This paper cites Robust speech recognition via large-scale weak supervision,.

A Non-autoregressive Model for Joint STT and TTS Robust speech recognition via large-scale weak supervision,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:28.450541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:28.450541Z digest=sha256:f574f169961d3e4c064a003720b917e66049739c8504fd42516759e2d78f225b

Observation 82b3c130-1de6-46c1-aa51-3fb9b5a1f987 · outbound

This paper cites ESPnet-SPK: full pipeline speaker embedding toolkit with reproducible recipes, self-supervised front-ends, and off-the-shelf models.

A Non-autoregressive Model for Joint STT and TTS ESPnet-SPK: full pipeline speaker embedding toolkit with reproducible recipes, self-supervised front-ends, and off-the-shelf models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:28.455695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:28.455695Z digest=sha256:d8dc62dda2bd4c486f5a5d37ea8c385165ebb6c59323a32650eff13ce88f4a81

Observation 96cce75b-7029-453f-9f93-4549ae9cf4fe · outbound

This paper cites The lj speech dataset,.

A Non-autoregressive Model for Joint STT and TTS The lj speech dataset,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:15:28.967736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:15:28.460414Z digest=sha256:350869582e413f75fe91f692c07a235830966dacc61ea254270de2ed06f89b8f

Observation 348305b0-a045-4aff-8f80-e15c448bc940 · outbound

This paper cites LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech.

A Non-autoregressive Model for Joint STT and TTS LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:28.464814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:28.464814Z digest=sha256:1255a5d8f41967e386a7b7e0cda5fa12f6b35f588f2c44628ce35c964cb669fe

Observation 442bc79f-1de7-4d20-913a-e655ee78199b · outbound

This paper cites Librispeech: an asr corpus based on public domain audio books,.

A Non-autoregressive Model for Joint STT and TTS Librispeech: an asr corpus based on public domain audio books,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:28.469522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:28.469522Z digest=sha256:227399530034cf5b30e3083f1c054ed671877d36a651e3692f837b2bcb2dcc02

Observation 7a0bd6be-8bba-440d-92d7-eeb78616a12b · outbound

This paper cites LibriTTS-R: A Restored Multi-Speaker Text-to-Speech Corpus.

A Non-autoregressive Model for Joint STT and TTS LibriTTS-R: A Restored Multi-Speaker Text-to-Speech Corpus

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:28.474447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:28.474447Z digest=sha256:d382c4060502ba53fb328569a5b530961b114c90d8c5c5c1739bce100b884225

Observation dba533ba-213b-44b6-b407-465ed4750d52 · outbound

This paper cites Conformer: Convolution-augmented Transformer for Speech Recognition.

A Non-autoregressive Model for Joint STT and TTS Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:28.480440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:28.480440Z digest=sha256:3bbfcf822670639248d72e09781363ce082c7e1585cb3da7f058926e316a42c8

Observation 18d15006-6a65-46be-87ac-9ce3b30c07ba · outbound

This paper cites ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification.

A Non-autoregressive Model for Joint STT and TTS ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:28.485951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:28.485951Z digest=sha256:c2354302172f3af440b658d75f3652936f630ec161fd23a29c29ed09b9e71ddf

Observation 60d7cb1f-ff83-4449-a370-e8cf45f47c44 · outbound

This paper cites Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,.

A Non-autoregressive Model for Joint STT and TTS Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:28.491768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:28.491768Z digest=sha256:b5118bfb7dbad53a3d1592bbf029229cbee8c4e5c7c4e7584063bc539d0b82ab

Observation 7701a5c4-ff0c-4c81-bba2-b716a10662c7 · outbound

This paper cites SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition.

A Non-autoregressive Model for Joint STT and TTS SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:28.496296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:28.496296Z digest=sha256:72e8984f44eb37a4a819f57464c8af676aec59072f0ab3f1a4f0ca58b4818da8

Observation 580a110f-fe7f-4fd9-9ad7-f31e81b52e7c · outbound

This paper cites Audio augmentation for speech recognition.

A Non-autoregressive Model for Joint STT and TTS Audio augmentation for speech recognition

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:28.500917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:28.500917Z digest=sha256:6a25a140092c40890af7fc4600e68b5b8b8680cf5572e55d0fd0853e299a030b

Observation a30e41ff-9fd7-4bbd-a147-5d29be044c01 · outbound

This paper cites Super-convergence: Very fast training of neural networks using large learning rates,.

A Non-autoregressive Model for Joint STT and TTS Super-convergence: Very fast training of neural networks using large learning rates,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:28.506312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:28.506312Z digest=sha256:1d6e19a880bea13cd226c83c52fd5b9afa1c4d11c403a2db5f4b76346664329a

Observation d0fdf2ab-b8e7-4cb6-b0e6-66a6fa579f6f · outbound

This paper cites Rethinking the inception architecture for computer vision,.

A Non-autoregressive Model for Joint STT and TTS Rethinking the inception architecture for computer vision,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:28.510887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:28.510887Z digest=sha256:5d45a344f3785e5eee0caf7e7a4a11a3816c28c15c7be566ff5cd0fd5223ee69

Observation f4c980a6-9f46-484e-95ff-654787756f49 · outbound

This paper cites Regularization of neural networks using dropconnect,.

A Non-autoregressive Model for Joint STT and TTS Regularization of neural networks using dropconnect,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:15:28.891532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:15:28.516216Z digest=sha256:1460664479011098224254749abdce03cc904fccd5d4478d65cb75feef6d05c2

Observation 004df12f-cf94-43a1-b93f-989747bbc029 · outbound

This paper cites Sequence noise injected training for end-to-end speech recognition,.

A Non-autoregressive Model for Joint STT and TTS Sequence noise injected training for end-to-end speech recognition,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:15:28.873661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:15:28.520889Z digest=sha256:8240e8f2996c99c999008e0c96ab001ec4db78a896c3975c939a2bc602e1c7ab

Observation 9a06238f-8bab-4481-971b-91cb02c8bc78 · outbound

This paper cites Relaxing the Conditional Independence Assumption of CTC-based ASR by Conditioning on Intermediate Predictions.

A Non-autoregressive Model for Joint STT and TTS Relaxing the Conditional Independence Assumption of CTC-based ASR by Conditioning on Intermediate Predictions

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:28.525732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:28.525732Z digest=sha256:3901d5ae20f98912636c10103ec388213af9e168587fd904f277a67e27dde6a5

Observation 25f1d255-6d72-4087-87ff-f4a4ff22e5fa · outbound

This paper cites FastSpeech 2: Fast and High-Quality End-to-End Text to Speech.

A Non-autoregressive Model for Joint STT and TTS FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:28.530654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:28.530654Z digest=sha256:c5129dbff5e3aced6d54565f6e19a08297217d3b225e958d7d830aef719ed794

Pith citing papers

No inbound Pith citation observations are available.