Pith. sign in

Paper Citation Record · LEDGER

A Non-autoregressive Model for Joint STT and TTS

As of 11 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 0 inbound Pith citation observations for arXiv:2501.09104.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.09104 v2

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T20:15:28.530654Z

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

38 of 38 outbound references displayed

  • verified exact3
  • verified fuzzy11
  • unresolved24
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 19382c3b-43d4-4f07-a5f9-f554ba6d7e2f · outbound

This paper cites Almost unsupervised text to speech and automatic speech recognition,.

A Non-autoregressive Model for Joint STT and TTS Almost unsupervised text to speech and automatic speech recognition,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:15:29.158504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:15:28.346059Z digest=sha256:d0b554816afc0ce000f51644fda2ab83bb4ab7e6a9f665f166a0349b34a6e7de

Observation 1cbe7a3b-e633-4f52-b960-bb91073d34ee · outbound

This paper cites Speecht5: Unified-modal encoder-decoder pre-training for spoken language processing,.

A Non-autoregressive Model for Joint STT and TTS Speecht5: Unified-modal encoder-decoder pre-training for spoken language processing,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:15:29.143256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:15:28.351495Z digest=sha256:d043074b501c24755fe54fb9f90215c176b41797d4d0e8914ea7f1dae3e1dd3d

Observation ce8b6928-7725-4f9a-ad72-f4fcab158215 · outbound

This paper cites LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT.

A Non-autoregressive Model for Joint STT and TTS LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:28.356241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:28.356241Z digest=sha256:d4d1341567c24044bc2c8563a6e6ddba964553f49f249db1dc7fcdda65e93b07

Observation e745e0e5-a8ee-4a6a-bcf3-0b8162d14bab · outbound

This paper cites SeamlessM4T: Massively Multilingual & Multimodal Machine Translation.

A Non-autoregressive Model for Joint STT and TTS SeamlessM4T: Massively Multilingual & Multimodal Machine Translation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:28.361491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:28.361491Z digest=sha256:1dc138ddd2fd130146233b50ab549109f5dcfb6c49ac69dd14378633c410ec4c

Observation 986c61e3-f5ce-4887-9e99-81673de043b5 · outbound

This paper cites SpeechVerse: A Large-scale Generalizable Audio Language Model.

A Non-autoregressive Model for Joint STT and TTS SpeechVerse: A Large-scale Generalizable Audio Language Model

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:28.366374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:28.366374Z digest=sha256:df57e543294dac079b44096ab09ec6c1e25eb04e3502cfb7e46df8cfb27a9067

Observation 5070c066-6467-47c6-9d22-90b4fb84ef2b · outbound

This paper cites Viola: Conditional language models for speech recognition, synthesis, and translation,.

A Non-autoregressive Model for Joint STT and TTS Viola: Conditional language models for speech recognition, synthesis, and translation,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:15:29.127478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:15:28.371562Z digest=sha256:b433c1a991e475c452f9815eb9b940debbc69710468a4da67c76237643091d3d

Observation 210853eb-d742-4853-9626-2bc877e506f7 · outbound

This paper cites FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs.

A Non-autoregressive Model for Joint STT and TTS FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:28.376868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:28.376868Z digest=sha256:9a7d6bf0262caf1d654e19434bf332c4d75427b2a2d4b0e85528693fb0b86339

Observation e92e1ccd-e614-476f-b7f5-bd0aca3b7af6 · outbound

This paper cites OWSM-CTC: An Open Encoder-Only Speech Foundation Model for Speech Recognition, Translation, and Language Identification.

A Non-autoregressive Model for Joint STT and TTS OWSM-CTC: An Open Encoder-Only Speech Foundation Model for Speech Recognition, Translation, and Language Identification

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:28.381704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:28.381704Z digest=sha256:fbbd0bd4c555f518a530ce1cf3343cf9f27cfc17747a0869abe5a3d5622741e0

Observation e185ac2e-fefd-4400-94fb-57479b7a0087 · outbound

This paper cites Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,.

A Non-autoregressive Model for Joint STT and TTS Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:28.387144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:28.387144Z digest=sha256:4aa4b8e667de3f95daa83139d051cc8b6dbb563845d307482c93be96e8e0b36f

Observation 9d7e56b1-3251-4f64-8795-59b5b151f0e1 · outbound

This paper cites Fastspeech: Fast, robust and controllable text to speech,.

A Non-autoregressive Model for Joint STT and TTS Fastspeech: Fast, robust and controllable text to speech,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:28.391635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:28.391635Z digest=sha256:269e80f86aafb2c26453a5d41096763611fbf99c0673625fb086555f0b67abc8

Observation 9135e5b5-2be5-4f6a-9626-61a1d0b3309e · outbound

This paper cites Joist: A joint speech and text streaming model for asr,.

A Non-autoregressive Model for Joint STT and TTS Joist: A joint speech and text streaming model for asr,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:15:29.090188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:15:28.396298Z digest=sha256:e81cb67c08cb2418d9c4fe9cc3cd80083141430375d468e608b9a7b01a6e49a5

Observation f8c5f476-a811-4f5a-a647-fdfd02816f7a · outbound

This paper cites Integrating text inputs for training and adapting rnn transducer asr models,.

A Non-autoregressive Model for Joint STT and TTS Integrating text inputs for training and adapting rnn transducer asr models,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:15:29.074509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:15:28.401786Z digest=sha256:db0179b6247e9801fc34ea525dc3b4e29b5f86919eb4f195d3022fe0908ee654

Observation 3c595c02-b884-4bc0-8e2e-5249f72347bb · outbound

This paper cites Semi-autoregressive streaming asr with label context,.

A Non-autoregressive Model for Joint STT and TTS Semi-autoregressive streaming asr with label context,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:15:29.057689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:15:28.406510Z digest=sha256:87104b8c4e667684d3e95be4fe311e4d0e193c18af56a7c058c2a264fd68662d

Observation c82be303-93f6-488a-b8f6-e55b71a57ae6 · outbound

This paper cites Align-Refine: Non-Autoregressive Speech Recognition via Iterative Realignment.

A Non-autoregressive Model for Joint STT and TTS Align-Refine: Non-Autoregressive Speech Recognition via Iterative Realignment

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-08-10T20:15:28.760841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:15:28.411135Z digest=sha256:89681810290020e41cc16d77eb48d7fbe1b82c8e9789f86ad2139f84af87a9a1

Observation 6d57ffeb-8354-48c8-b6a0-a6e84ac19d02 · outbound

This paper cites Mask CTC: Non-Autoregressive End-to-End ASR with CTC and Mask Predict.

A Non-autoregressive Model for Joint STT and TTS Mask CTC: Non-Autoregressive End-to-End ASR with CTC and Mask Predict

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-10T20:15:28.738910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:15:28.416162Z digest=sha256:cb8d5df92d51b64822b5ffd8f2525ccf6a07c26470548203dc0af2a222e39e7b

Observation a5baa8fd-78f8-46dc-bc2c-7badce02f133 · outbound

This paper cites BERT Meets CTC: New Formulation of End-to-End Speech Recognition with Pre-trained Masked Language Model.

A Non-autoregressive Model for Joint STT and TTS BERT Meets CTC: New Formulation of End-to-End Speech Recognition with Pre-trained Masked Language Model

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-10T20:15:28.715509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:15:28.421118Z digest=sha256:b174c1ea600d5e52d05dc60e25ad1b8dc5cd0ee5ebb71bf64a5dd2ac39053950

Observation 2e82dfc9-12b7-4d54-9ae0-f774b1688393 · outbound

This paper cites Bectra: Transducer-based end-to-end asr with bert-enhanced encoder,.

A Non-autoregressive Model for Joint STT and TTS Bectra: Transducer-based end-to-end asr with bert-enhanced encoder,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:15:29.040118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:15:28.426110Z digest=sha256:4c078c408d309e7d40f3d7ba9bc2badee0bc75d7bb6819d546024ecf38e7746c

Observation 52ffb3d8-360b-4cb8-bd8d-dddbaa6cada6 · outbound

This paper cites Mask-conformer: Augmenting conformer with mask-predict decoder,.

A Non-autoregressive Model for Joint STT and TTS Mask-conformer: Augmenting conformer with mask-predict decoder,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:15:29.022217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:15:28.430810Z digest=sha256:7903a3083a570ec9ea99da86a87e9158df3d8a580a0c4df583e8d6fa7508cb15

Observation 5e40e57e-16cc-46c1-93a4-c962b429327d · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech representations,.

A Non-autoregressive Model for Joint STT and TTS wav2vec 2.0: A framework for self-supervised learning of speech representations,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:28.435399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:28.435399Z digest=sha256:e7e535316152fb8ca1822af903a2ac919923b06316e32941bd61792fd39937ba

Observation afc07b86-e777-44ab-b611-e14c315d9f89 · outbound

This paper cites Layer normalization,.

A Non-autoregressive Model for Joint STT and TTS Layer normalization,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:28.440501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:28.440501Z digest=sha256:75f6b16e86bb6e05981fc6c10bd15c076fbf90cf743924791d1f0a575a775ab7

Observation 7279e4c5-b81f-4fdb-b014-702de35ec5c3 · outbound

This paper cites UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022.

A Non-autoregressive Model for Joint STT and TTS UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:28.445462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:28.445462Z digest=sha256:269c1d3dc735e596e3b1fc0d8bafe38fae3e4d78a0ab0349639a744dab5c74a0

Observation 3d529b1a-5af5-4d0c-a21c-e91333cc6ae9 · outbound

This paper cites Robust speech recognition via large-scale weak supervision,.

A Non-autoregressive Model for Joint STT and TTS Robust speech recognition via large-scale weak supervision,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:28.450541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:28.450541Z digest=sha256:f574f169961d3e4c064a003720b917e66049739c8504fd42516759e2d78f225b

Observation 82b3c130-1de6-46c1-aa51-3fb9b5a1f987 · outbound

This paper cites ESPnet-SPK: full pipeline speaker embedding toolkit with reproducible recipes, self-supervised front-ends, and off-the-shelf models.

A Non-autoregressive Model for Joint STT and TTS ESPnet-SPK: full pipeline speaker embedding toolkit with reproducible recipes, self-supervised front-ends, and off-the-shelf models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:28.455695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:28.455695Z digest=sha256:d8dc62dda2bd4c486f5a5d37ea8c385165ebb6c59323a32650eff13ce88f4a81

Observation 96cce75b-7029-453f-9f93-4549ae9cf4fe · outbound

This paper cites The lj speech dataset,.

A Non-autoregressive Model for Joint STT and TTS The lj speech dataset,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:15:28.967736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:15:28.460414Z digest=sha256:ced897a01a8acefd1c787eb4e1fbca54002c257b729e63bcb85b00b3b7fbb5ba

Observation 348305b0-a045-4aff-8f80-e15c448bc940 · outbound

This paper cites LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech.

A Non-autoregressive Model for Joint STT and TTS LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:28.464814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:28.464814Z digest=sha256:1255a5d8f41967e386a7b7e0cda5fa12f6b35f588f2c44628ce35c964cb669fe

Observation 442bc79f-1de7-4d20-913a-e655ee78199b · outbound

This paper cites Librispeech: an asr corpus based on public domain audio books,.

A Non-autoregressive Model for Joint STT and TTS Librispeech: an asr corpus based on public domain audio books,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:28.469522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:28.469522Z digest=sha256:227399530034cf5b30e3083f1c054ed671877d36a651e3692f837b2bcb2dcc02

Observation 7a0bd6be-8bba-440d-92d7-eeb78616a12b · outbound

This paper cites LibriTTS-R: A Restored Multi-Speaker Text-to-Speech Corpus.

A Non-autoregressive Model for Joint STT and TTS LibriTTS-R: A Restored Multi-Speaker Text-to-Speech Corpus

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:28.474447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:28.474447Z digest=sha256:d382c4060502ba53fb328569a5b530961b114c90d8c5c5c1739bce100b884225

Observation dba533ba-213b-44b6-b407-465ed4750d52 · outbound

This paper cites Conformer: Convolution-augmented Transformer for Speech Recognition.

A Non-autoregressive Model for Joint STT and TTS Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:28.480440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:28.480440Z digest=sha256:3bbfcf822670639248d72e09781363ce082c7e1585cb3da7f058926e316a42c8

Observation 18d15006-6a65-46be-87ac-9ce3b30c07ba · outbound

This paper cites ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification.

A Non-autoregressive Model for Joint STT and TTS ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:28.485951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:28.485951Z digest=sha256:c2354302172f3af440b658d75f3652936f630ec161fd23a29c29ed09b9e71ddf

Observation 60d7cb1f-ff83-4449-a370-e8cf45f47c44 · outbound

This paper cites Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,.

A Non-autoregressive Model for Joint STT and TTS Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:28.491768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:28.491768Z digest=sha256:b5118bfb7dbad53a3d1592bbf029229cbee8c4e5c7c4e7584063bc539d0b82ab

Observation 7701a5c4-ff0c-4c81-bba2-b716a10662c7 · outbound

This paper cites SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition.

A Non-autoregressive Model for Joint STT and TTS SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:28.496296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:28.496296Z digest=sha256:72e8984f44eb37a4a819f57464c8af676aec59072f0ab3f1a4f0ca58b4818da8

Observation 580a110f-fe7f-4fd9-9ad7-f31e81b52e7c · outbound

This paper cites Audio augmentation for speech recognition.

A Non-autoregressive Model for Joint STT and TTS Audio augmentation for speech recognition

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:28.500917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:28.500917Z digest=sha256:6a25a140092c40890af7fc4600e68b5b8b8680cf5572e55d0fd0853e299a030b

Observation a30e41ff-9fd7-4bbd-a147-5d29be044c01 · outbound

This paper cites Super-convergence: Very fast training of neural networks using large learning rates,.

A Non-autoregressive Model for Joint STT and TTS Super-convergence: Very fast training of neural networks using large learning rates,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:28.506312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:28.506312Z digest=sha256:1d6e19a880bea13cd226c83c52fd5b9afa1c4d11c403a2db5f4b76346664329a

Observation d0fdf2ab-b8e7-4cb6-b0e6-66a6fa579f6f · outbound

This paper cites Rethinking the inception architecture for computer vision,.

A Non-autoregressive Model for Joint STT and TTS Rethinking the inception architecture for computer vision,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:28.510887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:28.510887Z digest=sha256:5d45a344f3785e5eee0caf7e7a4a11a3816c28c15c7be566ff5cd0fd5223ee69

Observation f4c980a6-9f46-484e-95ff-654787756f49 · outbound

This paper cites Regularization of neural networks using dropconnect,.

A Non-autoregressive Model for Joint STT and TTS Regularization of neural networks using dropconnect,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:15:28.891532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:15:28.516216Z digest=sha256:ad86b48ecb90a6e46383dd1cfdb67cf0968ecfcb6e6d78a2f47f6cb910ef5b71

Observation 004df12f-cf94-43a1-b93f-989747bbc029 · outbound

This paper cites Sequence noise injected training for end-to-end speech recognition,.

A Non-autoregressive Model for Joint STT and TTS Sequence noise injected training for end-to-end speech recognition,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:15:28.873661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:15:28.520889Z digest=sha256:b9955fa3356deb2ff12b3362c5d5c0f6911d6587f7b15ff20b762b8d86e34bc0

Observation 9a06238f-8bab-4481-971b-91cb02c8bc78 · outbound

This paper cites Relaxing the Conditional Independence Assumption of CTC-based ASR by Conditioning on Intermediate Predictions.

A Non-autoregressive Model for Joint STT and TTS Relaxing the Conditional Independence Assumption of CTC-based ASR by Conditioning on Intermediate Predictions

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:28.525732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:28.525732Z digest=sha256:3901d5ae20f98912636c10103ec388213af9e168587fd904f277a67e27dde6a5

Observation 25f1d255-6d72-4087-87ff-f4a4ff22e5fa · outbound

This paper cites FastSpeech 2: Fast and High-Quality End-to-End Text to Speech.

A Non-autoregressive Model for Joint STT and TTS FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:28.530654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:28.530654Z digest=sha256:c5129dbff5e3aced6d54565f6e19a08297217d3b225e958d7d830aef719ed794

Pith citing papers

No inbound Pith citation observations are available.