Pith. sign in

Paper Citation Record · LEDGER

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition

As of 20 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 0 inbound Pith citation observations for arXiv:2505.16972.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.16972 v1

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:55:58.732055Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

57 of 57 outbound references displayed

  • verified exact1
  • verified fuzzy2
  • unresolved54
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 749b4bef-0a31-4856-a0bc-f8e3aeb8a981 · outbound

This paper cites online" 'onlinestring :=.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition online" 'onlinestring :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:53.174990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:53.174990Z digest=sha256:04c33cd47d0cbd3931b49f161cfcdd7a8a51b425cd15f2aff609df26fe83ad00

Observation d5a9b5a6-989a-4d16-b554-ddef40519e90 · outbound

This paper cites write newline.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:53.259170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:53.259170Z digest=sha256:6536e9705a525184978b66023f2ec72d783a38ce186ca5862cbb0c51be7f2ed8

Observation b244d73b-3bb0-4fdb-88d5-9139b5b7122f · outbound

This paper cites an unresolved cited work.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:56:03.915518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T14:55:53.443183Z digest=sha256:c9ef845e9d0e60374ecce48f9529c9222cd145a3da99858cde2413dbcda4d9e6

Observation d7b3252d-7557-4296-8eef-e011c7e64fa4 · outbound

This paper cites an unresolved cited work.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:56:03.815728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T14:55:53.583069Z digest=sha256:9887f86104fb73c14a2997041c2459a43e5a8d19164e7477d8994dbaa69defae

Observation 51065def-f9b0-4641-bc23-c1fe90918aba · outbound

This paper cites Common Voice: A Massively-Multilingual Speech Corpus.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Common Voice: A Massively-Multilingual Speech Corpus

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:53.700409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:53.700409Z digest=sha256:7e1a36c71405170162d9024d3b8449b51d30e794a9d057a6d55f1159944c070b

Observation 2c8ad1b7-942e-4d74-90aa-0c1b497eb2c8 · outbound

This paper cites Synthetic Data from Diffusion Models Improves ImageNet Classification.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Synthetic Data from Diffusion Models Improves ImageNet Classification

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:53.847905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:53.847905Z digest=sha256:6935c9fa1e8db4fb3e31f86cfd2da9c524f84207ff81afc0cb7f17a8fa99cfb1

Observation 07068995-31a3-45ad-8482-05f6f88c067f · outbound

This paper cites an unresolved cited work.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:56:03.670409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T14:55:53.982528Z digest=sha256:45f865073e7ef367a579ad0ec6c92dfd4884a08360234f40504db6bd35404e7b

Observation a52c2ee2-3e01-4230-88d4-43cd45c4dc80 · outbound

This paper cites an unresolved cited work.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:56:03.532208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T14:55:54.154721Z digest=sha256:4e4b024addf66d29ff01b668e8904445976036a6a185ea36fa17d77b207ee0d6

Observation 463240d4-c012-4afd-8079-e7ebc5705b0d · outbound

This paper cites wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:54.314575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:54.314575Z digest=sha256:c51529beb2efce07aa60e967a703a68162a56d624b428d79a5ef002321a378b4

Observation 888acafe-876e-438d-99b9-60c6a4d3cabd · outbound

This paper cites Seed-ASR: Understanding Diverse Speech and Contexts with LLM-based Speech Recognition.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Seed-ASR: Understanding Diverse Speech and Contexts with LLM-based Speech Recognition

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:54.427250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:54.427250Z digest=sha256:19a76ec555149e97db0b86a433e1fa0144822d1f74c14c00573ea92049ae233f

Observation 15b3e51d-1b6c-483c-9af9-a277b079a49c · outbound

This paper cites an unresolved cited work.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:56:03.383130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T14:55:54.498525Z digest=sha256:18da02d24ed60e426dbbb2d42ee0fa547a47e2aa07d4ead8ad8dbee8c02c6abd

Observation 35e732a5-8035-43a1-8e9c-66ff798c8af9 · outbound

This paper cites an unresolved cited work.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:56:03.244325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T14:55:54.566641Z digest=sha256:79f7468a8f038cbb5802e565ee0f0a69542ab4800558d6fd2414a62085020c8a

Observation fb412036-97d5-4785-9018-d1fe2474c66e · outbound

This paper cites XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:54.665380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:54.665380Z digest=sha256:f3573e01b109de15393002c2c1fbf05312d58d797106143c608092eeee21d7f2

Observation 9b65f5b2-3f11-49c6-800e-ef63fd2fe423 · outbound

This paper cites Towards Robust Speech Representation Learning for Thousands of Languages.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Towards Robust Speech Representation Learning for Thousands of Languages

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:54.752647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:54.752647Z digest=sha256:9605fdca1ba8749eae99a74f8687b3bd12d58f2bb33ca6bf83507748e66e6c00

Observation ae97234f-0160-422c-af4e-9c536861be1f · outbound

This paper cites Towards Achieving Human Parity on End-to-end Simultaneous Speech Translation via LLM Agent.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Towards Achieving Human Parity on End-to-end Simultaneous Speech Translation via LLM Agent

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:54.842142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:54.842142Z digest=sha256:dfa5216bad1c697f8021b68128a225be6b3bdd1a3eb757aae0ebf8f7f69ce842

Observation 851d97c0-f3eb-4d2c-b8c6-bb67bb7072d4 · outbound

This paper cites Seamless: Multilingual Expressive and Streaming Speech Translation.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:54.912281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:54.912281Z digest=sha256:21458933e0c38a584d744fcc822359e87960ce2adbf5e866c81dca6bc2c09303

Observation f2b799f4-8ca8-453e-9fba-cd4eaae2f49d · outbound

This paper cites an unresolved cited work.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:56:03.105039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T14:55:54.987167Z digest=sha256:34da810b9a7064439d0d018a894342dd2ae2555314a68381f8d6349f888d6425

Observation 3c0638db-9725-4da7-8d96-67e4b858b084 · outbound

This paper cites No Language Left Behind: Scaling Human-Centered Machine Translation.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition No Language Left Behind: Scaling Human-Centered Machine Translation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:55.065966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:55.065966Z digest=sha256:dc08c50f47cf7bd4e2a9fe048051fa80fb420332a6fe46108ed4f005cec7377c

Observation bb08feb4-246b-4e5e-b096-79b083b6457d · outbound

This paper cites CML-TTS A Multilingual Dataset for Speech Synthesis in Low-Resource Languages.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition CML-TTS A Multilingual Dataset for Speech Synthesis in Low-Resource Languages

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:55:59.809037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T14:55:55.134893Z digest=sha256:4fdf67d4429aa11b40160861d620179d8641d424689de4708339c7799a57fe63

Observation 42604363-4e17-44da-9d12-57971fd4abbc · outbound

This paper cites Understanding Back-Translation at Scale.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Understanding Back-Translation at Scale

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:55.225525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:55.225525Z digest=sha256:50b3eda75f5b2de32a3d06188377c21e217bf3503b2168ddbd66d384c4c065bc

Observation 614ab78c-b34f-489d-b9a8-3d7585b2a781 · outbound

This paper cites an unresolved cited work.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:56:02.983278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T14:55:55.298928Z digest=sha256:1e495793b10e9f01960730f9b0233a916393343f76549cf6b33f9555944ffbb9

Observation d695a88a-0cf4-4a2e-9a08-59f3d2684eed · outbound

This paper cites Hasegawa-Johnson, Shiyu Chang, and Yang Zhang.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Hasegawa-Johnson, Shiyu Chang, and Yang Zhang

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:56:02.806913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T14:55:55.388526Z digest=sha256:99af791cd628a221f83360cc6420e11d905d88d681dcd3e70cabffd4c64d966c

Observation d9558d3f-fb26-44cf-af03-08a9bafb6415 · outbound

This paper cites Textbooks Are All You Need.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Textbooks Are All You Need

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:55.478048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:55.478048Z digest=sha256:6e37149857b583dc97ec454c12cb87b722140c81e2f0fa95667fd6bb0734858b

Observation 373baf00-f961-45f3-b761-91202f53915b · outbound

This paper cites an unresolved cited work.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:56:02.684344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T14:55:55.542956Z digest=sha256:6721b3fe23cacc54157b9389265118edf856c865964150eadfe95852c604f086

Observation aa6c3185-dcb8-4aa5-93ae-5dd8727a85ab · outbound

This paper cites Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:55.604407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:55.604407Z digest=sha256:418c5184b512f84fa4e6412af30575a969dd01c6f11bb376e79dabdc6865d4a7

Observation 2a34c5b7-44a9-4d22-90ca-e3fe94a7f1f4 · outbound

This paper cites an unresolved cited work.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:56:02.482808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T14:55:55.661496Z digest=sha256:bbe118b9a93554a084a794f3f6289ab0ee825f7ea9f59e93f91b0e4f5a472a8d

Observation edee7b2d-1094-48e3-9e5f-16779cfea925 · outbound

This paper cites Speech Translation with Large Language Models: An Industrial Practice.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Speech Translation with Large Language Models: An Industrial Practice

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:55.748608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:55.748608Z digest=sha256:1f6113cbdd4f7e5672e6e4e590058ebc69e4ec68b53e49536df23ec4909ffaab

Observation 57c6908c-707b-4e18-a079-a69d3e6608b0 · outbound

This paper cites an unresolved cited work.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:56:02.381898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T14:55:55.855819Z digest=sha256:cfc36a24f680773ea7a777a70394cdbdfa26c00e28d7e670512753ee67909c58

Observation 968f3035-3656-43ed-aa5d-cb8f91ee8ba5 · outbound

This paper cites HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:55.933535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:55.933535Z digest=sha256:7251778d34441085878275047c78d92b0f28309afba2c2a4e27cef22bd35eae7

Observation b1a35c0c-15e6-4e3b-931c-20a8f00fe4df · outbound

This paper cites an unresolved cited work.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:56:02.230660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T14:55:56.043409Z digest=sha256:ee5e13f9a60927bdd3f0398a59f62d6f29ac97046130da40a508644472ca5053

Observation 704475c2-887a-4818-b96d-c4cba4812de1 · outbound

This paper cites Synthetic Data (Almost) from Scratch: Generalized Instruction Tuning for Language Models.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Synthetic Data (Almost) from Scratch: Generalized Instruction Tuning for Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:56.193027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:56.193027Z digest=sha256:695f43de5121c681a0dd1cdf16ea5d4dfa5d6454ec92dcc1117e6dba1cbf4306

Observation 3d92d0db-9563-4a93-aa41-6c2dace948a2 · outbound

This paper cites an unresolved cited work.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:56:02.087331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T14:55:56.305477Z digest=sha256:a26e566eecea16a1cb44a578a8b53b6cf5240726066e18704cc99b810e458703

Observation 0a8853c1-bc77-4cc2-b259-aa41ef251803 · outbound

This paper cites Massively Multilingual ASR: 50 Languages, 1 Model, 1 Billion Parameters.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Massively Multilingual ASR: 50 Languages, 1 Model, 1 Billion Parameters

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:56.404903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:56.404903Z digest=sha256:95a794156cadda8385576fb38c09e191af49d3556f01ff5f9282e9d952c6bca3

Observation 48978583-1258-41b0-8653-d9ddca2174a2 · outbound

This paper cites Scaling Speech Technology to 1,000+ Languages.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Scaling Speech Technology to 1,000+ Languages

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:56.498003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:56.498003Z digest=sha256:ed508405f60d7bf218d036e653792711f47a2759a9959a7f87aed18e0e36a762

Observation 875e3c76-392e-4d58-b291-005c49d5375b · outbound

This paper cites MLS: A Large-Scale Multilingual Dataset for Speech Research.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition MLS: A Large-Scale Multilingual Dataset for Speech Research

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:56.615438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:56.615438Z digest=sha256:a648a5608d5a66aa9b4b61fe553fb4e094a6e2722c47467fbcb9ede4ba846b80

Observation 99b7891b-7077-4c39-b587-7e6c1e7f1550 · outbound

This paper cites Less is More: Accurate Speech Recognition & Translation without Web-Scale Data.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Less is More: Accurate Speech Recognition & Translation without Web-Scale Data

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:56.665186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:56.665186Z digest=sha256:a01ba91014a7441bec8c87cddd640922d13d5988afd8dcb5be5fb4100211bdaf

Observation 7489012e-c329-4f5a-8be7-8331d9ac307c · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Robust Speech Recognition via Large-Scale Weak Supervision

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:56.715907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:56.715907Z digest=sha256:1d4292ecded0b90cf086c4ea522be25e893cdd5d7a6583b2cd61b1b326122501

Observation eb7f802f-3fe0-4d57-babc-bef3dcafee9f · outbound

This paper cites an unresolved cited work.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:56:01.957908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T14:55:56.774284Z digest=sha256:2c692ae4b2f0b0927f1d2381befd1829dda4c4cfbc35e30ac2a0c3e871584860

Observation 259d6162-e9a0-4a07-bb3d-090db4b295ae · outbound

This paper cites Blattmann, Dominik Lorenz, Patrick Esser, and Bj \"o rn Ommer.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Blattmann, Dominik Lorenz, Patrick Esser, and Bj \"o rn Ommer

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:56:01.810253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T14:55:56.858264Z digest=sha256:d02b255c0ec9fedaaabf2cb968f3cfebd6acd335355b69b6aa379cc0581ac005

Observation 31a069dc-96b6-4896-b176-9adfe6b9f8b1 · outbound

This paper cites an unresolved cited work.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:56:01.596909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T14:55:56.967930Z digest=sha256:f71c3c23a54f86e5de07dfebfbef11ec1264c2cedd98ba4522431690459ebe68

Observation 48d01bbb-d7de-4888-b977-9ccb38332a0e · outbound

This paper cites an unresolved cited work.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:56:01.413813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T14:55:57.112780Z digest=sha256:e99b9e2052af979f564f1f64d8affc2ade72bccda1668e2e2149d79943ea1bc0

Observation fa95a800-6595-417b-af9a-f94e5bea44ef · outbound

This paper cites Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:57.198647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:57.198647Z digest=sha256:b8974ab1c18339e5aa3dd67b0dd32aec031669642ccf8b0225d6d1336c4d1787

Observation 3e8cffa8-06ad-475f-8c14-b3629bb6e103 · outbound

This paper cites an unresolved cited work.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:56:01.264342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T14:55:57.312465Z digest=sha256:70188a6eea4da335dfcf4f4ff2d21f12f3a9104cfee7c3aef41a9439afb5ee8c

Observation f7a78e29-80ab-4c3d-b9aa-f77f7ab8bf71 · outbound

This paper cites StableRep: Synthetic Images from Text-to-Image Models Make Strong Visual Representation Learners.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition StableRep: Synthetic Images from Text-to-Image Models Make Strong Visual Representation Learners

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:57.433402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:57.433402Z digest=sha256:d3e72e7802d86bb688fd549369717704c8ae47b49751c642b787bdeb813ef8f0

Observation 6b8e309e-b3d5-48c3-8f9d-c4a615853a3d · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition LLaMA: Open and Efficient Foundation Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:57.537044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:57.537044Z digest=sha256:5b7952e8768d7614811d0aa72d64ddc9a554e2707653b2f3cbaaf5edaf0f3062

Observation e9b37409-a26b-4b71-919c-2247ad31a3d0 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:57.615225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:57.615225Z digest=sha256:89885c7cb34de6b7179ede021c4fb82358c30d49e7a35696d4a15f8e15f40c57

Observation 1b39a7ae-1242-48e5-a95f-f55e7ee62b94 · outbound

This paper cites Effective Data Augmentation With Diffusion Models.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Effective Data Augmentation With Diffusion Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:57.729459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:57.729459Z digest=sha256:4b7217be736fbb8015733439996d4f4f5832907d9391752c4ed0953f15bfe5c0

Observation b12b32fd-de54-4556-817b-b2c42d11dca1 · outbound

This paper cites VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:57.846781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:57.846781Z digest=sha256:e3e0e49da9a0a35c533179dd83947f6d39cfbce5138bd0166b37b47610fb54fa

Observation 9dd89f1e-ac5d-4aed-99c1-9262f5d8728a · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:57.998249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:57.998249Z digest=sha256:08b7196cf3af588c5856554bc427ebc91a69366273967741860f8829d9ec550d

Observation 14196f2b-ee01-4167-89f4-46fe289857e1 · outbound

This paper cites an unresolved cited work.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:56:01.063964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T14:55:58.086002Z digest=sha256:5c5505dcfb44940185ae3bbca0acdafe0f04b6a43cb5e4cdf23f8d11ad22222b

Observation 0b620978-fab8-4bfd-927c-dcf3d9c6e675 · outbound

This paper cites an unresolved cited work.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:56:00.913779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T14:55:58.222753Z digest=sha256:854c765ba612773be4f0bc559b6754023ca2179929ce379355682acdb046cf03

Observation cb572e1c-c837-40e5-b3bf-23355b8c2610 · outbound

This paper cites an unresolved cited work.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:56:00.730417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T14:55:58.293114Z digest=sha256:8c7a8c335c39d43236a191a13a514118b18130c1a4255d8e73b92ecc3f5ae3e4

Observation 298d3a94-989d-42fc-ba0b-9a1891f70374 · outbound

This paper cites Skywork: A More Open Bilingual Foundation Model.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Skywork: A More Open Bilingual Foundation Model

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:58.395604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:58.395604Z digest=sha256:44c5088fc69bf415c315d5c0d2eec8007e497af4ae630baf784ce84599605645

Observation 6b6f0483-c670-46c6-8a84-3915bb604f78 · outbound

This paper cites Enhancing Low-Resource ASR through Versatile TTS: Bridging the Data Gap.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Enhancing Low-Resource ASR through Versatile TTS: Bridging the Data Gap

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:58.468368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:58.468368Z digest=sha256:773c58531c006acb91cdb338e143921034e0fea7db38e2a23041dca897a1676b

Observation 230d8c32-385a-4d69-bc79-a1f6c7215863 · outbound

This paper cites an unresolved cited work.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:56:00.581979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T14:55:58.540112Z digest=sha256:3262e6b0b1a1f0edcd996f085b51f10b044e97c7889dcbc3e0b3cf41d400564c

Observation ed7f5382-0d3b-4d65-a29d-a64967f2ff51 · outbound

This paper cites an unresolved cited work.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:56:00.411285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T14:55:58.638358Z digest=sha256:9e3a7d65e3fece107805c9a899b50b6657ada3243631891190cecaeada1d129d

Observation 04e01395-1449-4357-81d7-0b687c7b1a02 · outbound

This paper cites an unresolved cited work.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:56:00.263139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T14:55:58.732055Z digest=sha256:c3f05f36f33a7dadc74b886ac2041580ce12594da91981d85b3f5331aced12e4

Pith citing papers

No inbound Pith citation observations are available.