{"work":{"id":"ed5fbefe-24bf-436f-98c1-68e9114360bf","openalex_id":"https://openalex.org/W4311000453","doi":"10.48550/arxiv.2212.04356","arxiv_id":"2212.04356","raw_key":null,"title":"Robust Speech Recognition via Large-Scale Weak Supervision","authors":null,"authors_text":"Zhang, Y","year":2022,"venue":"eess.AS","abstract":"We study the capabilities of speech processing systems trained simply to predict large amounts of transcripts of audio on the internet. When scaled to 680,000 hours of multilingual and multitask supervision, the resulting models generalize well to standard benchmarks and are often competitive with prior fully supervised results but in a zero-shot transfer setting without the need for any fine-tuning. When compared to humans, the models approach their accuracy and robustness. We are releasing models and inference code to serve as a foundation for further work on robust speech processing.","external_url":"https://arxiv.org/abs/2212.04356","cited_by_count":1161,"metadata_source":"pith","metadata_fetched_at":"2026-08-05T02:28:24.338817+00:00","pith_arxiv_id":"2212.04356","created_at":"2026-05-08T18:08:53.435901+00:00","updated_at":"2026-08-05T02:28:24.338817+00:00","title_quality_ok":true,"display_title":"Robust Speech Recognition via Large-Scale Weak Supervision","render_title":"Robust Speech Recognition via Large-Scale Weak Supervision"},"hub":{"state":{"work_id":"ed5fbefe-24bf-436f-98c1-68e9114360bf","tier":"super_hub","tier_reason":"100+ Pith inbound or 10,000+ external citations","pith_inbound_count":121,"external_cited_by_count":1161,"distinct_field_count":17,"first_pith_cited_at":"2023-05-02T17:38:21+00:00","last_pith_cited_at":"2026-07-09T17:24:30+00:00","author_build_status":"needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"needed","reader_status":"not_needed","recognition_status":"not_needed","updated_at":"2026-08-22T15:59:23.354822+00:00","tier_text":"super_hub"},"tier":"super_hub","role_counts":[{"context_role":"background","n":8},{"context_role":"method","n":8},{"context_role":"baseline","n":1}],"polarity_counts":[{"context_polarity":"background","n":8},{"context_polarity":"use_method","n":8},{"context_polarity":"baseline","n":1}],"runs":{"ask_index":{"job_type":"ask_index","status":"succeeded","result":{"title":"Robust Speech Recognition via Large-Scale Weak Supervision","claims":[{"claim_text":"We study the capabilities of speech processing systems trained simply to predict large amounts of transcripts of audio on the internet. When scaled to 680,000 hours of multilingual and multitask supervision, the resulting models generalize well to standard benchmarks and are often competitive with prior fully supervised results but in a zero-shot transfer setting without the need for any fine-tuning. When compared to humans, the models approach their accuracy and robustness. We are releasing models and inference code to serve as a foundation for further work on robust speech processing.","claim_type":"abstract","evidence_strength":"source_metadata"},{"claim_text":"phy (see Figure 1 for system overview). All data streams are synchronized via Lab Streaming Layer (LSL [40, 41]) with event markers inserted at critical training milestones, ensuring precise temporal alignment. 3.2. Analysis Streams Verbal CommunicationThe verbal content of the audio signal is extracted by faster-whisper [ 42], an optimized reimplementation of OpenAI's Whisper model [43]. Insults are detected by keyword recognition. For complex classification tasks, a transformer-based language ","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"First, systems must maintain low end-to-end latency to support interactive applications such as live captioning, meeting transcription, and voice-driven agents [5, 6]. Second, they must exhibit bounded computational and memory footprints over extended sessions; unbounded resource growth renders long-duration deployments infeasible on resource-constrained hardware [7, 8]. Third, streaming syst ems must produce stable incremental outputs, minimizing hypothesis revisions that degrade user experienc","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"Full construction details, including the prompt-audio pool composition and quality control pipeline, are provided in Appendix D. Evaluation Models and JudgingWe curate a model set with demonstrated strong audio understand- ing. Guided by MMSU [1] and MMAU-Pro [2], two benchmarks that emphasize paralinguistic and sound-mixture reasoning, we select open models (Qwen3-Omni [31], Mimo-Audio [32], Kimi-Audio [33]) and closed models (Gemini-3-Pro, Gemini-3-Flash, GPT-4o-Audio [34]). For Qwen3-Omni and","claim_type":"baseline","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"By leveraging a model to automati- cally label a larger unlabelled corpus, the training distribution is significantly expanded without incurring the cost of addi- tional manual labelling. This paper extends our preliminary work [13] regarding semi-supervised confidence detection, where the integration of neural embeddings from the Whisper encoder [14] with a set of handcrafted features was previously explored. Whisper was prioritised over acoustic self-supervised learning models such as Wav2Vec ","claim_type":"method","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"various aspects such as actions [46, 20, 21] (with InternVideo [46]), objects [45, 15], object annota- tions with positions [48], and more. While the majority of these models' outputs are comparatively independent, we utilize the pretrained T5 language model [34] to refine their descriptions for improved clarity. Moreover, we integrate the Whisper [33] speech recognition model into VideoChat-Text to capitalize on audio data within videos, further enhancing the richness of video descriptions. 3.1","claim_type":"method","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"Full details of both analyses are provided in Appendix A.2. 2.4 Listened decoder We trained a contrastive MEG-to-word decoder to identify words fromlistenedMEG responses. Word-level onset timestamps were obtained using forced alignment of the poem audio recordings to their ground truth transcripts, using WhisperX [30], which combines Whisper-based [31] segment 4 boundary detection with a Wav2Vec2 phoneme alignment model [ 32]. For each word onset, a 1-second MEG window was extracted spanning 200","claim_type":"method","confidence":0.85,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks Robust Speech Recognition via Large-Scale Weak Supervision because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (8 contexts).","role_counts":[{"n":8,"context_role":"background"},{"n":8,"context_role":"method"},{"n":1,"context_role":"baseline"}]},"error":null,"updated_at":"2026-07-03T03:43:19.408324+00:00"},"author_expand":{"job_type":"author_expand","status":"succeeded","result":{"authors_linked":[{"id":"fbeae6cf-f1aa-414c-8871-7056f33c24cc","orcid":null,"display_name":"Zhang"}]},"error":null,"updated_at":"2026-07-03T03:43:19.455441+00:00"},"context_extract":{"job_type":"context_extract","status":"succeeded","result":{"enqueued_papers":25},"error":null,"updated_at":"2026-05-14T17:59:32.412124+00:00"},"graph_features":{"job_type":"graph_features","status":"succeeded","result":{"co_cited":[{"title":"GPT-4 Technical Report","work_id":"b928e041-6991-4c08-8c81-0359e4097c7b","shared_citers":4},{"title":"Qwen3-Omni Technical Report","work_id":"ae43e594-8bab-4471-b6af-92a300f6a048","shared_citers":4},{"title":"Decoupled Weight Decay Regularization","work_id":"07ef7360-d385-4033-83f7-8384a6325204","shared_citers":3},{"title":"In: 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","work_id":"9da51225-b7bd-4032-b7db-ca577971dafe","shared_citers":3},{"title":"LLaMA: Open and Efficient Foundation Language Models","work_id":"c018fc23-6f3f-4035-9d02-28a2173b2b9d","shared_citers":3},{"title":"Qwen3-ASR Technical Report","work_id":"db50e258-4a3d-4141-ba30-f76f8f953880","shared_citers":3},{"title":"Scaling speech technology to 1,000+ languages","work_id":"d85c5ee3-0310-48b2-b8c8-a24601ce26c1","shared_citers":3},{"title":"The Llama 3 Herd of Models","work_id":"1549a635-88af-4ac1-acfe-51ae7bb53345","shared_citers":3},{"title":"AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning","work_id":"1f9d1d3b-a6d6-45a9-9f13-51393c03be8a","shared_citers":2},{"title":"BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding","work_id":"ed240a10-5b19-406c-baa5-30803f465785","shared_citers":2},{"title":"CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer","work_id":"f38fc088-12aa-4bf4-9ecd-08d3e797ccb7","shared_citers":2},{"title":"CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models","work_id":"3af84775-3b81-4078-b553-52739aae03ba","shared_citers":2},{"title":"Cosyvoice 3: Towards in-the-wild speech generation via scaling-up and post-training","work_id":"ce66e767-526a-419d-bb95-ac019edf4050","shared_citers":2},{"title":"Cosyvoice: A scalable multi- lingual zero-shot text-to-speech synthesizer based on supervised semantic tokens","work_id":"e5ad925a-4045-49b5-b301-208bcbf3eca8","shared_citers":2},{"title":"Ddp: Diffusion model for dense visual prediction","work_id":"b8a8bb9e-1d31-40e2-9cab-ae21e338dde6","shared_citers":2},{"title":"DeepSeek-V3 Technical Report","work_id":"57d2791d-2219-4c31-a077-afc04b12a75c","shared_citers":2},{"title":"Fish-speech: Leveraging large language models for advanced multilingual text-to- speech synthesis","work_id":"10ece45c-e63a-4ab2-b106-971d79be6fba","shared_citers":2},{"title":"Funaudiollm: V oice understanding and generation foundation models for natural interaction between humans and llms","work_id":"7cd6d289-dca2-414f-99e0-809f37c065fa","shared_citers":2},{"title":"Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities","work_id":"008df105-2fdd-45d8-857a-8e35868aecb6","shared_citers":2},{"title":"GPT-4o System Card","work_id":"f37bf1c7-4964-4e56-9762-d20da8d9009f","shared_citers":2},{"title":"Hannun, C","work_id":"eb57b6c8-18a7-4e51-b90a-2add68d4ee9e","shared_citers":2},{"title":"HunyuanVideo: A Systematic Framework For Large Video Generative Models","work_id":"881efa7e-7e73-4c66-9cc3-2803e551061c","shared_citers":2},{"title":"ImageBind One Embedding Space to Bind Them All","work_id":"b9701eca-d05e-4d2e-9045-6761df4ba175","shared_citers":2},{"title":"Jsonschemabench: A rigorous benchmark of structured outputs for language models","work_id":"0a01531a-03bb-463b-8bec-3777fe5a120b","shared_citers":2}],"time_series":[{"n":2,"year":2023},{"n":1,"year":2025},{"n":34,"year":2026}],"dependency_candidates":[]},"error":null,"updated_at":"2026-05-14T17:59:53.696225+00:00"},"identity_refresh":{"job_type":"identity_refresh","status":"succeeded","result":{"items":[{"title":"Qwen3 Technical Report","outcome":"unchanged","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","resolver":"local_arxiv","confidence":0.98,"old_work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e"}],"counts":{"fixed":0,"merged":0,"unchanged":1,"quarantined":0,"needs_external_resolution":0},"errors":[],"attempted":1},"error":null,"updated_at":"2026-05-14T17:59:24.605929+00:00"},"role_polarity":{"job_type":"role_polarity","status":"succeeded","result":{"title":"Robust Speech Recognition via Large-Scale Weak Supervision","claims":[{"claim_text":"We study the capabilities of speech processing systems trained simply to predict large amounts of transcripts of audio on the internet. When scaled to 680,000 hours of multilingual and multitask supervision, the resulting models generalize well to standard benchmarks and are often competitive with prior fully supervised results but in a zero-shot transfer setting without the need for any fine-tuning. When compared to humans, the models approach their accuracy and robustness. We are releasing models and inference code to serve as a foundation for further work on robust speech processing.","claim_type":"abstract","evidence_strength":"source_metadata"},{"claim_text":"phy (see Figure 1 for system overview). All data streams are synchronized via Lab Streaming Layer (LSL [40, 41]) with event markers inserted at critical training milestones, ensuring precise temporal alignment. 3.2. Analysis Streams Verbal CommunicationThe verbal content of the audio signal is extracted by faster-whisper [ 42], an optimized reimplementation of OpenAI's Whisper model [43]. Insults are detected by keyword recognition. For complex classification tasks, a transformer-based language ","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"First, systems must maintain low end-to-end latency to support interactive applications such as live captioning, meeting transcription, and voice-driven agents [5, 6]. Second, they must exhibit bounded computational and memory footprints over extended sessions; unbounded resource growth renders long-duration deployments infeasible on resource-constrained hardware [7, 8]. Third, streaming syst ems must produce stable incremental outputs, minimizing hypothesis revisions that degrade user experienc","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"Full construction details, including the prompt-audio pool composition and quality control pipeline, are provided in Appendix D. Evaluation Models and JudgingWe curate a model set with demonstrated strong audio understand- ing. Guided by MMSU [1] and MMAU-Pro [2], two benchmarks that emphasize paralinguistic and sound-mixture reasoning, we select open models (Qwen3-Omni [31], Mimo-Audio [32], Kimi-Audio [33]) and closed models (Gemini-3-Pro, Gemini-3-Flash, GPT-4o-Audio [34]). For Qwen3-Omni and","claim_type":"baseline","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"By leveraging a model to automati- cally label a larger unlabelled corpus, the training distribution is significantly expanded without incurring the cost of addi- tional manual labelling. This paper extends our preliminary work [13] regarding semi-supervised confidence detection, where the integration of neural embeddings from the Whisper encoder [14] with a set of handcrafted features was previously explored. Whisper was prioritised over acoustic self-supervised learning models such as Wav2Vec ","claim_type":"method","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"various aspects such as actions [46, 20, 21] (with InternVideo [46]), objects [45, 15], object annota- tions with positions [48], and more. While the majority of these models' outputs are comparatively independent, we utilize the pretrained T5 language model [34] to refine their descriptions for improved clarity. Moreover, we integrate the Whisper [33] speech recognition model into VideoChat-Text to capitalize on audio data within videos, further enhancing the richness of video descriptions. 3.1","claim_type":"method","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"Full details of both analyses are provided in Appendix A.2. 2.4 Listened decoder We trained a contrastive MEG-to-word decoder to identify words fromlistenedMEG responses. Word-level onset timestamps were obtained using forced alignment of the poem audio recordings to their ground truth transcripts, using WhisperX [30], which combines Whisper-based [31] segment 4 boundary detection with a Wav2Vec2 phoneme alignment model [ 32]. For each word onset, a 1-second MEG window was extracted spanning 200","claim_type":"method","confidence":0.85,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks Robust Speech Recognition via Large-Scale Weak Supervision because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (8 contexts).","role_counts":[{"n":8,"context_role":"background"},{"n":8,"context_role":"method"},{"n":1,"context_role":"baseline"}]},"error":null,"updated_at":"2026-07-03T03:43:19.405824+00:00"},"summary_claims":{"job_type":"summary_claims","status":"succeeded","result":{"title":"Robust Speech Recognition via Large-Scale Weak Supervision","claims":[{"claim_text":"We study the capabilities of speech processing systems trained simply to predict large amounts of transcripts of audio on the internet. When scaled to 680,000 hours of multilingual and multitask supervision, the resulting models generalize well to standard benchmarks and are often competitive with prior fully supervised results but in a zero-shot transfer setting without the need for any fine-tuning. When compared to humans, the models approach their accuracy and robustness. We are releasing models and inference code to serve as a foundation for further work on robust speech processing.","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks Robust Speech Recognition via Large-Scale Weak Supervision because it crossed a citation-hub threshold.","role_counts":[]},"error":null,"updated_at":"2026-05-14T18:00:11.376944+00:00"}},"summary":{"title":"Robust Speech Recognition via Large-Scale Weak Supervision","claims":[{"claim_text":"We study the capabilities of speech processing systems trained simply to predict large amounts of transcripts of audio on the internet. When scaled to 680,000 hours of multilingual and multitask supervision, the resulting models generalize well to standard benchmarks and are often competitive with prior fully supervised results but in a zero-shot transfer setting without the need for any fine-tuning. When compared to humans, the models approach their accuracy and robustness. We are releasing models and inference code to serve as a foundation for further work on robust speech processing.","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks Robust Speech Recognition via Large-Scale Weak Supervision because it crossed a citation-hub threshold.","role_counts":[]},"graph":{"co_cited":[{"title":"GPT-4 Technical Report","work_id":"b928e041-6991-4c08-8c81-0359e4097c7b","shared_citers":4},{"title":"Qwen3-Omni Technical Report","work_id":"ae43e594-8bab-4471-b6af-92a300f6a048","shared_citers":4},{"title":"Decoupled Weight Decay Regularization","work_id":"07ef7360-d385-4033-83f7-8384a6325204","shared_citers":3},{"title":"In: 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","work_id":"9da51225-b7bd-4032-b7db-ca577971dafe","shared_citers":3},{"title":"LLaMA: Open and Efficient Foundation Language Models","work_id":"c018fc23-6f3f-4035-9d02-28a2173b2b9d","shared_citers":3},{"title":"Qwen3-ASR Technical Report","work_id":"db50e258-4a3d-4141-ba30-f76f8f953880","shared_citers":3},{"title":"Scaling speech technology to 1,000+ languages","work_id":"d85c5ee3-0310-48b2-b8c8-a24601ce26c1","shared_citers":3},{"title":"The Llama 3 Herd of Models","work_id":"1549a635-88af-4ac1-acfe-51ae7bb53345","shared_citers":3},{"title":"AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning","work_id":"1f9d1d3b-a6d6-45a9-9f13-51393c03be8a","shared_citers":2},{"title":"BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding","work_id":"ed240a10-5b19-406c-baa5-30803f465785","shared_citers":2},{"title":"CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer","work_id":"f38fc088-12aa-4bf4-9ecd-08d3e797ccb7","shared_citers":2},{"title":"CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models","work_id":"3af84775-3b81-4078-b553-52739aae03ba","shared_citers":2},{"title":"Cosyvoice 3: Towards in-the-wild speech generation via scaling-up and post-training","work_id":"ce66e767-526a-419d-bb95-ac019edf4050","shared_citers":2},{"title":"Cosyvoice: A scalable multi- lingual zero-shot text-to-speech synthesizer based on supervised semantic tokens","work_id":"e5ad925a-4045-49b5-b301-208bcbf3eca8","shared_citers":2},{"title":"Ddp: Diffusion model for dense visual prediction","work_id":"b8a8bb9e-1d31-40e2-9cab-ae21e338dde6","shared_citers":2},{"title":"DeepSeek-V3 Technical Report","work_id":"57d2791d-2219-4c31-a077-afc04b12a75c","shared_citers":2},{"title":"Fish-speech: Leveraging large language models for advanced multilingual text-to- speech synthesis","work_id":"10ece45c-e63a-4ab2-b106-971d79be6fba","shared_citers":2},{"title":"Funaudiollm: V oice understanding and generation foundation models for natural interaction between humans and llms","work_id":"7cd6d289-dca2-414f-99e0-809f37c065fa","shared_citers":2},{"title":"Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities","work_id":"008df105-2fdd-45d8-857a-8e35868aecb6","shared_citers":2},{"title":"GPT-4o System Card","work_id":"f37bf1c7-4964-4e56-9762-d20da8d9009f","shared_citers":2},{"title":"Hannun, C","work_id":"eb57b6c8-18a7-4e51-b90a-2add68d4ee9e","shared_citers":2},{"title":"HunyuanVideo: A Systematic Framework For Large Video Generative Models","work_id":"881efa7e-7e73-4c66-9cc3-2803e551061c","shared_citers":2},{"title":"ImageBind One Embedding Space to Bind Them All","work_id":"b9701eca-d05e-4d2e-9045-6761df4ba175","shared_citers":2},{"title":"Jsonschemabench: A rigorous benchmark of structured outputs for language models","work_id":"0a01531a-03bb-463b-8bec-3777fe5a120b","shared_citers":2}],"time_series":[{"n":2,"year":2023},{"n":1,"year":2025},{"n":34,"year":2026}],"dependency_candidates":[]},"authors":[{"id":"fbeae6cf-f1aa-414c-8871-7056f33c24cc","orcid":null,"display_name":"Zhang","source":"manual","import_confidence":0.72}]}}