Pith. sign in

REVIEW 7 cited by

Earnings-22: A Practical Benchmark for Accents in the Wild

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2203.15591 v1 pith:VVFT6YGX submitted 2022-03-29 cs.CL

classification cs.CL
keywords speechbenchmarkearnings-22performanceacademicaccentedaccentscommercial
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Modern automatic speech recognition (ASR) systems have achieved superhuman Word Error Rate (WER) on many common corpora despite lacking adequate performance on speech in the wild. Beyond that, there is a lack of real-world, accented corpora to properly benchmark academic and commercial models. To ensure this type of speech is represented in ASR benchmarking, we present Earnings-22, a 125 file, 119 hour corpus of English-language earnings calls gathered from global companies. We run a comparison across 4 commercial models showing the variation in performance when taking country of origin into consideration. Looking at hypothesis transcriptions, we explore errors common to all ASR systems tested. By examining Individual Word Error Rate (IWER), we find that key speech features impact model performance more for certain accents than others. Earnings-22 provides a free-to-use benchmark of real-world, accented audio to bridge academic and industrial research.

Discussion (0). Sign in to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Synonymous Variational Perspective on the Rate-Distortion-Perception Tradeoff

    cs.IT 2026-04 unverdicted novelty 6.0 of 10

    Synset-based reconstruction and synonymous variational inference are claimed to derive the distributional divergence in RDP and unify it with classical rate-distortion theory.

  2. Whisper-CD: Accurate Long-Form Speech Recognition using Multi-Negative Contrastive Decoding

    cs.SD 2026-03 conditional novelty 6.0 of 10

    Multi-negative contrastive decoding with noise, silence and temporal-shift negatives cuts Whisper long-form WER by up to 24.3 pp while remaining faster than beam search.

  3. Edge-ASR: Towards Low-Bit Quantization of Automatic Speech Recognition Models

    cs.SD 2025-07 conditional novelty 6.0 of 10

    A benchmark of eight post-training quantization methods on Whisper and Moonshine edge speech models across seven datasets, finding 8-bit is safe and 3-bit weights are viable for larger models with advanced methods like SpQR.

  4. FastSLM: Hierarchical Temporal Abstraction for Efficient Long-Form Speech Adaptation

    eess.AS 2026-01 conditional novelty 5.0 of 10

    A hierarchical Q-Former compresses speech to about 1.67 tokens/sec, enabling hour-long audio processing with near-linear memory scaling and competitive benchmark scores.

  5. Adapting Whisper for Streaming Speech Recognition via Two-Pass Decoding

    cs.SD 2025-06 conditional novelty 5.0 of 10

    A U2-style two-pass adaptation with an 8,000-token CTC branch turns Whisper into a streaming ASR model that runs on CPUs in real time.

  6. Earnings25: A Comprehensive 500-Hour Speech Benchmark for Finance

    cs.CL 2026-07 accept novelty 4.0 of 10

    Earnings25 releases ~500 hours of 2025 earnings-call audio with aligned transcripts, speaker/industry metadata, and reproducible Whisper and Parakeet-TDT baselines.

  7. ASKD-Whisper: Adaptive Self-knowledge Distillation for Efficient and Low-Latency Automatic Speech Recognition

    cs.CL 2026-01 reject novelty 4.0 of 10

    A distillation method that decays teacher loss then applies self-distillation yields a Whisper-derived ASR model with 5x lower latency and slightly better average WER only on in-domain noisy datasets.

Pith tools