REVIEW 3 major objections 9 minor 23 references
Earnings25 is a 500-hour finance ASR benchmark that pairs full S&P 500 earnings calls with industry-balanced segments and rich metadata so systems can be judged beyond overall word error rate.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-30 11:39 UTC pith:YG6YBM2C
load-bearing objection Solid, usable finance ASR eval resource with real metadata value; the industry-gap story is thinner than the packaging because reference provenance is unspecified. the 3 major comments →
Earnings25: A Comprehensive 500-Hour Speech Benchmark for Finance
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The authors establish Earnings25 as a purpose-built finance-domain ASR benchmark: a 498-hour full-call S&P 500 set plus a 46-hour industry-balanced 290-segment set, both with aligned transcripts and structured metadata for speaker- and industry-aware evaluation beyond aggregate WER, together with standardized Whisper and Parakeet-TDT baselines.
What carries the argument
Earnings25’s dual test design—testset-full (long-form natural distribution) plus testset-segmented (one segment per industry via disproportionate stratified sampling)—with CTC forced alignment, quality filtering, and speaker/industry/call-structure metadata that make stratified scoring possible.
Load-bearing premise
The reference transcripts and the CTC forced-alignment cuts are accurate enough that measured word error rates and industry comparisons reflect true system behavior rather than alignment or transcription artifacts.
What would settle it
Re-transcribe a stratified sample of segments with independent human annotators, re-align with an alternate aligner, and check whether model rankings and the large WER gaps between terminology-dense industries (e.g., biotech/pharma) and the aggregate score reverse or shrink materially.
If this is right
- ASR papers on financial speech can report industry- and role-stratified WER on a shared public set instead of only corpus-level averages.
- Domain adaptation and jargon-handling methods can be stress-tested on long-tail sectors that frequency-weighted corpora under-represent.
- Speaker-aware and diarization-aware pipelines gain a metadata-rich earnings-call evaluation path with executive vs analyst roles.
- Baseline numbers for Whisper and Parakeet-TDT under fixed scoring become a reproducible reference point for later systems.
Where Pith is reading between the lines
- If industry-balanced evaluation becomes standard, model selection for finance products may shift toward systems that win on long-tail sectors rather than on high-frequency utilities and banks.
- The same stratified-segment recipe could be reused for other jargon-heavy meeting domains (legal, medical, earnings in other languages) where aggregate WER also masks failure modes.
- Open release of alignments and speaker tags invites secondary benchmarks for punctuation, numeral formatting, and role-conditioned language models on the same audio.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces Earnings25, an English finance-domain ASR benchmark with two test sets: testset-full (498 h of complete S&P 500 earnings calls from 2025 Q4; ~514 calls, 12 countries, 284 industry categories) and testset-segmented (46 h; 290 industry-balanced 5–10 min segments, one per industry, U.S.-only, sampled from >2,000 U.S. calls in 2025). Construction is documented: two-stage U.S. filter plus disproportionate stratified sampling, CTC forced alignment in NeMo with a 0.5 s max-word-duration cap, greedy block aggregation, boilerplate filtering, one segment per call, all seeded (2025). Aligned transcripts and metadata (speaker roles, industry labels, call structure) are released via Zenodo. Baselines for Whisper (base/medium/large-v2) and Parakeet-TDT-0.6B-v2 are reported under four consistent WER normalizations; the segmented set is slightly harder than the full set, and an illustrative industry cut (Table 5) shows biotech/pharma WER ≈15.3–15.4% vs. a 10.8% aggregate.
Significance. If reference quality is documented, Earnings25 is a useful complement to Earnings-21/22 and SPGISpeech: it adds industry-balanced evaluation, metadata for role- and industry-aware error analysis, and a genuinely reproducible pipeline — fixed seed 2025, seeded stratified sampling, logged decoding configurations, and four WER variants with identical normalization applied to reference and hypothesis. Public release of audio, transcripts, and alignments lowers the barrier to comparable financial-ASR evaluation. The baseline numbers are plausible and internally consistent (Table 4 model ordering matches expectations; Tables 1–3 cross-check arithmetically). The industry-stratified design is a real methodological contribution over frequency-weighted corpora, and the resource is likely to see use.
major comments (3)
- [§4 (Corpus Generation), §3, §5.3, Tables 4–5] Every reported number is measured against reference transcripts, yet §4 describes only alignment and segmentation: the source of the references (vendor service, internal human transcription, ASR-plus-correction), the transcription conventions (verbatim? disfluencies? number formatting?), and any QA or error-rate estimate are never stated. For a benchmark, the ground truth is the product, so this is load-bearing. It also interacts with the paper's own illustrative claim: Table 5 shows biotech (15.4%) and pharma (15.3%) well above the 10.8% aggregate, and if reference errors concentrate in jargon-dense industries (drug names, entity names), part of that gap is reference error rather than model error — the confound would manufacture the finding. This is fixable within scope: (i) state transcript provenance and style conventions; (ii) audit a stratified subsample (e.g., double-transcribe ~5–
- [Table 5, §6.2, Table 1] Statistical support for the industry-level analysis is thin and under-specified. Table 5 rests on 2–3 calls per subsector with no stated selection procedure and no uncertainty quantification; Table 3 shows 9 biotech and 6 pharma calls exist in testset-full, so the basis for the 2–3-call subset must be given (seeded subsample? convenience? worst cases?). Relatedly, testset-segmented contains exactly one ~9.5-minute segment per industry (Table 1), so per-industry WER on the flagship balanced track is a single-sample estimate and cannot support industry rankings. Please (i) state the Table 5 selection procedure, (ii) add bootstrap confidence intervals over calls, and (iii) clarify that the segmented track is intended to yield one balanced aggregate number rather than per-industry inference, or redesign accordingly.
- [§6.2] The attribution that higher WER on testset-segmented 'reflects the effect of industry stratification' is confounded. Relative to testset-full, the segmented set also (a) removes operator boilerplate, i.e., some of the easiest speech, which should raise WER; (b) spans Q1–Q4 2025 rather than Q4 only; (c) consists of mid-call excerpts lacking long-form decoding context; and (d) is U.S.-only, which should lower WER. The comparison conflates at least four factors, so the sentence as written overclaims. Either soften it to a hypothesis or add a small control — e.g., a frequency-weighted U.S.-only sample processed with identical filtering and excerpting — to isolate the stratification effect.
minor comments (9)
- [§4.2.1] Name the specific NeMo CTC checkpoint used for forced alignment, and justify the 0.5 s maximum word-duration cap (could it truncate genuinely long tokens such as slowly pronounced company names?). A small alignment-quality spot check would substantiate the 'reliable timing' claim in §4.3.1.
- [§3, §5.2] State the industry taxonomy and version underlying the 284/290 categories (e.g., BICS level, GICS). The granularity includes categories such as 'Adult Nightclubs,' so the source matters for reproducing the stratified sampling and for cross-study comparability.
- [§5.1, Tables 2–3] The text says testset-full covers 'all S&P 500 companies,' but Tables 2–3 sum to 514 calls. Please reconcile (intra-quarter index changes? dual share classes? multiple calls per company?).
- [§3] Clarify the overlap between the two test sets: can a testset-segmented segment be excerpted from a call that also appears in testset-full? This matters for disjointness claims when both tracks are used together.
- [§6.1] Specify Whisper decoding precisely (temperature, greedy vs. beam and beam size, condition-on-previous-text) or point to the exact configuration file in the released code; 'loaded from a fixed model configuration' is not sufficient for exact replication.
- [Table 4] Consider adding bootstrap confidence intervals over calls/segments. The full-vs-segmented gaps (e.g., large-v2: 0.14039 → 0.15333) are modest relative to plausible call-level variance.
- [§8] Clarify what is actually distributed via the Zenodo DOI — is the audio itself included, or only transcripts/metadata/alignments? The phrase 'terms of the original content providers' leaves unclear what obligations users assume and whether the benchmark is usable out of the box; this affects the practical reproducibility claim.
- [§2.2] The metadata-gap claim should be tempered with respect to SPGISpeech 2.0 [7], which provides speaker-tagged transcription; differentiate Earnings25's speaker-role metadata from SPGISpeech 2.0's speaker tags explicitly.
- [General] Typographical: missing spaces in 'forevaluating' (§1.2) and 'testset-fullconsists' (§3); letter-spaced 'W A V' (§4.2.1, Table 1). Align industry labels across tables — Table 5 uses 'Pharma'/'Biotech' while Table 3 uses 'Large Pharma'/'Biotech'.
Circularity Check
No circularity: benchmark resource with external-model baselines, not a derivation claiming first-principles predictions.
full rationale
Earnings25 is a dataset and evaluation-protocol paper. Its load-bearing content is (i) construction of two held-out test sets via industry-stratified sampling, CTC forced alignment, and segment extraction, and (ii) reporting of WER for external pretrained checkpoints (Whisper base/medium/large-v2 and Parakeet-TDT-0.6B-v2) under fixed decoding and NeMo text normalization. Reported WERs are measurements of those models against reference transcripts; they are not fitted parameters renamed as predictions, nor quantities defined in terms of the quantities they supposedly derive. Citations (wav2vec 2.0, Whisper, SPGISpeech, Earnings-21/22, NeMo, CTC, TDT) are to external systems or prior public benchmarks and are not used as uniqueness theorems or ansatz carriers that force the paper's results. Sampling and stratification choices affect what distribution is measured but do not make the measured WERs true by construction. Concerns about reference-transcript provenance or single-segment-per-industry variance are validity/correctness issues, not circularity. No self-definitional loop, fitted-input-as-prediction, or load-bearing self-citation chain is present. Score 0; steps empty.
Axiom & Free-Parameter Ledger
free parameters (4)
- segment_duration_range =
5–10 minutes
- alignment_max_word_duration =
0.5 seconds
- boundary_padding =
0.2 seconds
- sampling_random_seed =
2025
axioms (4)
- domain assumption Equal per-industry stratified sampling reveals domain-specific ASR failures better than natural sector-frequency evaluation.
- domain assumption CTC forced alignment (NeMo) recovers sufficiently accurate word-level timestamps from reference transcripts for segmentation and speaker-aware analysis.
- domain assumption WER and NeMo-normalized / lowercased / punctuation-stripped variants are adequate primary metrics for comparing finance ASR systems on this audio.
- domain assumption Reference transcripts paired with the audio are accurate enough to serve as evaluation ground truth.
invented entities (1)
-
Earnings25 (testset-full + testset-segmented)
independent evidence
Cite this review
Pith. "Pith review of Earnings25: A Comprehensive 500-Hour Speech Benchmark for Finance." pith.science (2026). https://pith.science/paper/YG6YBM2C
@misc{pith2026260723813,
author = {Pith},
title = {Pith review of: Earnings25: A Comprehensive 500-Hour Speech Benchmark for Finance},
year = {2026},
howpublished = {\url{https://pith.science/paper/YG6YBM2C}},
note = {Machine review of arXiv:2607.23813}
}
read the original abstract
We introduce Earnings25, a finance-domain benchmark for evaluating automatic speech recognition (ASR) on English-language earnings calls under realistic conditions. Earnings25 comprises two complementary test sets: (i) testset-full, 498 hours of full English-language S&P 500 earnings calls from Q4 2025, and (ii) testset-segmented, a 46-hour industry-balanced set of 290 segments sampled from English-language U.S. earnings calls in 2025. The benchmark provides aligned transcripts and structured metadata, including speaker roles, industry labels, and call structure, enabling speaker- and industry-aware evaluation beyond aggregate word error rate (WER). We report reproducible baselines for Whisper and Parakeet-TDT using standardized scoring.
Reference graph
Works this paper leans on
-
[1]
Introduction 1.1. Financial speech understanding is challenging Earnings calls are a critical channel for corporate communica- tion and pose unique challenges for automatic speech recogni- tion (ASR). They combine spontaneous dialogue with scripted remarks, heavy financial jargon, company and product names, frequent numeric expressions, rapid turn-taking,...
-
[2]
Prior Work 2.1. Conversational and meeting speech benchmarks Earnings calls share several conversational properties with meeting and telephone speech, including rapid turn-taking, dis- fluencies, and speaker overlap. Standard benchmarks such asSwitchboard(1992) [3],CallHome(1997) [4], andAMI (2007) [5] capture aspects of these phenomena, but generally lac...
Pith/arXiv arXiv 1992
-
[3]
earnings calls across 2025 (Q1–Q4), comprising 290 seg- ments—one per industry—to ensure balanced domain cover- age
Corpus Definition Earnings25 comprises two complementary test sets designed to support both long-form and segment-level evaluation, as summarized in Table 1.testset-fullconsists of 498 hours of complete earnings-call recordings from S&P 500 companies in 2025 Q4, preserving full conversational context with typical call durations of approximately one hour.t...
2025
-
[4]
This sec- tion details the industry-stratified sampling, alignment, and seg- mentation procedures
Corpus Generation We sample full earnings calls, apply CTC-based forced align- ment to obtain word-level timestamps, and aggregate shorter speech segments into 5–10 minute evaluation units. This sec- tion details the industry-stratified sampling, alignment, and seg- mentation procedures. 4.1. Industry-stratified sampling of the full earnings-call universe...
2025
-
[5]
Corpus Analysis 5.1. Geographic coverage and accent distribution testset-fullincludes earnings calls from all S&P 500 companies and spans a broad geographic footprint, covering English earn- ings calls from 12 countries. While U.S.-domiciled companies dominate the corpus (reflecting index composition), the dataset also includes calls from companies headqu...
2025
-
[6]
Transcription Experiments 6.1. Baseline Models and Evaluation Protocol We benchmark representative ASR systems spanning sequence- to-sequence and transducer architectures: OpenAIWhisper models [2] (base, medium, large-v2) and NVIDIA NeMo’s Parakeet-TDT-0.6B-v2[12]. No external language model is used. Whisper inference:Decoding options are loaded from a fi...
arXiv 2025
-
[7]
Regarding limitations, Earnings25 focuses on English- language earnings calls and primarily reflects speech from U.S.- domiciled companies
Limitations and Conclusion Earnings25 provides value along three dimensions: (1) domain-specific evaluation, offering a challenging benchmark based on S&P 500 earnings calls with broad industry coverage and finance-specific terminology; (2)reproducible baselines, with standardized evaluations for contemporary ASR models, including Whisper and Parakeet-TDT...
-
[8]
Data Access and Licensing Earnings25 is released for research and benchmarking pur- poses. We redistribute the audio recordings, transcripts, metadata, annotations, and evaluation splits through Zenodo: https://doi.org/10.5281/zenodo.18762168. The transcripts, an- notations, metadata, evaluation splits, and alignments are re- leased under the Creative Com...
-
[9]
The tool was not used to generate ex- perimental results, analyses, or conclusions, and no generative AI system is listed as an author
Generative AI Use Disclosure During preparation of this manuscript, the authors used a gener- ative AI assistant only for language editing and polishing (e.g., grammar and phrasing). The tool was not used to generate ex- perimental results, analyses, or conclusions, and no generative AI system is listed as an author
-
[10]
wav2vec 2.0: A framework for self-supervised learn- ing of speech representations,
A. Baevski, Y . Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learn- ing of speech representations,” inAdvances in Neu- ral Information Processing Systems, vol. 33, 2020. [On- line]. Available: https://proceedings.neurips.cc/paper/2020/hash/ 92d1e1eb1cd6f9fba3227870bb6d7f07-Abstract.html
2020
-
[11]
Robust speech recognition via large- scale weak supervision,
A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, “Robust speech recognition via large- scale weak supervision,” 2022. [Online]. Available: https: //arxiv.org/abs/2212.04356
Pith/arXiv arXiv 2022
-
[12]
SWITCHBOARD: Telephone speech corpus for research and development,
J. J. Godfrey, E. Holliman, and J. McDaniel, “SWITCHBOARD: Telephone speech corpus for research and development,” inProc. ICASSP 1992, 1992, pp. 517–520
1992
-
[13]
CALLHOME american english speech,
“CALLHOME american english speech,” Linguistic Data Consortium (LDC), Catalog No. LDC97S42, 1997. [Online]. Available: https://catalog.ldc.upenn.edu/LDC97S42
1997
-
[14]
Unleashing the killer corpus: Experiences in creating the multimodal AMI meeting corpus,
J. Carletta, “Unleashing the killer corpus: Experiences in creating the multimodal AMI meeting corpus,”Language Resources and Evaluation, vol. 41, no. 2, pp. 181–190, 2007
2007
-
[15]
SPGISpeech: 5,000 Hours of Transcribed Financial Audio for Fully Formatted End-to-End Speech Recognition,
P. K. O’Neill, V . Lavrukhin, S. Majumdar, V . Noroozi, Y . Zhang, O. Kuchaiev, J. Balam, Y . Dovzhenko, K. Freyberg, M. D. Shul- man, B. Ginsburg, S. Watanabe, and G. Kucsko, “SPGISpeech: 5,000 Hours of Transcribed Financial Audio for Fully Formatted End-to-End Speech Recognition,” inInterspeech 2021, 2021, pp. 1434–1438
2021
-
[16]
SPGIS- peech 2.0: Transcribed multi-speaker financial audio for speaker- tagged transcription,
R. Grossman, T. Park, K. Dhawan, A. Titus, S. Zhi, Y . Shchadilova, W. Wang, J. Balam, and B. Ginsburg, “SPGIS- peech 2.0: Transcribed multi-speaker financial audio for speaker- tagged transcription,” inInterspeech 2025, 2025, pp. 4048–4052
2025
-
[17]
Earnings-21: A Practical Benchmark for ASR in the Wild,
M. Del Rio, N. Delworth, R. Westerman, M. Huang, N. Bhandari, J. Palakapilly, Q. McNamara, J. Dong, P. ˙Zelasko, and M. Jett ´e, “Earnings-21: A Practical Benchmark for ASR in the Wild,” in Interspeech 2021, 2021, pp. 3465–3469
2021
-
[18]
Earnings-22: A Practical Benchmark for Accents in the Wild,
M. D. Rio, P. Ha, Q. McNamara, C. Miller, and S. Chandra, “Earnings-22: A Practical Benchmark for Accents in the Wild,”
-
[20]
NeMo: a toolkit for building AI applications using neural modules,
O. Kuchaiev, J. Li, B. Ginsburget al., “NeMo: a toolkit for building AI applications using neural modules,” 2019. [Online]. Available: https://arxiv.org/abs/1909.09577
Pith/arXiv arXiv 2019
-
[21]
Con- nectionist temporal classification: Labelling unsegmented se- quence data with recurrent neural networks,
A. Graves, S. Fern ´andez, F. Gomez, and J. Schmidhuber, “Con- nectionist temporal classification: Labelling unsegmented se- quence data with recurrent neural networks,” inProceedings of the 23rd International Conference on Machine Learning, 2006, pp. 369–376
2006
-
[22]
Efficient sequence transduction by jointly predicting tokens and durations,
H. Xu, F. Jia, S. Majumdar, H. Huang, S. Watanabe, and B. Ginsburg, “Efficient sequence transduction by jointly predicting tokens and durations,” 2023. [Online]. Available: https://arxiv.org/abs/2304.06795
Pith/arXiv arXiv 2023
-
[23]
Nemo in- verse text normalization: From development to production,
Y . Zhang, E. Bakhturina, K. Gorman, and B. Ginsburg, “Nemo in- verse text normalization: From development to production,”arXiv preprint arXiv:2104.05055, 2021
Pith/arXiv arXiv 2021
-
[2022]
Available: https://arxiv.org/abs/2203.15591
[Online]. Available: https://arxiv.org/abs/2203.15591
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.