Pith. sign in

REVIEW 3 major objections 9 minor 23 references

Earnings25 is a 500-hour finance ASR benchmark that pairs full S&P 500 earnings calls with industry-balanced segments and rich metadata so systems can be judged beyond overall word error rate.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

Earnings25 releases ~500 hours of 2025 earnings-call audio with aligned transcripts, speaker/industry metadata, and reproducible Whisper and Parakeet-TDT baselines.

T0 review reviewed 2026-07-30 challenge →

load-bearing objection Solid, usable finance ASR eval resource with real metadata value; the industry-gap story is thinner than the packaging because reference provenance is unspecified. the 3 major comments →

arxiv 2607.23813 v1 pith:YG6YBM2C submitted 2026-07-26 cs.CL cs.AI

Earnings25: A Comprehensive 500-Hour Speech Benchmark for Finance

classification cs.CL cs.AI
keywords automatic speech recognitionASR benchmarkfinancial speechearnings callsindustry-aware evaluationlong-form speechword error ratespeaker metadata
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper introduces Earnings25, a public evaluation resource for automatic speech recognition on English earnings calls under realistic conditions. It offers two complementary test sets: nearly 500 hours of complete S&P 500 calls from late 2025, and a 46-hour set of 290 industry-stratified segments drawn from U.S. calls across the year. Aligned transcripts come with speaker roles, industry labels, and call structure, so researchers can measure errors by speaker type and sector rather than a single aggregate score. Reproducible baselines for Whisper and Parakeet-TDT show that industry balance and terminology-dense sectors raise error rates that a frequency-weighted overall metric can hide. A sympathetic reader cares because finance speech mixes jargon, numbers, accents, and rapid Q&A in ways general ASR benchmarks do not capture, and progress needs a shared, metadata-rich yardstick.

Core claim

The authors establish Earnings25 as a purpose-built finance-domain ASR benchmark: a 498-hour full-call S&P 500 set plus a 46-hour industry-balanced 290-segment set, both with aligned transcripts and structured metadata for speaker- and industry-aware evaluation beyond aggregate WER, together with standardized Whisper and Parakeet-TDT baselines.

What carries the argument

Earnings25’s dual test design—testset-full (long-form natural distribution) plus testset-segmented (one segment per industry via disproportionate stratified sampling)—with CTC forced alignment, quality filtering, and speaker/industry/call-structure metadata that make stratified scoring possible.

Load-bearing premise

The reference transcripts and the CTC forced-alignment cuts are accurate enough that measured word error rates and industry comparisons reflect true system behavior rather than alignment or transcription artifacts.

What would settle it

Re-transcribe a stratified sample of segments with independent human annotators, re-align with an alternate aligner, and check whether model rankings and the large WER gaps between terminology-dense industries (e.g., biotech/pharma) and the aggregate score reverse or shrink materially.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • ASR papers on financial speech can report industry- and role-stratified WER on a shared public set instead of only corpus-level averages.
  • Domain adaptation and jargon-handling methods can be stress-tested on long-tail sectors that frequency-weighted corpora under-represent.
  • Speaker-aware and diarization-aware pipelines gain a metadata-rich earnings-call evaluation path with executive vs analyst roles.
  • Baseline numbers for Whisper and Parakeet-TDT under fixed scoring become a reproducible reference point for later systems.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If industry-balanced evaluation becomes standard, model selection for finance products may shift toward systems that win on long-tail sectors rather than on high-frequency utilities and banks.
  • The same stratified-segment recipe could be reused for other jargon-heavy meeting domains (legal, medical, earnings in other languages) where aggregate WER also masks failure modes.
  • Open release of alignments and speaker tags invites secondary benchmarks for punctuation, numeral formatting, and role-conditioned language models on the same audio.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 9 minor

Summary. The paper introduces Earnings25, an English finance-domain ASR benchmark with two test sets: testset-full (498 h of complete S&P 500 earnings calls from 2025 Q4; ~514 calls, 12 countries, 284 industry categories) and testset-segmented (46 h; 290 industry-balanced 5–10 min segments, one per industry, U.S.-only, sampled from >2,000 U.S. calls in 2025). Construction is documented: two-stage U.S. filter plus disproportionate stratified sampling, CTC forced alignment in NeMo with a 0.5 s max-word-duration cap, greedy block aggregation, boilerplate filtering, one segment per call, all seeded (2025). Aligned transcripts and metadata (speaker roles, industry labels, call structure) are released via Zenodo. Baselines for Whisper (base/medium/large-v2) and Parakeet-TDT-0.6B-v2 are reported under four consistent WER normalizations; the segmented set is slightly harder than the full set, and an illustrative industry cut (Table 5) shows biotech/pharma WER ≈15.3–15.4% vs. a 10.8% aggregate.

Significance. If reference quality is documented, Earnings25 is a useful complement to Earnings-21/22 and SPGISpeech: it adds industry-balanced evaluation, metadata for role- and industry-aware error analysis, and a genuinely reproducible pipeline — fixed seed 2025, seeded stratified sampling, logged decoding configurations, and four WER variants with identical normalization applied to reference and hypothesis. Public release of audio, transcripts, and alignments lowers the barrier to comparable financial-ASR evaluation. The baseline numbers are plausible and internally consistent (Table 4 model ordering matches expectations; Tables 1–3 cross-check arithmetically). The industry-stratified design is a real methodological contribution over frequency-weighted corpora, and the resource is likely to see use.

major comments (3)
  1. [§4 (Corpus Generation), §3, §5.3, Tables 4–5] Every reported number is measured against reference transcripts, yet §4 describes only alignment and segmentation: the source of the references (vendor service, internal human transcription, ASR-plus-correction), the transcription conventions (verbatim? disfluencies? number formatting?), and any QA or error-rate estimate are never stated. For a benchmark, the ground truth is the product, so this is load-bearing. It also interacts with the paper's own illustrative claim: Table 5 shows biotech (15.4%) and pharma (15.3%) well above the 10.8% aggregate, and if reference errors concentrate in jargon-dense industries (drug names, entity names), part of that gap is reference error rather than model error — the confound would manufacture the finding. This is fixable within scope: (i) state transcript provenance and style conventions; (ii) audit a stratified subsample (e.g., double-transcribe ~5–
  2. [Table 5, §6.2, Table 1] Statistical support for the industry-level analysis is thin and under-specified. Table 5 rests on 2–3 calls per subsector with no stated selection procedure and no uncertainty quantification; Table 3 shows 9 biotech and 6 pharma calls exist in testset-full, so the basis for the 2–3-call subset must be given (seeded subsample? convenience? worst cases?). Relatedly, testset-segmented contains exactly one ~9.5-minute segment per industry (Table 1), so per-industry WER on the flagship balanced track is a single-sample estimate and cannot support industry rankings. Please (i) state the Table 5 selection procedure, (ii) add bootstrap confidence intervals over calls, and (iii) clarify that the segmented track is intended to yield one balanced aggregate number rather than per-industry inference, or redesign accordingly.
  3. [§6.2] The attribution that higher WER on testset-segmented 'reflects the effect of industry stratification' is confounded. Relative to testset-full, the segmented set also (a) removes operator boilerplate, i.e., some of the easiest speech, which should raise WER; (b) spans Q1–Q4 2025 rather than Q4 only; (c) consists of mid-call excerpts lacking long-form decoding context; and (d) is U.S.-only, which should lower WER. The comparison conflates at least four factors, so the sentence as written overclaims. Either soften it to a hypothesis or add a small control — e.g., a frequency-weighted U.S.-only sample processed with identical filtering and excerpting — to isolate the stratification effect.
minor comments (9)
  1. [§4.2.1] Name the specific NeMo CTC checkpoint used for forced alignment, and justify the 0.5 s maximum word-duration cap (could it truncate genuinely long tokens such as slowly pronounced company names?). A small alignment-quality spot check would substantiate the 'reliable timing' claim in §4.3.1.
  2. [§3, §5.2] State the industry taxonomy and version underlying the 284/290 categories (e.g., BICS level, GICS). The granularity includes categories such as 'Adult Nightclubs,' so the source matters for reproducing the stratified sampling and for cross-study comparability.
  3. [§5.1, Tables 2–3] The text says testset-full covers 'all S&P 500 companies,' but Tables 2–3 sum to 514 calls. Please reconcile (intra-quarter index changes? dual share classes? multiple calls per company?).
  4. [§3] Clarify the overlap between the two test sets: can a testset-segmented segment be excerpted from a call that also appears in testset-full? This matters for disjointness claims when both tracks are used together.
  5. [§6.1] Specify Whisper decoding precisely (temperature, greedy vs. beam and beam size, condition-on-previous-text) or point to the exact configuration file in the released code; 'loaded from a fixed model configuration' is not sufficient for exact replication.
  6. [Table 4] Consider adding bootstrap confidence intervals over calls/segments. The full-vs-segmented gaps (e.g., large-v2: 0.14039 → 0.15333) are modest relative to plausible call-level variance.
  7. [§8] Clarify what is actually distributed via the Zenodo DOI — is the audio itself included, or only transcripts/metadata/alignments? The phrase 'terms of the original content providers' leaves unclear what obligations users assume and whether the benchmark is usable out of the box; this affects the practical reproducibility claim.
  8. [§2.2] The metadata-gap claim should be tempered with respect to SPGISpeech 2.0 [7], which provides speaker-tagged transcription; differentiate Earnings25's speaker-role metadata from SPGISpeech 2.0's speaker tags explicitly.
  9. [General] Typographical: missing spaces in 'forevaluating' (§1.2) and 'testset-fullconsists' (§3); letter-spaced 'W A V' (§4.2.1, Table 1). Align industry labels across tables — Table 5 uses 'Pharma'/'Biotech' while Table 3 uses 'Large Pharma'/'Biotech'.

Circularity Check

0 steps flagged

No circularity: benchmark resource with external-model baselines, not a derivation claiming first-principles predictions.

full rationale

Earnings25 is a dataset and evaluation-protocol paper. Its load-bearing content is (i) construction of two held-out test sets via industry-stratified sampling, CTC forced alignment, and segment extraction, and (ii) reporting of WER for external pretrained checkpoints (Whisper base/medium/large-v2 and Parakeet-TDT-0.6B-v2) under fixed decoding and NeMo text normalization. Reported WERs are measurements of those models against reference transcripts; they are not fitted parameters renamed as predictions, nor quantities defined in terms of the quantities they supposedly derive. Citations (wav2vec 2.0, Whisper, SPGISpeech, Earnings-21/22, NeMo, CTC, TDT) are to external systems or prior public benchmarks and are not used as uniqueness theorems or ansatz carriers that force the paper's results. Sampling and stratification choices affect what distribution is measured but do not make the measured WERs true by construction. Concerns about reference-transcript provenance or single-segment-per-industry variance are validity/correctness issues, not circularity. No self-definitional loop, fitted-input-as-prediction, or load-bearing self-citation chain is present. Score 0; steps empty.

Axiom & Free-Parameter Ledger

4 free parameters · 4 axioms · 1 invented entities

As a benchmark paper, load-bearing content is design choices and tooling assumptions, not a mathematical derivation. Claims rest on stratified sampling design, CTC alignment fidelity, reference-transcript quality, and standard WER-style scoring—not on fitted physical constants or new particles.

free parameters (4)
  • segment_duration_range = 5–10 minutes
    Blocks are constrained to a hand-chosen 5–10 minute window that defines testset-segmented units and affects conversational context available to ASR.
  • alignment_max_word_duration = 0.5 seconds
    Hard cap used to mitigate trailing-silence/low-confidence alignment errors; directly shapes word boundaries and extracted segments.
  • boundary_padding = 0.2 seconds
    Fixed pad added before audio extraction to absorb alignment jitter; chosen by authors, not derived.
  • sampling_random_seed = 2025
    Seed fixes industry draws and per-call segment selection; reproducibility knob rather than data fit, but still an author-chosen control.
axioms (4)
  • domain assumption Equal per-industry stratified sampling reveals domain-specific ASR failures better than natural sector-frequency evaluation.
    Core design claim in §1.2, §4.1, and §6.2; motivates testset-segmented and interpretation that higher segmented WER is a feature.
  • domain assumption CTC forced alignment (NeMo) recovers sufficiently accurate word-level timestamps from reference transcripts for segmentation and speaker-aware analysis.
    Pipeline in §4.2 is load-bearing for all segment cuts and timing metadata; errors would bias the benchmark units.
  • domain assumption WER and NeMo-normalized / lowercased / punctuation-stripped variants are adequate primary metrics for comparing finance ASR systems on this audio.
    §6.1 scoring protocol; standard in ASR but still an evaluative assumption about what 'better' means for earnings calls.
  • domain assumption Reference transcripts paired with the audio are accurate enough to serve as evaluation ground truth.
    Implicit throughout corpus definition and experiments; provenance/QA process is not fully detailed in the manuscript.
invented entities (1)
  • Earnings25 (testset-full + testset-segmented) independent evidence
    purpose: Provide a public, metadata-rich finance ASR evaluation suite with full-call and industry-balanced segment tracks.
    The named benchmark and its two splits are the paper’s primary constructed artifact, not a pre-existing standard object.

reviewed 2026-07-30 · how reviews work

0 comments
Cite this review

Pith. "Pith review of Earnings25: A Comprehensive 500-Hour Speech Benchmark for Finance." pith.science (2026). https://pith.science/paper/YG6YBM2C

@misc{pith2026260723813,
  author       = {Pith},
  title        = {Pith review of: Earnings25: A Comprehensive 500-Hour Speech Benchmark for Finance},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/YG6YBM2C}},
  note         = {Machine review of arXiv:2607.23813}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

We introduce Earnings25, a finance-domain benchmark for evaluating automatic speech recognition (ASR) on English-language earnings calls under realistic conditions. Earnings25 comprises two complementary test sets: (i) testset-full, 498 hours of full English-language S&P 500 earnings calls from Q4 2025, and (ii) testset-segmented, a 46-hour industry-balanced set of 290 segments sampled from English-language U.S. earnings calls in 2025. The benchmark provides aligned transcripts and structured metadata, including speaker roles, industry labels, and call structure, enabling speaker- and industry-aware evaluation beyond aggregate word error rate (WER). We report reproducible baselines for Whisper and Parakeet-TDT using standardized scoring.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

23 extracted references · 1 canonical work pages

  1. [1]

    Introduction 1.1. Financial speech understanding is challenging Earnings calls are a critical channel for corporate communica- tion and pose unique challenges for automatic speech recogni- tion (ASR). They combine spontaneous dialogue with scripted remarks, heavy financial jargon, company and product names, frequent numeric expressions, rapid turn-taking,...

  2. [2]

    Prior Work 2.1. Conversational and meeting speech benchmarks Earnings calls share several conversational properties with meeting and telephone speech, including rapid turn-taking, dis- fluencies, and speaker overlap. Standard benchmarks such asSwitchboard(1992) [3],CallHome(1997) [4], andAMI (2007) [5] capture aspects of these phenomena, but generally lac...

  3. [3]

    earnings calls across 2025 (Q1–Q4), comprising 290 seg- ments—one per industry—to ensure balanced domain cover- age

    Corpus Definition Earnings25 comprises two complementary test sets designed to support both long-form and segment-level evaluation, as summarized in Table 1.testset-fullconsists of 498 hours of complete earnings-call recordings from S&P 500 companies in 2025 Q4, preserving full conversational context with typical call durations of approximately one hour.t...

  4. [4]

    This sec- tion details the industry-stratified sampling, alignment, and seg- mentation procedures

    Corpus Generation We sample full earnings calls, apply CTC-based forced align- ment to obtain word-level timestamps, and aggregate shorter speech segments into 5–10 minute evaluation units. This sec- tion details the industry-stratified sampling, alignment, and seg- mentation procedures. 4.1. Industry-stratified sampling of the full earnings-call universe...

  5. [5]

    Corpus Analysis 5.1. Geographic coverage and accent distribution testset-fullincludes earnings calls from all S&P 500 companies and spans a broad geographic footprint, covering English earn- ings calls from 12 countries. While U.S.-domiciled companies dominate the corpus (reflecting index composition), the dataset also includes calls from companies headqu...

  6. [6]

    Transcription Experiments 6.1. Baseline Models and Evaluation Protocol We benchmark representative ASR systems spanning sequence- to-sequence and transducer architectures: OpenAIWhisper models [2] (base, medium, large-v2) and NVIDIA NeMo’s Parakeet-TDT-0.6B-v2[12]. No external language model is used. Whisper inference:Decoding options are loaded from a fi...

  7. [7]

    Regarding limitations, Earnings25 focuses on English- language earnings calls and primarily reflects speech from U.S.- domiciled companies

    Limitations and Conclusion Earnings25 provides value along three dimensions: (1) domain-specific evaluation, offering a challenging benchmark based on S&P 500 earnings calls with broad industry coverage and finance-specific terminology; (2)reproducible baselines, with standardized evaluations for contemporary ASR models, including Whisper and Parakeet-TDT...

  8. [8]

    We redistribute the audio recordings, transcripts, metadata, annotations, and evaluation splits through Zenodo: https://doi.org/10.5281/zenodo.18762168

    Data Access and Licensing Earnings25 is released for research and benchmarking pur- poses. We redistribute the audio recordings, transcripts, metadata, annotations, and evaluation splits through Zenodo: https://doi.org/10.5281/zenodo.18762168. The transcripts, an- notations, metadata, evaluation splits, and alignments are re- leased under the Creative Com...

  9. [9]

    The tool was not used to generate ex- perimental results, analyses, or conclusions, and no generative AI system is listed as an author

    Generative AI Use Disclosure During preparation of this manuscript, the authors used a gener- ative AI assistant only for language editing and polishing (e.g., grammar and phrasing). The tool was not used to generate ex- perimental results, analyses, or conclusions, and no generative AI system is listed as an author

  10. [10]

    wav2vec 2.0: A framework for self-supervised learn- ing of speech representations,

    A. Baevski, Y . Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learn- ing of speech representations,” inAdvances in Neu- ral Information Processing Systems, vol. 33, 2020. [On- line]. Available: https://proceedings.neurips.cc/paper/2020/hash/ 92d1e1eb1cd6f9fba3227870bb6d7f07-Abstract.html

  11. [11]

    Robust speech recognition via large- scale weak supervision,

    A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, “Robust speech recognition via large- scale weak supervision,” 2022. [Online]. Available: https: //arxiv.org/abs/2212.04356

  12. [12]

    SWITCHBOARD: Telephone speech corpus for research and development,

    J. J. Godfrey, E. Holliman, and J. McDaniel, “SWITCHBOARD: Telephone speech corpus for research and development,” inProc. ICASSP 1992, 1992, pp. 517–520

  13. [13]

    CALLHOME american english speech,

    “CALLHOME american english speech,” Linguistic Data Consortium (LDC), Catalog No. LDC97S42, 1997. [Online]. Available: https://catalog.ldc.upenn.edu/LDC97S42

  14. [14]

    Unleashing the killer corpus: Experiences in creating the multimodal AMI meeting corpus,

    J. Carletta, “Unleashing the killer corpus: Experiences in creating the multimodal AMI meeting corpus,”Language Resources and Evaluation, vol. 41, no. 2, pp. 181–190, 2007

  15. [15]

    SPGISpeech: 5,000 Hours of Transcribed Financial Audio for Fully Formatted End-to-End Speech Recognition,

    P. K. O’Neill, V . Lavrukhin, S. Majumdar, V . Noroozi, Y . Zhang, O. Kuchaiev, J. Balam, Y . Dovzhenko, K. Freyberg, M. D. Shul- man, B. Ginsburg, S. Watanabe, and G. Kucsko, “SPGISpeech: 5,000 Hours of Transcribed Financial Audio for Fully Formatted End-to-End Speech Recognition,” inInterspeech 2021, 2021, pp. 1434–1438

  16. [16]

    SPGIS- peech 2.0: Transcribed multi-speaker financial audio for speaker- tagged transcription,

    R. Grossman, T. Park, K. Dhawan, A. Titus, S. Zhi, Y . Shchadilova, W. Wang, J. Balam, and B. Ginsburg, “SPGIS- peech 2.0: Transcribed multi-speaker financial audio for speaker- tagged transcription,” inInterspeech 2025, 2025, pp. 4048–4052

  17. [17]

    Earnings-21: A Practical Benchmark for ASR in the Wild,

    M. Del Rio, N. Delworth, R. Westerman, M. Huang, N. Bhandari, J. Palakapilly, Q. McNamara, J. Dong, P. ˙Zelasko, and M. Jett ´e, “Earnings-21: A Practical Benchmark for ASR in the Wild,” in Interspeech 2021, 2021, pp. 3465–3469

  18. [18]

    Earnings-22: A Practical Benchmark for Accents in the Wild,

    M. D. Rio, P. Ha, Q. McNamara, C. Miller, and S. Chandra, “Earnings-22: A Practical Benchmark for Accents in the Wild,”

  19. [20]

    NeMo: a toolkit for building AI applications using neural modules,

    O. Kuchaiev, J. Li, B. Ginsburget al., “NeMo: a toolkit for building AI applications using neural modules,” 2019. [Online]. Available: https://arxiv.org/abs/1909.09577

  20. [21]

    Con- nectionist temporal classification: Labelling unsegmented se- quence data with recurrent neural networks,

    A. Graves, S. Fern ´andez, F. Gomez, and J. Schmidhuber, “Con- nectionist temporal classification: Labelling unsegmented se- quence data with recurrent neural networks,” inProceedings of the 23rd International Conference on Machine Learning, 2006, pp. 369–376

  21. [22]

    Efficient sequence transduction by jointly predicting tokens and durations,

    H. Xu, F. Jia, S. Majumdar, H. Huang, S. Watanabe, and B. Ginsburg, “Efficient sequence transduction by jointly predicting tokens and durations,” 2023. [Online]. Available: https://arxiv.org/abs/2304.06795

  22. [23]

    Nemo in- verse text normalization: From development to production,

    Y . Zhang, E. Bakhturina, K. Gorman, and B. Ginsburg, “Nemo in- verse text normalization: From development to production,”arXiv preprint arXiv:2104.05055, 2021

  23. [2022]

    Available: https://arxiv.org/abs/2203.15591

    [Online]. Available: https://arxiv.org/abs/2203.15591

This paper was first reviewed by grok-4.5 on July 30, 2026.