{"id":"ce094771-522f-4cda-8820-74f6ad74c141","arxiv_id":"2506.11086","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"State-of-the-art text-to-speech models are often unintelligible when reading mathematical expressions aloud, with accuracy varying sharply by expression category and model.","lead":"This study played audio of mathematical expressions, generated by five commercial text-to-speech systems, and asked 49 listeners to write down what they heard. Most systems were frequently unintelligible for categories like matrices and summations, and expert human recordings scored much higher.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The expert-rendition gap is confounded: TTS and reference audio were produced from different verbalization text, so the claimed gap cannot be attributed to TTS prosody without a same-text control.","rationale":"The L1 listening test is well designed for comparing TTS models: each expression is given identical MXText across all five models, with 120 MX and three ratings per audio, and the manual pronunciation-correctness check plus ANOVA with correctly pronounced MX provides partial support that LLM choice is not the main driver. The soft spot is RQ2 and the abstract's 'significantly worse than expert rendition' claim: the reference audio and the TTS audio are generated from different text, so the gap conflates wording choice with synthesis/prosody. This is the single most load-bearing concern because it directly affects the headline attribution to TTS models. The paper's descriptive intelligibility results remain useful, and the reader's CONDITIONAL verdict already captures the need for controls; my recommendation is to keep the verdict unchanged while explicitly adding a same-text reference condition to the release/analysis requirements.","tokens_in":9499,"tokens_out":6733,"duration_ms":82038,"concrete_test":"For a subset of the 35 L2 MX (or all 35), ask the same four experts to read the exact LLM-generated MXText aloud, without paraphrasing, and record this as same-text reference audio. Rerun the L2 protocol with seven audio items per MX: the five TTS AudioMX, the original expert reference, and the same-text reference, rated under the same MUSHRA-style protocol by a comparable listener panel. Recompute Table 3 as mean differences relative to the same-text reference. If the mean gap for Advanced Algebraic, Calculus, Summation, and Sq. Roots and Trig falls below 10 points on the 0-100 scale, the reported gap is primarily an LLM-verbalization artifact; if the gap remains above 20 points, the conclusion that TTS synthesis lacks MX prosody survives.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract claim that 'for most categories, performance of TTS models is significantly worse than that of expert rendition' rests on the L2 comparison in Section 3.3 and Table 3. In L2, each TTS AudioMX is synthesized from LLM-generated MXText (Section 2.2.2), while the hidden reference RAudioMX is recorded by experts using their own chosen wording. Section 5.2 explicitly states that 'Most experts used descriptors like whole divided by' and introduced pauses that are absent from the TTS renderings. The reported gap is therefore a system-level difference between (LLM front-end + TTS) and (expert verbalizer + human speech), not a TTS-model deficit. The L1 results are less affected because all five TTS models receive identical MXText, so cross-model differences in Table 2 are attributable to the TTS; however, absolute intelligibility levels and the RQ2 gap still depend on text quality. The post-hoc check in Section 4 that filtering to correctly pronounced MXText does not change the inference addresses the LLM front-end only partially: it does not compare TTS audio with expert audio generated from the same text. A controlled same-text reference condition is needed before the conclusion that 'public API or open source general purpose TTS models do not include the complexities of prosody particular to MX' (Section 4) can be drawn.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper evaluates the intelligibility and perceived quality of five text-to-speech (TTS) systems for mathematical expressions (MX). Because TTS models cannot directly process LaTeX, the authors use LLMs (QWEN2.5-7b and GPT4) to generate English verbalizations (MXText). Two listening tests are conducted: L1 measures transcription correctness (CoC, LCER, TeXBLEU) and MOS for 120 MX across 8 categories; L2 compares TTS audio against human expert reference recordings using a MUSHRA-style rating. The paper reports that intelligibility varies significantly by TTS model and MX category, and that for most categories there is a large gap between TTS and expert renditions, concluding that current TTS models lack MX-appropriate prosody.","tokens_in":9664,"tokens_out":7378,"duration_ms":78809,"significance":"The study addresses a real gap: there is little prior work on perceptual evaluation of TTS for mathematical content. The L1 results, based on 1800 listener ratings and three transcription metrics, provide useful evidence that intelligibility is model- and category-dependent. The paper also introduces LCER/TeXBLEU for MX transcription evaluation and shows that ASR-based cascade metrics are poorly correlated with human transcription. However, the RQ2 conclusion about the expert-rendition gap is not yet supported because the TTS and reference audio are not produced from the same text; a same-text control is needed. The post-hoc pronunciation correction check in Section 4 addresses the LLM front-end only for L1, not for L2.","major_comments":[{"comment":"The L2 comparison conflates verbalization text with TTS performance. The TTS AudioMX is synthesized from LLM-generated MXText (Section 2.2.2), while the hidden reference RAudioMX is recorded by experts using their own wording (explicitly stated in Section 5.2). Consequently, the mean differences in Table 3 reflect differences in both the spoken text and the speech rendering, and the conclusion in Section 4 that 'public API or open source general purpose TTS models do not include the complexities of prosody particular to MX' is not warranted by the current design. A same-text control—either experts reading the MXText or TTS models synthesizing the expert verbalizations—is required to attribute the gap to TTS. Without it, the RQ2 gap should be described as a property of the LLM+TTS cascade.","section":"Section 3.3 / Table 3"},{"comment":"The abstract states that 'for most categories, performance of TTS models is significantly worse than that of expert rendition,' but no statistical test is reported for the L2 differences in Table 3. The ANOVA mentioned in Section 4 is on L1 LCER scores and does not support the L2 claim. Please add appropriate significance tests (e.g., paired tests across listeners or MX) on the per-MX reference-to-TTS differences.","section":"Abstract and Section 4"},{"comment":"The manual filtering to correctly pronounced MXText does not address the L2 confound: it only shows that, in L1, the LLM factor becomes insignificant when incorrect pronunciations are removed. The conclusion that 'when the pronunciation is correct, LCER and MOS are statistically equivalent across both LLMs' is not a substitute for a same-text comparison with expert audio, and it does not support the abstract's claim about the gap with expert rendition.","section":"Section 4 (post-hoc check)"}],"minor_comments":[{"comment":"The text says 'We sample 35 MX from 5 categories (excluding matrices)' but Table 3 reports 7 categories; please clarify the number of categories actually used in L2.","section":"Section 3.3"},{"comment":"Please report inter-annotator agreement for the manual CoC evaluation; the current description says four evaluators mark each transcript but does not state how disagreements were resolved.","section":"Section 3.2"},{"comment":"Please specify how the 120 MX were assigned to QWEN vs. GPT4 (e.g., random, category-balanced); the ANOVA on the 'LLM' factor assumes the factor is crossed with category.","section":"Section 2.2.2"},{"comment":"Typo 'ANoV A' should be 'ANOVA'; please also report degrees of freedom and effect sizes for the significant factors.","section":"Section 4"},{"comment":"The sentence 'we consolidate MX from different folders of HME to include categories such as matrices, de-duplicate and obtain a dataset of 3141 MX' is awkward; consider splitting into two sentences.","section":"Section 2.1"},{"comment":"Minor language issues: 'a few of who did at most 3 batches' should be 'a few of whom'; 'have age range' should be 'are in the age range'.","section":"Section 3.1"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses a worthwhile question and the L1 experiment is a solid empirical contribution. However, the RQ2 claim about the gap with expert renditions is over-stated given the confound between TTS and reference text. The authors should either add a same-text control or carefully weaken the abstract and conclusions. The L2 significance statements also need proper statistical backing."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is the first paper I know of that actually puts TTS audio of math expressions in front of listeners and asks them to transcribe, and that alone makes it worth knowing. The L1 experiment—five TTS models, same LLM-generated verbalization for each expression, 49 listeners, three metrics—is a genuinely useful category-by-category map. The result that intelligibility varies by category and model, with numerics near ceiling and matrices/summation near floor, is credible and probably reproducible. The qualitative error analysis (pause placement changing 'a2-b2 over a' into 'a2 - b2/a', 'sin' mispronounced as [sIn]) is the kind of detail that will help people building math-aware TTS. The Whisper ASR control showing high TTS-to-ASR fidelity but low correlation with human transcription is a nice sanity check; it shows the problem is in the human-perception/prosody domain, not just garbled audio.\n\nThe soft spot is exactly where the stress-test note lands, and I think it lands. The L2 'gap with expert rendition' compares TTS audio synthesized from LLM-generated MXText with expert recordings made from the expert's own wording. Section 5.2 admits experts used 'whole divided by' and deliberate pauses. So Table 3 quantifies (LLM verbalizer + TTS) versus (expert verbalizer + human speech). The conclusion that 'general purpose TTS models do not include the complexities of prosody particular to MX' overshoots. You would need expert audio spoken from the same MXText, or TTS audio from expert verbalizations, to isolate prosody. The L1 cross-model comparisons are not affected, because all five TTS models read identical text, so Table 2 differences are attributable to TTS. Good.\n\nSmaller issues: no confidence intervals or inter-annotator agreement on CoC; the 'significantly worse' in the abstract has no significance test attached in the paper; no data/code release; and the pronunciation-error rate (12.5%) is acknowledged and post-hoc filtered, but the absolute intelligibility numbers still inherit whatever LLM quirks remain.\n\nBottom line: the core observation holds—current TTS, even with a good LLM front-end, is often not intelligible for math, and category matters. The paper is not circular and does not overclaim its novelty. If I were refereeing it, I'd ask for a matched-text expert condition (or a clear qualification of what the L2 gap means), statistical detail, and release of audio/text transcriptions; with those, it would be a solid contribution. Worth a serious referee, and I'd cite it for the L1 results.","headline":"First real listening-test data on TTS reading math, with a solid cross-model comparison in L1; the expert-rendition gap claim is weaker than the abstract suggests because it compares different text, not just different speech.","tokens_in":10281,"tokens_out":2258,"would_cite":true,"duration_ms":27471,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Listening tests of five state-of-the-art text-to-speech models show that spoken mathematical expressions are often transcribed incorrectly, with accuracy varying by model and expression category and trailing human expert recordings in…","keywords":["text-to-speech","mathematical expressions","intelligibility","listening test","LaTeX","large language models","mean opinion score","prosody"],"falsifier":"Give the same five TTS models hand-crafted expert English pronunciations (the kind the human experts used) for the same 120 expressions, rerun the transcription test, and compare counts-of-correct. If transcription accuracy jumps to near the expert-audio level, the intelligibility gap is caused by the LLM-generated text, not by TTS synthesis; if it stays low, the synthesis itself is the bottleneck.","tokens_in":9252,"feed_emoji":"🎧","tokens_out":5567,"duration_ms":59356,"temperature":0.7,"pith_summary":"This paper asks whether state-of-the-art text-to-speech (TTS) systems can read mathematical expressions aloud so that listeners can write them down correctly. The authors take LaTeX expressions, use two large language models to turn them into English pronunciations, feed those into five TTS models, and have listeners transcribe the audio. Across 120 expressions in eight categories, transcription is often far from perfect, and for most categories it is significantly worse than audio recorded by human experts. The results imply that current TTS models lack the prosody and phrasing that mathematics needs, and that evaluation of TTS for math must use human listeners, not automatic speech recognition.","feed_headline":"Five TTS models stumble reading math aloud","feed_subtitle":"Listening tests show transcription accuracy below expert recordings for most expression categories.","key_machinery":"The evaluation cascade: LaTeX math is converted to a spoken-English string (MXText) by an LLM, then synthesized by one of five TTS models into audio (AudioMX), which listeners transcribe back to LaTeX. The intelligibility measurement rests on three complementary metrics: expert-judged count-of-correct, a normalized Levenshtein-based LaTeX character error rate, and TeXBLEU, a structural-grammar similarity score. A MUSHRA-style hidden-reference test quantifies the gap against expert recordings.","core_discovery":"The paper's central claim is that TTS output for mathematical expressions is not necessarily intelligible, and that the intelligibility gap varies both with the TTS model and with the type of expression. Using three transcription metrics—count-of-correct, LaTeX character error rate, and TeXBLEU—the authors find that no model is consistently best; matrices and summations are hardest, while numerics are nearly perfect. Listeners' opinion scores overstate their actual comprehension, and the choice of LLM that produces the pronunciation has little effect once the pronunciation is correct. In a hidden-reference comparison, expert human recordings score higher than every TTS model for every expression, with large gaps for calculus, roots, and summation.","pith_inferences":["Because expert reference audio was recorded without the LLM in the loop, the reported expert-versus-TTS gap is a property of the LLM+TTS cascade; a TTS model fed an expert's spoken text might narrow the gap substantially.","The roughly 12.5% judged-incorrect LLM pronunciations mean some intelligibility failures are upstream of the TTS; isolating TTS prosody from text errors would require controlling the text input.","The category-level difficulty ranking (numerics easiest, matrices and summations hardest) could serve as a stress-test battery for future math-aware TTS models.","The low correlation between ASR metrics and listener transcription suggests a blind spot: systems optimized for ASR word error rates may not be optimizing for human comprehension of structured notation."],"forward_implications":["No single public TTS model can be trusted for math audio; systems must be chosen or tuned by expression category.","User perception of understanding is not a reliable substitute for transcription: MOS scores ran ahead of actual correctness.","Automatic speech recognition metrics on a TTS-ASR pipeline do not predict human intelligibility, so ASR-based evaluation is insufficient.","Accessible math content, such as audio textbooks for vision-impaired readers, needs purpose-built prosody rather than stock voices.","Future work should fine-tune or train TTS models with math-specific prosody, with the category-level baselines here as a benchmark."],"supporting_citations":[{"why":"Supplies the handwritten mathematical expression dataset from which the 120 test expressions and eight categories were sampled.","marker":"[5]"},{"why":"Provides the MathBridge corpus and the argument that word-error rate alone is insufficient for evaluating spoken math.","marker":"[4]"},{"why":"Prior text-to-speech-for-math system that assumed TTS is error-free, which this paper's evaluation stands against.","marker":"[15]"},{"why":"Prior math-to-speech vision-based system, also assuming error-free TTS, used as contrast for the evaluation gap.","marker":"[16]"},{"why":"Defines TeXBLEU, one of the three metrics used to score listener transcriptions.","marker":"[25]"},{"why":"Supplies the Levenshtein distance underlying the LaTeX character error rate metric.","marker":"[26]"},{"why":"Specifies the MUSHRA hidden-reference methodology adopted for the expert-comparison listening test.","marker":"[27]"},{"why":"The automatic speech recognition model used to show that ASR-based metrics do not track human intelligibility.","marker":"[29]"}],"fun_headline_variants":["TTS models still can't read math aloud clearly","Math TTS falls short of human experts","Matrices and summations stump TTS models","Text-to-speech fails at math intelligibility"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper assumes that the LLM-generated English pronunciation fairly represents the mathematical content, so that transcription failures can be blamed on the TTS model rather than on the wording fed to it.","fun_headline_variants_meta":{"raw":{"variants":["TTS models still can't read math aloud clearly","Math TTS falls short of human experts","Matrices and summations stump TTS models","Text-to-speech fails at math intelligibility"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00018,"raw_usage":{"total_tokens":1249,"prompt_tokens":837,"completion_tokens":412,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":453,"completion_tokens_details":{"reasoning_tokens":353}},"tokens_in":453,"tokens_out":412,"duration_ms":4505,"temperature":1.0,"reasoning_tokens":353,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T10:35:16.621666+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Give the same five TTS models hand-crafted expert English pronunciations (the kind the human experts used) for the same 120 expressions, rerun the transcription test, and compare counts-of-correct. If transcription accuracy jumps to near the expert-audio level, the intelligibility gap is caused by the LLM-generated text, not by TTS synthesis; if it stays low, the synthesis itself is the bottleneck.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the handwritten mathematical expression dataset from which the 120 test expressions and eight categories were sampled."},{"cited_title":"These met- rics address RQ1","cited_arxiv_id":null,"evidence_quote":"Provides the MathBridge corpus and the argument that word-error rate alone is insufficient for evaluating spoken math."},{"cited_title":"Mathematical for- mulas in text to speech system,","cited_arxiv_id":null,"evidence_quote":"Prior text-to-speech-for-math system that assumed TTS is error-free, which this paper's evaluation stands against."},{"cited_title":"Mathvision: An accessible intelligent agent for visually impaired people to understand mathematical equa- tions,","cited_arxiv_id":null,"evidence_quote":"Prior math-to-speech vision-based system, also assuming error-free TTS, used as contrast for the evaluation gap."},{"cited_title":"MathReader : Text-to-Speech for Mathematical Documents","cited_arxiv_id":"2501.07088","evidence_quote":"Defines TeXBLEU, one of the three metrics used to score listener transcriptions."},{"cited_title":"Techniques for automatically correcting words in text,","cited_arxiv_id":null,"evidence_quote":"Supplies the Levenshtein distance underlying the LaTeX character error rate metric."},{"cited_title":"Towards a prosodic model for synthe- sized speech of mathematical expressions in mathml,","cited_arxiv_id":null,"evidence_quote":"Specifies the MUSHRA hidden-reference methodology adopted for the expert-comparison listening test."},{"cited_title":"Hamex-a handwritten and audio dataset of mathematical expressions,","cited_arxiv_id":null,"evidence_quote":"The automatic speech recognition model used to show that ASR-based metrics do not track human intelligibility."}],"review_version":1}