{"id":"02a31ef0-4e7d-4538-842d-9749ce7431b0","arxiv_id":"2507.16835","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"In a large production dataset, the Google STT plus GPT-4.1 plus Cartesia TTS interview stack outperformed four alternatives, while LLM-judged quality correlated weakly with user satisfaction.","lead":"This paper compared five speech-to-text, language model, and text-to-speech combinations inside a production AI interview system. It found that one stack, Google STT with GPT-4.1 and Cartesia TTS, scored best on both automated and human ratings, but that the automated scores barely predicted how satisfied users were.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Non-overlapping deployment periods leave C3's superiority as a plausible ranking, not an established causal result; the paper's own §5.1 limitation is the load-bearing gap.","rationale":"The paper's intended contribution is an empirical ranking of production STT/LLM/TTS stacks. For that ranking to support the stated conclusion, the observed differences must reflect the components rather than when each configuration happened to be deployed. The paper itself names this temporal confound in §5.1, and there is no randomization, no concurrent deployment, and no released dataset or code that would allow an external check. The reader's weakest assumption identifies exactly this gap, and my independent reading of the full text agrees that it is the most load-bearing threat to the central claim. I do not see the paper as internally inconsistent in its statistical procedures, and the LLM-as-a-judge method, while not independently validated in the paper, is a scalable tool whose limitations are partially acknowledged. The weaker post-hoc result for user satisfaction against C1 reinforces the need for conditionality but is secondary to the temporal confound. Because the reader already issued a CONDITIONAL verdict, my stress test does not move the verdict; it confirms that conditionality is the right level of confidence. If the requested concrete test fails to preserve C3's advantage, the verdict should move toward rejection of the headline ranking, but on the current evidence a conditional acceptance remains appropriate.","tokens_in":7161,"tokens_out":3536,"duration_ms":47777,"concrete_test":"Use deployment timestamps to build a time-matched or overlapping sample: restrict the analysis to interviews from all five configurations that occurred during the same calendar window, then additionally match on job role, candidate experience level, and interview language. Rerun the Welch ANOVA and Games-Howell post-hoc tests on accuracy, CQ overall, TQ overall, and rate star for this matched cohort. If C3's advantage over the other configurations, especially C1 on rate star, does not survive the time-matched comparison, the claimed superiority is an artifact of deployment period rather than of the STT/LLM/TTS combination.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that configuration C3 outperforms the other stacks on both LLM-judge metrics and user satisfaction rests on comparisons made across non-overlapping deployment periods. Section 5.1 explicitly acknowledges that these periods introduce potential confounds from seasonal variation, candidate pool changes, and system improvements over time. Because deployments were not randomized or concurrent, the higher means for C3 in Table 3 and the significant Welch ANOVA results in Table 4 cannot distinguish component-level effects from period-level effects. For instance, if C3 ran later in calendar time, general improvements to the production pipeline or a shift in candidate seniority could inflate its scores independently of the STT/LLM/TTS choices. The paper's own limitation statement, not an external criticism, is the weakest point in the causal chain from observed scores to the headline ranking. A second precision issue compounds this: §4.2 reports that C3 was significantly better than C2, C4, and C5 on rate star, but not C1, so the abstract's claim that C3 outperforms alternatives in user satisfaction is stronger than the pairwise evidence supports. Together these issues make the stack-level superiority claim plausible but not established.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a production-scale comparison of five STT x LLM x TTS configurations used in an AI-conducted job interview system. Using LLM-as-a-Judge scores (Claude 3.5 Sonnet) and post-interview user ratings, it reports that the Google STT + GPT-4.1 + Cartesia TTS stack (C3) achieves the highest means on conversational quality, technical quality, accuracy, and user satisfaction, with Welch ANOVA and Games-Howell post-hoc tests indicating significant differences. The paper also reports weak correlations between LLM-judge metrics and user ratings and proposes a component-level decomposition claiming that STT quality is the dominant factor. The manuscript includes the full evaluation prompt and detailed statistical tables, but several load-bearing limitations are acknowledged in Section 5.1.","tokens_in":7387,"tokens_out":2510,"duration_ms":30937,"significance":"If the causal ranking were established, this would be a valuable large-scale empirical comparison of cascaded speech-language systems in a real deployment, with practical guidance for component selection and a useful dual-evaluation methodology. The paper's strengths include the scale of the production data, the transparent disclosure of the full LLM-judge prompt, and the use of standard statistical procedures with effect sizes. However, the central claims about C3's superiority and the component-level attribution depend on comparisons across non-overlapping deployment periods, a non-factorial design, and an LLM judge whose agreement with human judgments is not demonstrated in this manuscript. The weak objective-user correlation is an interesting and falsifiable finding, but it does not by itself compensate for the causal identification gaps.","major_comments":[{"comment":"The manuscript's own limitation statement acknowledges that non-overlapping deployment periods introduce confounds from seasonal variation, candidate pool changes, and system improvements over time. Because C3's superiority is reported entirely through comparisons across these non-overlapping periods, the observed higher means and significant Welch ANOVA results cannot be attributed specifically to the STT/LLM/TTS choices. The abstract's phrasing that C3 'outperforms alternatives' is stronger than the evidence supports unless the authors provide deployment date ranges, concurrent control measurements, or a sensitivity analysis showing that period-level factors cannot explain the effect. This is the load-bearing gap for the paper's headline claim.","section":"Section 5.1 (Temporal Confounds); Tables 3-4"},{"comment":"The post-hoc description states that for rate star, C3 was significantly better than C2, C4, and C5, but not C1. The abstract nonetheless claims that C3 outperforms alternatives in user satisfaction scores. Since the pairwise comparison against C1 is not significant, the user-satisfaction component of the central claim is not fully supported. The paper should either report the full post-hoc matrix for rate star, including the C3 vs. C1 comparison, and temper the abstract, or provide additional evidence that the lack of significance is due to sample size or other identifiable factors.","section":"Section 4.2 and Abstract (User Satisfaction Claim)"},{"comment":"RQ3 asks whether LLM-based evaluation can reliably assess voice-based AI interactions, and the abstract claims a 'validated evaluation methodology.' The only validation support is a citation to the authors' prior work [1]; this manuscript reports no agreement statistics between the Claude 3.5 Sonnet judge and human raters, no inter-rater reliability, and no calibration analysis. The prompt itself contains specific instructions (e.g., replacing 'Mid-level' with 'Experienced'; scoring only the interviewer's responses, not the candidate's), which may introduce systematic biases that are not evaluated. Without judge-validation evidence with confidence intervals or agreement coefficients, the construct validity of the primary outcome metrics remains unestablished.","section":"Section 3.4 and Appendix A (LLM-as-a-Judge Validation)"},{"comment":"The component contribution analysis claims that STT has the dominant impact and that TTS contributes smaller but meaningful improvements. However, the five configurations in Table 1 form a non-factorial design: Google STT appears only with GPT-4.1 or GPT-4o, while Whisper appears only with GPT-4o or Groq2, and Cartesia TTS appears in only one configuration. The reported 'components effects' are therefore confounded with the specific combinations in which each component appears. The statement that Google STT 'consistently outperforms' Whisper STT is not supported by matched comparisons that hold the other two components fixed. A factorial or matched-pair design, or a clear regression model with interaction terms, is needed before these component-level conclusions can be drawn.","section":"Section 4.5 and Table 1 (Component Contribution Analysis)"}],"minor_comments":[{"comment":"The abstract says 'over 300,000 AI-conducted job interviews' while Section 3.3 describes 'over 5,000 AI-conducted interviews, sampled and segmented.' Please clarify whether 300,000 refers to the total population and 5,000 to the analyzed sample, and state the sampling procedure.","section":"Abstract and Section 3.3"},{"comment":"There are typographical errors: 'Speech-to-T ext' and 'T ext-to-Speech' should be 'Speech-to-Text' and 'Text-to-Speech.'","section":"Section 3.1"},{"comment":"The soft skills row reports both a Welch p-value of 1.000 and a Kruskal-Wallis result with p=0.136; the table's 'Test Used' column says Kruskal-Wallis, but the text in Section 4.2 uses the Welch p-value. Please specify which test was used for the soft skills metric and clarify the discrepancy.","section":"Table 5"},{"comment":"The sentence 'The analysis tells important findings about component contributions' appears to be missing a verb; it should likely read 'reveals important findings' or 'provides important findings.'","section":"Section 4.5"},{"comment":"The figures are referenced in the text but the figure captions do not include the numerical values or confidence intervals. Adding error bars or boxplot annotations would help readers assess the overlap in distributions, especially given the small effect sizes.","section":"Section 4.4 and Figures 1-2"}],"recommendation":"major_revision","confidential_remarks":"The paper is a real-world production study, and the authors are transparent about some limitations. The main concern is that the temporal confound, acknowledged in Section 5.1, directly undermines the causal ranking that the abstract states as fact. Additionally, the LLM-as-a-Judge is described as 'validated' through a citation to a prior paper without presenting any agreement statistics in this manuscript, which is problematic for a paper whose RQ3 is about judge reliability. These issues are fixable in principle with additional data or with a substantially more cautious framing, so I do not recommend rejection, but the present version does not support its strongest claims."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a useful, readable production study of five STT×LLM×TTS stacks in a real AI-interview system, built on a big dataset (over 5,000 sampled interviews from 300,000+). The descriptive ranking and the weak correlation between LLM-judge scores and user satisfaction are worth knowing. But the headline causal claim — that Google STT + GPT-4.1 + Cartesia is the best stack — is undercut by non-overlapping deployment periods, which the authors themselves acknowledge in §5.1. Treat this as a plausible ranking, not an established result.\n\nWhat is genuinely good: the modular swap design yields pairwise controlled comparisons (C1 vs C3 isolates TTS, C2 vs C4 isolates STT), the statistical analysis is appropriate (Levene, Welch ANOVA, Games-Howell), and the appendix contains the full judge prompt. The weak correlation observation (most correlations with user ratings below 0.11) is a genuinely interesting caution for anyone building automated evaluation pipelines. The paper also flags its own limitations: self-selected ratings, domain specificity, limited TTS variety.\n\nThe soft spots, in proportion: the temporal confound is the big one. Configurations were deployed sequentially, not concurrently or randomly, so C3's higher means could reflect candidate season, pipeline improvements, or candidate pool shifts rather than component quality. The abstract overstates the user-satisfaction claim: the post-hoc test shows C3 was not significantly better than C1 on rate star, only better than C2, C4, and C5. The LLM judge is validated by citing the authors' prior paper [1], with no agreement statistics reported here. No data release, and per-configuration sample sizes are missing, which weakens effect-size interpretation. The component contribution analysis is not a full factorial, though the pairwise contrasts help.\n\nWho this is for: product teams choosing voice AI components, and researchers interested in LLM-as-a-judge limits. It deserves serious referee time — the scale and honest framing outweigh the confound — but it needs revision: soften the causal language, add concurrent or controlled deployment evidence, and validate the judge in-paper. I'd send it to review.","headline":"A useful, large-scale production comparison of STT×LLM×TTS stacks, but the headline ranking is plausible rather than established because deployments were non-overlapping and the LLM judge is not validated in-paper.","tokens_in":7896,"tokens_out":2562,"would_cite":true,"duration_ms":29221,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that a Google STT + GPT-4.1 + Cartesia TTS stack outperforms four other production configurations for AI-conducted interviews on both automated quality metrics and user ratings.","keywords":["speech-to-text","text-to-speech","large language models","LLM-as-a-Judge","conversational AI","cascaded architecture","user satisfaction","AI interviews"],"falsifier":"Re-run the five configurations in randomized, overlapping deployment windows, or add interview date as a covariate to the Welch ANOVA; if the advantage of Google STT + GPT-4.1 + Cartesia TTS disappears or shrinks to non-significance, the reported ordering is an artifact of timing rather than component quality.","tokens_in":6983,"feed_emoji":"🎙️","tokens_out":4928,"duration_ms":55971,"temperature":0.7,"pith_summary":"The paper asks which combinations of speech-to-text, large language model, and text-to-speech components work best in a voice-based AI interviewer. Analyzing transcripts from thousands of real job interviews, it claims that a stack with Google's STT, GPT-4.1, and Cartesia's TTS beats four other production configurations on LLM-judged conversational and technical quality and on candidate ratings. It also claims that these objective quality metrics correlate only weakly with user satisfaction, so automated scores do not capture what makes candidates happy. A sympathetic reader would take the paper's three contributions as a production-scale comparison, a validated LLM-judge evaluation method, and evidence that component choice in cascaded voice AI matters most at the transcription stage.","feed_headline":"Best AI-interview voice stack: Google STT, GPT-4.1, Cartesia TTS","feed_subtitle":"Production data from 300,000 interviews shows this combo tops both automated quality scores and candidate ratings.","key_machinery":"The argument is carried by a production system whose STT, LLM, and TTS modules can be swapped independently, creating five naturally occurring configurations; by LLM-as-a-Judge evaluation, in which a Claude 3.5 Sonnet model scores each transcript on conversational sub-metrics (dialogue flow, response building, acknowledgement) and technical sub-metrics (skill alignment, logical progression, question clarity); and by statistical machinery of Levene's tests, Welch's ANOVA, Games-Howell post-hoc tests, and Pearson correlations that turn score distributions into comparisons. The same five configurations are re-sliced by component to attribute differences to STT, LLM, or TTS.","core_discovery":"The central claim is that component choice in a cascaded STT x LLM x TTS pipeline materially changes both measured interview quality and user satisfaction. Using five production configurations on over 5,000 interviews drawn from a system running about 1,500 interviews per day, the paper reports that Google STT + GPT-4.1 + Cartesia TTS scored highest on accuracy (8.12), conversational quality (8.78), technical quality (8.57), and average user rating (4.53), with statistically significant Welch ANOVA differences for every main metric (soft skills excluded because n=10). A second claim is that automated LLM-judge metrics and candidate star ratings are nearly orthogonal, with most correlations below 0.11, which the authors interpret as evidence that user experience depends on factors like perceived empathy and voice naturalness that technical quality scores miss. A third claim is that decomposition of the five configurations shows STT has the dominant effect, the LLM a moderate effect, and TTS a smaller but systematic effect.","pith_inferences":["Going beyond the paper, the weak correlation between automated quality and user satisfaction suggests that optimizing purely for objective conversational quality may not raise user ratings; factors such as perceived empathy, voice naturalness, and interaction design may need independent measurement.","The five configurations are not a full factorial design, so a fully crossed experiment with multiple STT, LLM, and TTS options would disentangle interaction effects that the current component decomposition can only approximate.","If transcription errors indeed cascade, then testing robustness under noisy audio or accented speech could show even larger STT effects than the average-case comparison reported here.","Using an ensemble of multiple LLM judges, as the paper names for future work, might produce scores that correlate more strongly with user satisfaction than a single judge's scores."],"forward_implications":["Engineers building cascaded voice AI systems should prioritize STT quality, since the paper attributes the largest performance differences to the transcription component.","The Google STT + GPT-4.1 + Cartesia TTS combination is the current best-performing production stack for AI-conducted interviews.","Objective LLM-judge scores should be supplemented with direct user feedback, because the paper finds these metrics capture largely different aspects of system performance.","LLM-as-a-Judge can serve as a scalable, cost-effective evaluation method for voice-based conversational AI, though its scores will not fully predict user satisfaction.","Improving the TTS component yields smaller but consistent gains, which may still matter for user experience in deployed voice systems."],"supporting_citations":[{"why":"Provides the prior validation of AI-assisted recruitment evaluations against human data, grounding the paper's dual evaluation approach.","marker":"[1]"},{"why":"A unified speech-to-dialogue benchmark toolkit that motivates systematic evaluation of pipeline combinations, providing the methodological contrast for this production-scale study.","marker":"[2]"},{"why":"An asynchronous agent architecture for real-time voice agents that supports the production system's parallel STT and TTS design.","marker":"[3]"},{"why":"Establishes LLM-as-a-Judge as a scalable and reliable evaluation method, which the paper relies on for its automated scoring framework.","marker":"[7]"}],"fun_headline_variants":["Google STT, GPT-4.1, Cartesia TTS tops AI interview stacks","300k interviews show best AI interview stack: Google, GPT-4.1, Cartesia","User ratings and technical metrics diverge in AI interviews","STT choice matters most in AI interview pipeline quality"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The comparison assumes that the five configurations were tested under comparable conditions; because deployments did not overlap in time, seasonal candidate differences or unrelated system improvements could produce the same pattern.","fun_headline_variants_meta":{"raw":{"variants":["Google STT, GPT-4.1, Cartesia TTS tops AI interview stacks","300k interviews show best AI interview stack: Google, GPT-4.1, Cartesia","User ratings and technical metrics diverge in AI interviews","STT choice matters most in AI interview pipeline quality"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000237,"raw_usage":{"total_tokens":1497,"prompt_tokens":925,"completion_tokens":572,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":541,"completion_tokens_details":{"reasoning_tokens":492}},"tokens_in":541,"tokens_out":572,"duration_ms":6141,"temperature":1.0,"reasoning_tokens":492,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T17:01:50.738484+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the five configurations in randomized, overlapping deployment windows, or add interview date as a covariate to the Welch ANOVA; if the advantage of Google STT + GPT-4.1 + Cartesia TTS disappears or shrinks to non-significance, the reported ordering is an artifact of timing rather than component quality.","supporting_citations":[{"cited_title":"Better together: Quantifying the benefits of ai-assisted recruitment, 2025","cited_arxiv_id":null,"evidence_quote":"Provides the prior validation of AI-assisted recruitment evaluations against human data, grounding the paper's dual evaluation approach."},{"cited_title":"ESPnet-SDS: A unified all-in-one speech-to-dialogue system","cited_arxiv_id":null,"evidence_quote":"A unified speech-to-dialogue benchmark toolkit that motivates systematic evaluation of pipeline combinations, providing the methodological contrast for this production-scale study."},{"cited_title":"Machine learning and information theory concepts towards an AI Mathematician","cited_arxiv_id":"2403.04571","evidence_quote":"An asynchronous agent architecture for real-time voice agents that supports the production system's parallel STT and TTS design."},{"cited_title":"{index}\" - LLM model: {model} - interview_transcript:","cited_arxiv_id":null,"evidence_quote":"Establishes LLM-as-a-Judge as a scalable and reliable evaluation method, which the paper relies on for its automated scoring framework."}],"review_version":1}