{"id":"61d9c866-8b13-401d-865b-fc42067597ba","arxiv_id":"2411.08135","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Group-annotated MuTox reveals that speech-aware inference reduces false-positive bias against group mentions in English and Spanish, while transcript correction barely changes it.","lead":"The authors add demographic group annotations to a multilingual speech toxicity dataset and compare speech-based and text-based toxicity detectors. They find that models that can hear the audio at inference time produce fewer false positives on samples that mention demographic groups, especially ambiguous ones.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The core evidence for 'speech at inference reduces bias' is MUTOX vs MUTOX-ASR, but MUTOX-ASR is a jointly speech-text-trained model with audio dropped at test time, not a text-only baseline; this confound, plus an undefined FPR definition for ambiguous samples, undermines the abstract's central…","rationale":"The paper makes a real contribution with a carefully constructed group-annotation dataset, and the MUTOX versus MUTOX-ASR design is a legitimate ablation for the role of inference-time modality availability. The reader's weakest-assumption analysis correctly identifies that MUTOX-ASR is not a genuine text-only system, so the comparison conflates the benefit of speech with the cost of an unexpected missing input. Because the abstract and introduction generalize to 'text-based classifiers,' the central claim as stated is stronger than what the ablation supports. The ambiguous-sample FPR issue compounds this: 'Cannot say' and 'No consensus' are not binary ground-truth labels, so reporting an FPR on them requires an explicit and justified denominator, which the paper does not provide. Both issues are addressable through re-analysis of the released annotations and one additional baseline, so the conditional verdict is appropriate and no verdict change is needed.","tokens_in":14917,"tokens_out":9133,"duration_ms":94690,"concrete_test":"Using the released annotations, recompute group-mention FPR for MUTOX, MUTOX-ASR, and a genuinely text-only classifier trained without speech data (e.g., fine-tuned on MUTOX text transcripts), and report FPR on ambiguous samples both with ambiguous labels excluded and with an explicit binarization rule. If a true text-only model does not show the elevated FPR seen for MUTOX-ASR, the speech-benefit claim is confounded by modality dropout; if the ambiguous FPR changes when the denominator is defined, the 'particularly for ambiguous' claim is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The only comparison isolating inference-time speech access is MUTOX versus MUTOX-ASR (Table 2). MUTOX-ASR is not a text-only or cascaded detector; it is the same jointly speech-text-trained model with audio removed at inference. The higher FPR for MUTOX-ASR may reflect train/test modality mismatch rather than an intrinsic property of text-based inference. The abstract's phrasing that 'text-based biases are mitigated by speech-based systems' overstates what this ablation supports. Additionally, Fig. 3b/c report 'FPR' on samples labeled 'Cannot say' or 'No consensus', but standard FPR requires binary gold labels; no passage defines how ambiguous samples enter the denominator. If ambiguous samples are treated as negatives, MUTOX's zero FPR on that subset becomes an artifact of an unstated binarization rule rather than evidence about speech context.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a new set of group annotations for the English and Spanish test partitions of the MuTox speech toxicity dataset, and uses them to compare four toxicity classifiers: ETOX (wordlist), Detoxify (text-only neural network), MuTox-ASR (jointly speech-text trained, but with only text at inference), and MuTox (jointly trained, with raw speech and text at inference). The authors report that MuTox, the only model with access to speech at inference, shows reduced false-positive rates on utterances mentioning demographic groups and zero false positives on ambiguous samples, while MuTox-ASR shows elevated false-positive rates on group mentions. They also find that correcting ASR transcripts has little effect on false-positive rates, concluding that improving classifiers rather than transcription pipelines is more helpful for reducing group bias.","tokens_in":15116,"tokens_out":4536,"duration_ms":44957,"significance":"If the central claim holds, the paper makes a useful contribution: it provides the first fairness-audit annotations for a multilingual speech toxicity dataset and offers practical guidance for speech-first toxicity detection. The annotation protocol is rigorous and the public release of the group annotations is a concrete asset for the community. However, the main empirical conclusion rests on a single ablation—MuTox versus MuTox-ASR—that is confounded by a train/test modality mismatch, and the computation of false-positive rates on ambiguous samples is underspecified. These issues must be addressed before the abstract's causal claim about 'access to speech data during inference' can be accepted.","major_comments":[{"comment":"This is the central confound and must be fixed or the claims must be reframed to only state that MuTox-ASR, a speech-trained model evaluated without audio, exhibits higher FPR under this modality dropout, which is a different and weaker claim.","section":"§4.1, Table 2, §4.3"},{"comment":"The same issue affects the comparison of group versus no-group ambiguous samples in Fig. 3c.","section":"§4.2, §4.3, Fig. 3b/3c"},{"comment":"The 'no significance testing' limitation may be acceptable for a purely exploratory dataset-release paper, but the paper makes strong comparative claims (e.g., 'eliminates false positives', 'reduces bias') that go beyond description and need some measure of uncertainty.","section":"§8 and Fig. 4"}],"minor_comments":[{"comment":"This would help readers assess whether the FPR differences are robust or an artifact of a particular operating point.","section":"§4.2"},{"comment":"Consider adding a caveat in the abstract or conclusion.","section":"§5, Fig. 5"},{"comment":"No further action needed.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The paper is by Meta FAIR authors and evaluates Meta's own MuTox dataset and MuTox/MuTox-ASR models, with annotations produced by the authors themselves. This creates an independence concern that is not disqualifying but should be made explicit in the paper, ideally with a statement about the relationship and a call for independent audit of the released annotations. Additionally, the central comparison should be supplemented with a genuinely text-only baseline; if that is not possible, the authors should temper the abstract's causal language. The dataset contribution itself is solid and likely worth publishing after the technical issues are addressed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Punchline: the group annotations are the real contribution; the bias-reduction claim does not hold up as stated.\n\nThe paper adds a careful set of group and toxicity labels for the MuTox English and Spanish test sets, with a three-stage annotation process and public release. That is a genuinely useful resource, and the paper's survey of the literature supports its claim to be the first group-level fairness audit for speech toxicity. I would gladly use it.\n\nThe empirical comparison is the soft spot. The key contrast is MUTOX versus MUTOX-ASR. MUTOX-ASR is not a text-only or cascaded detector; it is the same jointly speech-text-trained model with audio dropped at inference. So the difference could reflect the cost of an unexpected missing modality, not the intrinsic bias of text-based inference. The paper body is more careful about this than the abstract, which says 'text-based biases are mitigated by speech-based systems' - that overstates the ablation.\n\nAlso, the FPR on 'ambiguous' samples (Cannot say / No consensus) is never defined. Standard FPR needs a binary gold; these labels are not binary. If the authors treat all non-'Yes' as negative, then zero FPR on that subset is partly an artifact of that coding. They need to state the rule and show sensitivity to it.\n\nThey honestly note the absence of statistical testing. That is a real limit because per-group cells in Fig. 4 are small; the apparent differences could be noise. Confidence intervals or raw counts would help.\n\nThe transcription-error finding (correcting ASR transcripts barely changes FPR) is a nice secondary result, appropriately scoped to English/Spanish.\n\nBottom line: the dataset deserves a home, but the headline empirical claim needs a genuine text-only baseline, explicit handling of ambiguous labels, and some measure of uncertainty. I would accept this for peer review but require those revisions.","headline":"The group annotations are the real contribution; the bias-reduction claim does not hold up as stated because the key comparison is confounded.","tokens_in":15642,"tokens_out":4634,"would_cite":true,"duration_ms":43987,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Access to raw speech at inference time reduces false positives on demographic group mentions in toxicity detection.","keywords":["speech toxicity detection","group mention bias","false positive rate","MUTOX dataset","group annotations","multimodal inference","ambiguous samples","transcription error"],"falsifier":"Synthesize neutral-prosody renderings of the same transcripts and feed them to MUTOX: if its false-positive rate on group-mention clips rises to the level of MUTOX-ASR, the benefit is prosodic; if it stays low, the inference-time-speech explanation would need revision.","tokens_in":14721,"feed_emoji":"🗣️","tokens_out":10088,"duration_ms":84851,"temperature":0.7,"pith_summary":"This paper asks whether speech-based toxicity detectors are less biased against mentions of demographic groups than text-based ones. To answer it, the authors annotate 1,954 English and Spanish clips from MUTOX for toxicity, group mentions, and transcript accuracy, then compare four classifiers: a wordlist, a text neural network, and the MUTOX speech model with and without access to audio at inference. They report that the full speech model produces fewer false positives on group-mention clips than the same model fed only text transcripts, and that the difference concentrates on clips human annotators found ambiguous. They also find that correcting ASR transcripts does little to reduce the bias, pointing to the classifier rather than the transcription pipeline as the main lever. If right, the result argues for keeping raw audio available at test time in speech toxicity systems and focusing fairness effort on classifiers.","feed_headline":"Speech audio cuts false toxicity flags on group mentions","feed_subtitle":"Hearing the audio, not just transcripts, lowers false positives on demographic mentions—especially in ambiguous clips.","key_machinery":"The central object is the paired comparison MUTOX versus MUTOX-ASR: the same classifier, trained jointly on speech and text with SONAR embeddings (a sentence-level multilingual representation space), run either with raw audio plus ASR text (MUTOX) or with only the ASR text's SONAR embedding (MUTOX-ASR). This pairing is meant to isolate inference-time access to speech. The other machinery is a new annotation layer over MUTOX's test set, covering toxicity, group mentions, and corrected transcripts, which turns the dataset into a bias audit instrument, along with the false-positive-rate measurement on group-mention and ambiguous subsets.","core_discovery":"The paper's central discovery is that a toxicity classifier trained on speech and text and given both at inference (MUTOX) has a lower false-positive rate on samples mentioning demographic groups than the same model given only text at inference (MUTOX-ASR), while text-only baselines show the opposite pattern. On ambiguous clips, those labeled Cannot say or No consensus, MUTOX and DETOXIFY have zero false positives, whereas MUTOX-ASR's false-positive rate rises when groups are mentioned. The paper interprets this as evidence that speech carries prosodic and contextual cues that let the model avoid treating neutral group mentions as toxicity, and it argues that the effect comes from inference-time access, not from training, because both MUTOX variants were trained identically. It also shows that replacing ASR transcripts with annotator-corrected transcripts barely changes the false-positive rate, which the paper takes to mean transcription is not the driver of group bias.","pith_inferences":["Beyond the paper, if prosody is the active signal, resynthesizing the same utterances with flat intonation should make MUTOX's false positives approach MUTOX-ASR's; running that manipulation would test the mechanism directly.","Beyond the paper, the released group annotations could double as an ASR fairness audit, checking whether clips mentioning marginalized groups are disproportionately mistranscribed.","Beyond the paper, the transcription result likely generalizes only where ASR is strong; for lower-resourced languages, better transcripts may still be the cheaper fix."],"forward_implications":["Deployed speech toxicity detectors that can listen to raw audio at test time should produce fewer false positives on benign group-mention speech than cascaded ASR-to-text systems, with the largest gains on clips annotators find ambiguous.","Models trained jointly on speech and text should not be converted to text-only inference; removing audio at test time appears to make them lean harder on group mentions as toxicity cues.","For English and Spanish, better ASR transcription is unlikely to reduce group-mention false positives; modifying the classifier is the more direct lever.","The released MUTOX group annotations provide a reusable benchmark for auditing future speech toxicity systems across gender, race and ethnicity, and religion in two languages."],"supporting_citations":[{"why":"Supplies the MUTOX dataset and the MUTOX speech-text model that the paper audits and compares.","marker":"Costa-jussà et al., 2024"},{"why":"Defines the false-positive-rate bias metric and documents group-mention bias in text toxicity classifiers.","marker":"Dixon et al., 2018"},{"why":"Provides Civil Comments group annotations and nuanced bias metrics that the paper's annotation scheme extends to speech.","marker":"Borkan et al., 2019"},{"why":"Provides SONAR embeddings, the shared representation underlying both MUTOX variants and the modality-dropout comparison.","marker":"Duquenne et al., 2023"},{"why":"Provides the ASR model whose transcripts appear in MUTOX and are corrected in the transcription-error analysis.","marker":"Radford et al., 2023"},{"why":"Supplies DETOXIFY, the text-only neural baseline in the classifier comparison.","marker":"Hanu, 2020"},{"why":"Supplies ETOX, the wordlist-based multilingual baseline in the classifier comparison.","marker":"Costa-jussà et al., 2023"}],"fun_headline_variants":["Speech audio reduces toxicity false flags on group mentions","Hearing tone cuts false toxicity flags on demographic mentions","Audio input, not transcripts, lowers group-bias false positives","Toxicity AI less biased when it hears speech, not just text","Speech data at inference trims bias in toxicity detection"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The main conclusion rests on comparing the speech model with a version of itself that is denied audio at test time; if that denial is not a fair proxy for a genuinely text-based system, the measured bias reduction could be an artefact of the model being trained with speech but tested without it.","fun_headline_variants_meta":{"raw":{"variants":["Speech audio reduces toxicity false flags on group mentions","Hearing tone cuts false toxicity flags on demographic mentions","Audio input, not transcripts, lowers group-bias false positives","Toxicity AI less biased when it hears speech, not just text","Speech data at inference trims bias in toxicity detection"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00051,"raw_usage":{"total_tokens":2428,"prompt_tokens":839,"completion_tokens":1589,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":455,"completion_tokens_details":{"reasoning_tokens":1508}},"tokens_in":455,"tokens_out":1589,"duration_ms":18245,"temperature":1.0,"reasoning_tokens":1508,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T21:55:41.159800+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Synthesize neutral-prosody renderings of the same transcripts and feed them to MUTOX: if its false-positive rate on group-mention clips rises to the level of MUTOX-ASR, the benefit is prosodic; if it stays low, the inference-time-speech explanation would need revision.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the false-positive-rate bias metric and documents group-mention bias in text toxicity classifiers."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides Civil Comments group annotations and nuanced bias metrics that the paper's annotation scheme extends to speech."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the ASR model whose transcripts appear in MUTOX and are corrected in the transcription-error analysis."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies DETOXIFY, the text-only neural baseline in the classifier comparison."}],"review_version":1}