{"id":"0e2ab531-8f75-47ef-9cd3-b11cb5c930b7","arxiv_id":"2607.09675","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Male speakers account for 77% of speaking time on U.S. news and talk radio, with female shares never exceeding 35.8% in any topic and falling to 10.8% in talk-show segments.","lead":"Over 1,400 hours from 74 U.S. news/talk radio stations show male speakers taking 77% of airtime, with the gap holding through commute hours and every topic category. The public VANPY pipeline makes similar large-scale voice audits feasible for media and other domains.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.5","headline":"Topic-level gender claims rest on a 46%-accurate zero-shot classifier whose errors concentrate on the very categories used for the strongest disparity statements.","rationale":"The Reader correctly isolates the weakest assumption: the zero-shot topic pipeline. Gender accuracy is high enough that the 77% overall figure and the hour-of-day patterns survive ordinary scrutiny; the topic percentages do not. My concern is essentially the same as the Reader’s, only sharpened by the observation that the documented confusions fall precisely on the high-volume public-discourse categories that the paper uses to argue structural importance. No circularity or internal inconsistency is present; the limitation is purely measurement. Because the core airtime result remains usable while the topic claims require better labels (and ideally multi-day data), the appropriate verdict stays CONDITIONAL. I therefore leave the Reader’s verdict unchanged and mark agreement as “agree.”","tokens_in":16616,"tokens_out":584,"duration_ms":6546,"concrete_test":"Re-label a stratified sample of ≥500 segments (or the existing 222 plus additional draws) with two independent human annotators using the same 20-topic taxonomy; recompute female speaking-time share per topic using only the human-consensus labels. If the female share for Talk Show Segments rises above ~20% or Entertainment News falls below ~30%, or if any topic reverses rank order relative to Figure 6, the topic-level half of the strongest claim is not supported by the data.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper’s strongest claim packages three results: overall 77% male airtime, commute-hour gaps, and male dominance in every topic (lowest female share 10.8% in Talk Show Segments, highest 35.8% in Entertainment News). Gender classification is validated at 97.8% on 370 segments and the station/hour aggregates are therefore comparatively solid. Topic claims are not. Section 3.2 reports that facebook/bart-large-mnli zero-shot labels (candidate list generated by Claude from a single station’s transcript) achieve only 46% accuracy and weighted F1 0.49 on 222 stratified segments. Discussion §5 further notes that confusions are systematic among Community Affairs, Politics & Government, Miscellaneous, and General News—the categories that carry the bulk of total hours and the largest reported male-to-female ratios—and that advertisements and short segments are frequently mis-assigned. Because the per-topic percentages are simple aggregates of these noisy labels, the headline numbers 10.8% and 35.8% (and the claim of “male dominance across all examined content categories”) are not known to be robust to label noise. The overall airtime figure does not inherit this problem; the topic half of the strongest claim does.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper measures gender representation in U.S. news and talk radio from a single 24-hour multi-station recording (74 stations after filtering, >1,400 hours). Using the authors’ VANPY pipeline (diarization, gender classification, Whisper STT, zero-shot topic labels), it reports that male speakers account for 77% (SD 6.8%) of total speaking time, that the gap persists through commute hours (~7.5–9.5 min/h female vs ~30.6–32.9 min/h male), and that male speakers dominate every topic category, with female share lowest in Talk Show Segments (10.8%) and highest in Entertainment News (35.8%). Gender classification is manually validated at 97.8% on 370 segments; topic labels are validated at 46% accuracy / weighted F1 0.49 on 222 segments. VANPY is released as a reusable framework.","tokens_in":16930,"tokens_out":1324,"duration_ms":11910,"significance":"If the overall airtime and daypart results hold, the paper supplies a large-scale, audio-derived quantification of gender exposure on the most-listened U.S. radio format, going beyond staffing counts or small manual samples. The station-level and hourly distributions (Figs. 2–5) and the public VANPY pipeline are genuine contributions: gender classification is validated at high accuracy, the measurement is purely observational, and the tooling is reusable for other audio domains. The topic-level half of the claim is weaker and currently overstated relative to the reported label accuracy, so the paper’s lasting value is primarily the aggregate speaking-time and temporal results plus the open pipeline.","major_comments":[{"comment":"§3.2 and §5: Topic-level gender claims (Fig. 6; abstract and §4.3 numbers 10.8% Talk Show Segments, 35.8% Entertainment News, and “male dominance across all examined content categories”) rest on facebook/bart-large-mnli zero-shot labels whose manual check yields only 46% accuracy and weighted F1 0.49 on 222 segments. Discussion notes systematic confusions among Community Affairs, Politics & Government, Miscellaneous, and General News—the categories that carry the bulk of hours and the largest reported male-to-female ratios—plus ad and short-segment misassignment. These percentages are simple aggregates of noisy labels and are not shown to be robust. Either re-label a larger stratified sample (or a high-confidence subset), report uncertainty under label noise, or demote topic results to exploratory and remove the specific 10.8%/35.8% figures from the abstract and strongest claims.","section":null},{"comment":"§3.1 and axiom of representativeness: The entire corpus is one contiguous 24-hour window (2024-07-31 23:00–2024-08-01 23:00 UTC). Station programming, host schedules, and topic mix vary by day of week and news cycle; a single midweek summer day cannot by itself support claims about “typical” allocation or “consistent” patterns. At minimum, state this limitation prominently in the abstract/results and avoid language that generalizes beyond the sampled day; ideally add a second day or a multi-day subsample for a subset of stations to test stability of the 77% figure and daypart pattern.","section":null},{"comment":"§3.2 diarization and call-in quality: Gender accuracy is high (97.8%), but the authors acknowledge overlap segments, mixed-gender segments, and degraded call-in audio as error sources. Because talk radio is rich in call-ins and multi-speaker turns, residual diarization error can still bias duration aggregates (especially if female call-ins are more often truncated or mis-segmented). Report the fraction of duration discarded by YAMNet/non-speech filtering and by multi-speaker/ambiguous segments, and show that the 77% result is stable under reasonable exclusion thresholds.","section":null}],"minor_comments":[{"comment":"Abstract vs body: Abstract states female commute time “approximately 7.5–9.5 minutes per hour” and male “30.6–32.9”; §4.2 text gives broader ranges (female 7.3–12.9, male ~30–33). Align the numbers and cite the exact hours used for “commute periods.”","section":null},{"comment":"Fig. 6 and topic taxonomy: Candidate list was generated by Claude from one station’s transcript then generalized to 20 labels (Appendix B). State whether any post-hoc merging was done and whether “Commercial Breaks” were excluded from speaking-time denominators.","section":null},{"comment":"Typos and consistency: “consistant” (abstract), “allotted duration” vs “air-time,” mixed “female speaking time was approximately” vs “remained approximately.” Standardize terminology (speaking time / airtime / representation).","section":null},{"comment":"§2.2 related work: GMMP 2025 and French TV/radio studies are well cited; a short explicit comparison of the 23% female speaking-time figure to those benchmarks would help readers place the result.","section":null},{"comment":"Reproducibility: VANPY is said to be public; add a frozen commit/DOI, the exact model checkpoints (Whisper large, BART-MNLI, ECAPA-TDNN, gender model), and the station list already in Appendix A as a machine-readable file.","section":null}],"recommendation":"major_revision","confidential_remarks":"The overall 77% male airtime result and the daypart pattern look solid enough for a methods-plus-measurement paper in cs.CY once the topic claims are demoted or re-validated and the single-day limitation is stated clearly. I would not reject on the topic-accuracy issue alone if the authors reframe; I would reject if they insist on keeping 10.8%/35.8% as headline findings without better labels. Scope fit is good for a computational social science / media analytics venue."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The usable result here is the overall speaking-time gap: male voices take 77% (SD 6.8%) of airtime across 74 U.S. news/talk stations in a 24-hour snapshot of >1,400 hours. Gender classification was checked at 97.8% on 370 stratified segments, the station-level spread is wide but consistently male-heavy, and the commute-hour minute counts line up with the same pattern. That is a cleaner, listener-facing measurement than earlier head-count or schedule studies, and the open VANPY pipeline is a real practical contribution.\n\nWhat is new is the scale and the direct-from-audio method on the dominant U.S. format, not the existence of a gender gap (already shown for music radio, commercial FM talent, French broadcast news, and GMMP). The paper is honest about diarization failures on overlaps and call-ins and about the single-day window. Citation pattern is fine; self-citation of VANPY is tooling, not circular.\n\nThe soft spot is exactly where the stress-test says: topic claims. Zero-shot BART labels (Claude-generated 20-way list from one station) hit only 46% accuracy / 0.49 weighted F1 on 222 segments, with systematic confusions among Community Affairs, Politics, General News, and Miscellaneous—the categories that carry most hours and the biggest reported ratios. So the headline 10.8% (Talk Show) and 35.8% (Entertainment) numbers, and the blanket “male dominance across all topics,” are not known to be robust. Temporal “structural bias” claims also rest on one day. Those are real but secondary weaknesses; they do not sink the core airtime figure.\n\nThis is for media-studies and computational social-science people who want a large, reproducible U.S. radio baseline and a pipeline they can re-run. It deserves a serious referee. I would cite the 77% and the station/hour aggregates; I would not cite the per-topic percentages without better labels or multi-day data. Engage.","headline":"Solid 77% male airtime measurement on 74 U.S. news/talk stations; topic percentages are too noisy to trust at face value.","tokens_in":17515,"tokens_out":532,"would_cite":true,"duration_ms":5615,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Male speakers account for 77% of speaking time on U.S. news and talk radio, with the gap holding across the day and every content topic.","keywords":["gender representation","radio broadcasting","speaking time","news and talk radio","speaker diarization","media bias","automated audio analysis","topic analysis"],"falsifier":"Have human coders re-label a large stratified sample of the same segments and recompute male and female speaking shares by topic; if the topic-level gaps shrink or reverse under gold labels, those claims fail (the overall airtime gap can still be checked separately against gender-classification accuracy).","tokens_in":17482,"feed_emoji":"📻","tokens_out":890,"duration_ms":19713,"temperature":0.7,"pith_summary":"This paper measures who actually speaks on U.S. news and talk radio by processing more than 1,400 hours of audio from 74 stations recorded over one day. It finds that male voices take about three-quarters of all speaking time, that the imbalance remains during morning and evening commute hours, and that men dominate every topic category examined—including politics, community affairs, and talk shows. Female share peaks only at about 36% even in entertainment news and falls to roughly 11% in talk-show segments. Radio still reaches most Americans weekly, so the voices that fill the air shape who is heard as an authority. The authors also release the automated pipeline they used so the same measurements can be repeated on other audio.","feed_headline":"Men take 77% of airtime on US news-talk radio","feed_subtitle":"Gap holds in commute hours and every topic, from politics to talk shows","key_machinery":"An end-to-end audio pipeline that diarizes speakers, classifies gender from voice embeddings, transcribes speech, and assigns zero-shot topic labels, converting continuous multi-station radio streams into quantified speaking-time and topic shares by gender.","core_discovery":"Across 74 U.S. news and talk stations and more than 1,400 hours of filtered speech, male speakers account for 77% of total speaking time (standard deviation 6.8%). During commute periods female speaking time stays near 7.5–9.5 minutes per hour while male speaking time is about 30–33 minutes per hour. Male dominance appears in every topic category, with female representation lowest in talk-show segments (10.8%) and highest in entertainment news (35.8%).","pith_inferences":["If airtime exposure shapes perceived authority, the measured gap may reinforce who is treated as a legitimate voice on politics and community issues.","A single 24-hour sample leaves open whether the 77% figure is stable across weeks or seasons; multi-day replications would test that.","Merging short or interrupted segments into longer content windows could reassign some topic labels without changing the overall speaking-time gap.","The few relatively balanced public or diversity-oriented stations suggest programming and ownership choices, not only talent pools, can move the ratio."],"forward_implications":["Typical news-and-talk listeners hear male voices roughly three times as often as female voices overall.","Peak listening windows (commute hours) expose audiences to persistently male-dominated speech.","Public-discourse topics such as politics, community affairs, and general news carry strong male speaking-time majorities.","Even the most balanced stations in the sample stay short of parity, and most stations sit well below 30% female speaking time.","The same measurement pipeline can be applied to podcasts, other broadcast streams, or corporate audio to track representation at scale."],"fun_headline_variants":["Men claim 77% of speaking time on US news-talk radio","Women get under 23% of airtime across US news-talk stations","Male voices take 77% of hours on US news and talk radio","US news-talk radio: men speak 77% of total airtime","Talk shows allot women just 10.8% of speaking time"],"cache_read_input_tokens":128,"weakest_assumption_plain":"The per-topic gender findings rest on automatic topic labels whose accuracy, by the authors’ own manual check, is only about 46 percent.","fun_headline_variants_meta":{"raw":{"variants":["Men claim 77% of speaking time on US news-talk radio","Women get under 23% of airtime across US news-talk stations","Male voices take 77% of hours on US news and talk radio","US news-talk radio: men speak 77% of total airtime","Talk shows allot women just 10.8% of speaking time"]},"model":"grok-4.5","effort":"low","cost_usd":0.00405,"raw_usage":{"total_tokens":1315,"prompt_tokens":912,"num_sources_used":0,"completion_tokens":80,"cost_in_usd_ticks":40500000,"prompt_tokens_details":{"text_tokens":912,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":323,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":912,"tokens_out":80,"duration_ms":3432,"temperature":1.0,"reasoning_tokens":323,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-14T18:12:39.300260+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Have human coders re-label a large stratified sample of the same segments and recompute male and female speaking shares by topic; if the topic-level gaps shrink or reverse under gold labels, those claims fail (the overall airtime gap can still be checked separately against gender-classification accuracy).","supporting_citations":[],"review_version":1}