{"id":"bfdd9ab9-775c-45f0-94f8-19d5ae1c62f3","arxiv_id":"2607.23808","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"Indic DiarBench is an open ~108-hour joint diarization+ASR benchmark spanning all 22 scheduled Indian languages with human-corrected speaker-attributed transcripts and commercial/LLM baselines.","lead":"The paper releases Indic DiarBench, ~108 hours of human-annotated multi-speaker audio covering all 22 scheduled Indian languages for joint diarization and ASR. It gives the field a shared test set where commercial APIs and multimodal models still fail on overlap, code-mixing, and low-resource languages.","discovery_kind":"extension","skeptic_critique":{"model":"moonshotai/kimi-k3","headline":"The \"all 22 languages\" coverage is nominal for 12 of them: ~1.5 near-field hours each, on which the paper's per-language rankings and family-level conclusions rest without any uncertainty quantification.","rationale":"I agree with the reader that metric-trustworthiness is the weak region, but I locate the load-bearing mechanism differently. The reader emphasized annotation label noise; on inspection, the near-field subset (~53 hrs, the largest) has structurally reliable speaker labels because participants were non-co-located with per-speaker microphones, and annotators were barred from introducing new speakers — so the gold-standard-noise concern applies mainly to the ~55 hrs of far-field/ITW audio, and even there visual cues and superchecks provide partial cover. The concern with no mitigation is sample size: half the claimed language coverage is ~1.5 hrs each, and the paper's most distinctive analytical findings (language difficulty rankings tied to overlap, the family-level cpWER gap, the claim that systems \"remain weak in low-resource settings\") rest on those smallest slices with no uncertainty analysis. This does not undermine the resource's existence or release — the benchmark is real, public, and fills a genuine gap — so the CONDITIONAL verdict stands, but the conditions should explicitly include per-language uncertainty quantification, alongside the reader's IAA and conflict-disclosure conditions. If the bootstrap test shows wide intervals, the fix is cheap (hedge §5, report CIs); if tight, the paper is stronger than its current presentation shows.","tokens_in":9757,"tokens_out":2213,"duration_ms":93972,"concrete_test":"Using the released corpus and scoring scripts, bootstrap-resample (1,000 draws, session-level) the per-language DER/cpWER for Santali, Urdu, Dogri, Telugu, and Maithili and report 95% CIs. If the CIs for the ~1.5-hr languages exceed ~10pp wide, the §5 easiest/hardest-language rankings and the Dravidian-vs-Indo-Aryan 5pp cpWER gap are not statistically supported and should be retracted or hedged; if they are tight, the concern does not land. Also publish the ITW-condition overlap average to reconcile §5's 6.5% with Table 2.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The strongest claim says the benchmark \"spans all 22 scheduled Indian languages\" and supports \"baselines showing current systems remain weak under overlap and low-resource settings.\" The spanning claim is technically true but load-bearing in a thin sense: Table 2 shows 12 languages (Assamese, Bodo, Dogri, Kashmiri, Konkani, Maithili, Manipuri, Nepali, Sanskrit, Santali, Sindhi, Urdu) exist only as ~1.1–1.6 hours of near-field audio. At ~1.5 hours of multi-speaker speech, a DER or cpWER estimate is over perhaps a few hundred speaker turns; per-language point estimates plausibly carry ±5–15pp uncertainty. Yet §5 draws language-level conclusions on exactly these languages — \"Santali (6.5% overlap ... 9.7% DER) and Urdu (12.8% DER) are among the easiest\" vs. \"Telugu ... most challenging\" — and a cross-family claim that \"Dravidian languages show near-field cpWER roughly 5 percentage points above Indo-Aryan languages at comparable DER,\" all without confidence intervals or significance tests. The duration-weighted aggregates in Table 3 are dominated by the 8–10 well-covered languages, so the headline \"all 22\" framing and the low-resource-language findings are the parts of the claim least supported by the data volume. Notably, the reader's flagged concern (annotation label noise) is partly mitigated for the largest subset: near-field meetings used one mic per non-co-located speaker, giving near-ground-truth speaker timing for ~53 hrs, and annotators could not invent speakers there. The sample-size issue, by contrast, has no such mitigation in the text. A secondary internal tension: §5 states in-the-wild recordings \"have the lowest overlap (6.5%),\" but Table 2's 6.5% belongs to Santali, which has no ITW audio; the ITW-specific overlap average is never shown, so this comparative claim is unverifiable from the paper as written.","agreement_with_reader":"partial"},"referee_report":{"model":"moonshotai/kimi-k3","summary":"The paper introduces Indic DiarBench, an open-access benchmark for joint speaker diarization and speaker-attributed ASR covering all 22 scheduled Indian languages. The corpus comprises ~108 hours of multi-speaker audio in three conditions: near-field meetings (~53 h, one microphone per non-co-located speaker, all 22 languages), far-field meetings (~27 h, 8 languages), and in-the-wild YouTube audio (~28 h, 10 languages). Annotations are produced by a five-stage human-in-the-loop pipeline (multi-system ASR bootstrap, professional correction, dual-format code-mixed transcription, QC, expert supercheck) yielding RTTM speaker timing and speaker-attributed transcripts. The authors evaluate seven systems (commercial APIs, multimodal LLMs, and the Sarvam pipeline) using DER (no collar, overlap included), cpWER, and WDER, with duration-weighted aggregates, DER decomposition, per-language heatmaps, and an overlap-correlation analysis. Headline findings: the Indic-specialized Sarvam pipeline leads (16.0% DER, 38.8% cpWER), multimodal LLMs trade strong transcription for very high missed detection, and overlap ratio correlates strongly with both DER and cpWER.","tokens_in":10179,"tokens_out":3277,"duration_ms":87830,"significance":"If the data and annotations are as described, this is a useful and genuinely novel resource: the first joint diarization + speaker-attributed ASR benchmark spanning all 22 scheduled Indian languages, a clear gap left by DISPLACE (which decoupled the ASR track and covered fewer languages). Strengths worth naming: the corpus and evaluation protocols are publicly released (reproducible resource); the near-field subset's one-mic-per-speaker, non-co-located design gives near-ground-truth speaker timing for ~53 hours and structurally prevents speaker-label invention there; the dual-format (native-script / Romanized) WER convention is a principled, stated choice for code-mixed evaluation; and the baseline suite produces concrete, falsifiable reference numbers across seven named systems with a DER error decomposition. The metrics (DER, cpWER, WDER) are standard community definitions, so there is no circularity concern. The benchmark is likely to be used and cited by the Indic speech community regardless of the reservations below.","major_comments":[{"comment":"No quantitative evidence of gold-label quality is provided. The five-stage pipeline is described, but there is no inter-annotator agreement measurement (e.g., DER/cpWER between double-annotated files, or kappa on speaker labels) on any subset. This matters most exactly where the benchmark is most valuable: the in-the-wild subset, where annotators may add/merge/remove speakers, and the high-overlap segments the paper itself says 'often require multiple rounds of review.' Without a label-noise estimate, the reader cannot tell whether, e.g., the 4–6 pp WDER gaps between systems in Table 3 exceed annotation noise. Please double-annotate a stratified sample (including high-overlap and in-the-wild files) and report agreement in the same units as the benchmark metrics.","section":"§3.2 (Annotation Pipeline)"},{"comment":"Per-language and language-family conclusions are drawn on very thin data without uncertainty quantification. Twelve languages (Assamese, Bodo, Dogri, Kashmiri, Konkani, Maithili, Manipuri, Nepali, Sanskrit, Santali, Sindhi, Urdu) exist only as ~1.1–1.6 h of near-field audio — a few hundred speaker turns each — yet §5 ranks them ('Santali ... 9.7% DER and Urdu ... 12.8% DER are among the easiest') and asserts a cross-family effect ('Dravidian languages show near-field cpWER roughly 5 percentage points above Indo-Aryan languages at comparable DER'). The Malayalam near-field row (1.3 h) also feeds the Dravidian average. Point estimates at this volume plausibly carry ±5–15 pp uncertainty. Please add bootstrap confidence intervals (or at minimum per-language session counts and a significance caveat) and soften the family-level claim accordingly. The duration-weighted aggregates in Table 3 are","section":"§5 (Performance across languages) and Table 2"},{"comment":"The text states 'Telugu emerges as the most challenging language, exhibiting the highest overlap (24.7%)', but Table 2 lists Telugu's overlap as 20.4%; 24.7% is Maithili's value (and Maithili has only near-field audio, so its table value equals its near-field value). Either the §5 numbers come from a near-field-only overlap computation not shown, or the value is misattributed. Since the 'Telugu most challenging' claim partly rests on this figure, please reconcile the text with Table 2 and state explicitly which condition each quoted overlap number refers to.","section":"§5 (Performance across languages) vs. Table 2"},{"comment":"The top-ranked system (Sarvam) is a product of the first authors' employer, and its reference [19] is a blog post. This is disclosed via affiliations and is not disqualifying, but the comparison's credibility requires stronger reproducibility commitments than are currently stated: exact model/API versions and access dates for all systems, the prompting/endpoint configuration used for GPT-4o and Gemini 3 Pro, and release of the scoring scripts and system outputs alongside the dataset. Additionally, excluding diarization-only models (e.g., Pyannote) is defensible for the joint task, but a DER-only baseline would contextualize the DER column and cost little; at minimum, the paper should note that Table 3 DERs are not comparable to diarization-only literature for that reason.","section":"§4 (Models) and Table 3"}],"minor_comments":[{"comment":"The heatmaps are difficult to parse at print size: language codes are non-standard abbreviations (Brx, Doi, Kok, Sat, etc.) without a legend, and several cells exceed 100% (e.g., Gemini cpWER 103, 116) which deserves a one-line explanation (insertions on low-word-count languages).","section":"Figure 2"},{"comment":"DER is computed 'without a forgiveness collar and including overlapping speech' — good — but please state the scoring tool/version (e.g., dscore / pyannote.metrics) and confirm the same convention for the Miss/FA/Conf decomposition, since collar conventions differ across prior benchmarks.","section":"§4 (Metrics)"},{"comment":"The 'Overlap %' column is described as 'averaged over all conditions,' but for 12 languages only one condition exists; clarify whether overlap is time-weighted or session-averaged, and how the total row (12.8%) is computed.","section":"§3.3 / Table 2"},{"comment":"Stating that no speaker IDs are released for the in-the-wild subset is appreciated; please also state whether the in-the-wild audio itself is redistributed or released as YouTube URLs + timestamps, since link rot will affect reproducibility.","section":"§3.3 (Limitations)"},{"comment":"Typo: 'Sshubam' likely should be 'Shubam'.","section":"Acknowledgments"},{"comment":"The DISPLACE '24 row lists 38 hours; the text (§2) says 158 hours with 38 labelled. Align the table cell with the labelled-hours figure and note the distinction in the caption.","section":"Table 1"}],"recommendation":"major_revision","confidential_remarks":"The first three authors are Sarvam AI employees and the best-performing system in Table 3 is Sarvam's commercial pipeline. Affiliations are disclosed and the benchmark data itself appears independently collected (with AI4Bharat collaboration), so I do not see grounds for action beyond the reproducibility requests in my major comments (versioned configs, released scoring scripts and system outputs). The editor may wish to verify that the HuggingFace release actually contains the annotations and protocols promised in §1 before final acceptance. Note also the reader's flagged concern about annotation label noise is partly mitigated by the near-field recording design (one mic per non-co-located speaker for ~53 h), but the in-the-wild ~28 h has no such structural safeguard, which is why I am asking for measured inter-annotator agreement rather than accepting the pipeline description at face value."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"This is a resource paper that does the job it claims. Nobody had an open joint diarization + speaker-attributed ASR eval set spanning all 22 scheduled Indian languages; DISPLACE was narrower and decoupled the tracks. They shipped ~108h across near-field, far-field, and in-the-wild, with human-corrected RTTM and dual code-mix transcripts, plus a public HF release. That is the new artifact.\n\nWhat works: collection and annotation are described carefully enough to trust the bulk of the gold (especially near-field, one mic per non-co-located speaker, annotators cannot invent speakers). Metrics are the right community ones (DER no collar, cpWER, WDER). Error decomposition in Table 3 is useful—Gemini’s DER is mostly missed short turns; Sarvam is balanced; Azure is conservative VAD. Overlap vs error correlation is honest. Duration-weighted aggregates are the number to cite.\n\nSoft spots in proportion. Twelve languages are only ~1.1–1.6h near-field. Point claims like “Santali easiest / Telugu hardest” and the Dravidian vs Indo-Aryan 5pp cpWER line rest on thin turn counts with no CIs—the stress-test is right there; the “all 22” framing is true but load-bearing in a thin sense, and Table 3 is dominated by the well-covered languages. No IAA is a real but standard omission for a gold resource, not a reason to discard it. Sarvam topping its own table is COI texture, not circular math; a blunt disclosure line would help. Minor: the ITW “lowest overlap 6.5%” sentence does not line up cleanly with Table 2.\n\nMath and citations are fine for this genre—no invented metrics, prior English/Mandarin/DISPLACE work is properly placed. For anyone building or evaluating multilingual conversational ASR for India this is citable infrastructure. I would send it to referees, ask for IAA, uncertainty on the thin languages, and clearer COI, and engage with the release rather than wait for a perfect second version.","headline":"Real infrastructure gap filled: all-22 Indic joint diarization+ASR labels with a public release and sane baselines—thin hours on 12 languages just mean you should not over-read the per-language rankings.","tokens_in":11220,"tokens_out":553,"would_cite":true,"duration_ms":27988,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Indic DiarBench is the first open joint diarization-and-ASR benchmark spanning all 22 scheduled Indian languages, with roughly 108 hours of human-corrected multi-speaker speech.","keywords":["speaker diarization","Indian languages","multilingual benchmark","speaker-attributed ASR","code-mixing","conversational speech","DER","cpWER"],"falsifier":"Independent re-annotation of a large high-overlap subset that systematically changes speaker boundaries or word sequences and thereby reverses the reported ranking of systems on DER or cpWER would show the released labels are not yet a stable joint benchmark.","tokens_in":10977,"feed_emoji":"🗣️","tokens_out":912,"duration_ms":32242,"temperature":0.7,"pith_summary":"Progress in Indian-language speech recognition has mostly targeted clean single-speaker audio, yet real meetings and conversations involve overlapping talkers, English code-mixing, and wide dialectal range. This paper releases Indic DiarBench: about 108 hours of natural multi-speaker recordings that cover every one of India's 22 scheduled languages, drawn from near-field virtual meetings, far-field rooms, and in-the-wild YouTube conversations. Every file carries human-corrected, time-aligned transcripts that attribute each word to a speaker. The authors evaluate commercial speech APIs and multimodal models on the same mixed single-channel audio and show that even the strongest systems still produce large errors under high overlap and on lower-resource languages. The corpus, labels, and protocols are released openly so the field can measure joint speaker diarization and recognition under conditions that match Indian conversational speech.","feed_headline":"First joint diarization-ASR bench for all 22 Indian languages","feed_subtitle":"108 hours of human-corrected multi-speaker audio show systems still fail on overlap and low-resource speech.","key_machinery":"Indic DiarBench—the multilingual multi-condition corpus together with the joint metrics DER (acoustic segmentation), cpWER, and WDER (speaker-attributed transcription)—is the central object. It forces every system to be scored on identical mixed single-channel audio so diarization mistakes and recognition mistakes are measured together rather than in isolation.","core_discovery":"No prior open resource jointly evaluates speaker diarization and speaker-attributed ASR across all 22 scheduled Indian languages under realistic multi-speaker conditions. Indic DiarBench supplies roughly 108 hours of such audio with human-corrected, time-aligned speaker transcripts, and the accompanying baselines show that current commercial APIs and multimodal models remain far from reliable—especially when speakers overlap or the language is lower-resource.","pith_inferences":["Because multimodal models show high missed-detection rates yet competitive transcription once segments are found, hybrid stacks that keep a strong diarizer in front of an LLM decoder are a natural architecture to test next.","The reported correlation between overlap ratio and error for the best Indic system implies that overlap-aware separation or training will move aggregate scores more than language-specific fine-tuning alone.","Weighting far-field and YouTube subsets more heavily in future leaderboards would better reflect true single-microphone difficulty than near-field virtual meetings alone."],"forward_implications":["Comparable joint diarization-plus-ASR numbers can now be reported on all 22 scheduled Indian languages instead of English-only or single-speaker sets.","Public baselines identify high-overlap segments and lower-resource languages as the dominant remaining failure modes.","Dual native-script and Romanized English reference transcripts allow fairer scoring of code-mixed system output.","Open RTTM and segment-level labels enable development of tightly coupled diarization-ASR pipelines rather than cascaded ones.","The same collection and annotation protocol can be extended to finish in-the-wild coverage for the remaining twelve languages."],"fun_headline_variants":["Indic DiarBench: first joint diarization-ASR set for all 22 Indian languages","108h open bench stresses overlap and low-resource Indian speech systems","Human-corrected multi-speaker audio across India's 22 scheduled languages","Baselines lag on code-mix, overlap in full 22-language diarization-ASR test","Open Indic DiarBench exposes gaps in joint speaker ID and ASR"],"cache_read_input_tokens":128,"weakest_assumption_plain":"The multi-stage human correction pipeline is assumed to yield speaker labels and transcripts accurate enough that differences in system error rates reflect true capability rather than leftover annotation noise, especially on heavily overlapping speech.","fun_headline_variants_meta":{"raw":{"variants":["Indic DiarBench: first joint diarization-ASR set for all 22 Indian languages","108h open bench stresses overlap and low-resource Indian speech systems","Human-corrected multi-speaker audio across India's 22 scheduled languages","Baselines lag on code-mix, overlap in full 22-language diarization-ASR test","Open Indic DiarBench exposes gaps in joint speaker ID and ASR"]},"model":"grok-4.5","effort":"low","cost_usd":0.001623,"raw_usage":{"total_tokens":805,"prompt_tokens":693,"num_sources_used":0,"completion_tokens":92,"cost_in_usd_ticks":16228000,"prompt_tokens_details":{"text_tokens":693,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":20,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":693,"tokens_out":92,"duration_ms":3447,"temperature":1.0,"reasoning_tokens":20,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-30T11:49:42.491803+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Independent re-annotation of a large high-overlap subset that systematically changes speaker boundaries or word sequences and thereby reverses the reported ranking of systems on DER or cpWER would show the released labels are not yet a stable joint benchmark.","supporting_citations":[],"review_version":1}