{"id":"6c348d27-fbdb-4e07-a101-5cba30e9da64","arxiv_id":"2508.05055","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":1,"one_line_summary":"MOVER is a five-stage combination method that fuses diarization and ASR outputs from multiple meeting recognition systems, improving tcpWER by 9.55% and 8.51% relative on CHiME-8 DASR and NOTSOFAR-1.","lead":"Meeting transcription systems turn meeting audio into labeled transcripts of who said what. MOVER merges the outputs of several such systems into one combined transcript, and reports roughly 8.5 to 9.5 percent fewer errors on two public benchmarks.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Central claim is uncheckable: the supplied full text is unreadable mojibake and carries a hep-ex header for a different arXiv ID, leaving only the abstract's self-reported tcpWER gains as evidence.","rationale":"The reader identified speaker-alignment and segment-grouping reliability as the weakest assumption. I agree that this is the key internal assumption on which the central claim depends: if alignment errors are correlated with the individual systems' errors, the combination gain can disappear. However, the more immediate load-bearing obstacle is that the supplied full text is unreadable and carries a header for a different arXiv ID, so even the existence of the alignment equations, the experimental setup, and the actual per-system numbers cannot be checked. This is a verifiability concern rather than a demonstrated correctness flaw, so the correct verdict is the same as the reader's: UNVERDICTED. The verdict should not move to ACCEPT or CONDITIONAL because the evidence is insufficient, and it should not move to REJECT because no internal inconsistency has been established. The header mismatch is flagged as an in-scope textual observation, not as evidence of author misconduct.","tokens_in":15550,"tokens_out":2522,"duration_ms":30277,"concrete_test":"Obtain the clean PDF or source of arXiv:2508.05055 directly from arXiv and re-extract the text. If the full text remains unreadable, the central claim is unverifiable and the verdict should remain UNVERDICTED. If it is readable, locate the speaker-alignment and segment-grouping sections, then rerun the five-stage pipeline on the official CHiME-8 and NOTSOFAR-1 evaluation outputs using the released system hypotheses, and recompute tcpWER against the reported 9.55%/8.51% improvements. A focused sub-check: measure whether alignment failures are correlated with ASR/diarization errors by counting cases where the combined output is worse than the best input system; if such cases are systematic, the headline gain would not generalize.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is an empirical one: a five-stage pipeline, MOVER, fuses hypotheses that differ in both time boundaries and speaker labels, and thereby beats every individual system by 9.55% and 8.51% relative tcpWER on CHiME-8 DASR and NOTSOFAR-1 multi-channel. For this claim to hold, the speaker-alignment and segment-grouping stages must relate hypotheses across different time intervals and speaker labels without introducing errors that are correlated with the systems' errors. If alignment merges segments from different speakers, or if grouping locks in a wrong speaker mapping, the combination can degrade below the best single system. The abstract offers no analysis of such failure modes, no sensitivity analysis, and no diversity measure across the fused systems. More fundamentally, the submitted full text is not parseable: it is mojibake, and the visible running header reads 'arXiv:2508.05058v1 [hep-ex] 7 Aug 2025', i.e., a different paper. Thus the five-stage method, the equations, the baseline descriptions, the per-system results, and the experimental configurations are all unavailable for inspection. The two headline improvements are plausible, but they currently rest entirely on an uncheckable summary. This is a verifiability failure, not an accusation of misconduct, and it blocks any method-level assessment of the alignment assumption.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes MOVER (Meeting recognizer Output Voting Error Reduction), a five-stage system combination method for meeting recognition that fuses hypotheses that differ in both time intervals and speaker labels, combining diarization and ASR outputs jointly. The abstract reports relative tcpWER improvements of 9.55% and 8.51% over state-of-the-art systems on the CHiME-8 DASR task and the multi-channel track of the NOTSOFAR-1 task, respectively. However, the submitted full text is unreadable mojibake with a running header that identifies a different arXiv paper (hep-ex 2508.05058), so the five-stage algorithm, equations, experimental configurations, baselines, and per-system results cannot be inspected. The central empirical claim is therefore supported only by the abstract's self-reported numbers.","tokens_in":15824,"tokens_out":5213,"duration_ms":54806,"significance":"If the reported gains are genuine, MOVER would be a meaningful contribution, as it is claimed to be the first combination method that accommodates systems differing simultaneously in both diarization and ASR hypotheses. The choice of tcpWER on two official public benchmarks is appropriate and provides a standard, comparable metric. At the same time, the significance cannot be evaluated from the submitted materials: no method description, no experimental details, no code, and no supplementary reproducibility artifacts are visible. The abstract-level claims are plausible but entirely unverifiable in the current form.","major_comments":[{"comment":"The supplied full text is not readable: it consists of mojibake, and the running header reads 'arXiv:2508.05058v1 [hep-ex] 7 Aug 2025,' which is not the identifier of this paper. As a result, the five-stage algorithm, all equations, the experimental setup, the baseline descriptions, the per-system results, and the tables cannot be inspected. The central empirical claim therefore rests entirely on the abstract. This is a verifiability failure that blocks any method-level assessment; it must be fixed before the paper can be evaluated.","section":"Full text (entire submission)"},{"comment":"The headline numbers (relative tcpWER improvements of 9.55% and 8.51%) are reported without statistical support: no error bars, confidence intervals, significance tests, or per-system breakdowns. For a combination method, one needs to see the tcpWER of each individual system and the fused output to confirm that the combination beats every component. Please provide a full results table with individual system scores and some measure of variance (e.g., bootstrap confidence intervals).","section":"Abstract"},{"comment":"The component systems and the 'state-of-the-art' baseline are not identified. If one of the combined systems is also the SOTA baseline, then the relative improvement may be partly circular because the baseline contributes to the fused output. Please disclose the identities of all combined systems and whether the baseline is among them, and discuss the degree of output diversity across systems.","section":"Abstract"},{"comment":"The five-stage pipeline involves speaker alignment, segment grouping, and word/timing fusion, which inevitably contain thresholds and voting weights. No sensitivity analysis or ablation of the five stages is visible. The abstract gives no indication that the method is robust to parameter choices. Please provide an ablation study (e.g., leave-one-stage-out) and a sensitivity analysis of the key thresholds.","section":"Method (unreadable)"}],"minor_comments":[{"comment":"The acronym tcpWER is used without definition at first use; please expand (e.g., 'total concatenated-permutation Word Error Rate') for readers outside the meeting-recognition subfield.","section":"Abstract"},{"comment":"The claim of being the 'first approach' is strong and should be carefully positioned against prior system-combination work (e.g., ROVER, DOVER, DOVER-Lap, and other recent fusion methods).","section":"Introduction (wherever visible)"},{"comment":"The arXiv identifier in the running header does not match this submission. Please ensure the correct PDF is uploaded; the current material appears to be a corrupted or mislabeled file.","section":"Full text"}],"recommendation":"major_revision","confidential_remarks":"The provided PDF is unreadable and carries a header from a different arXiv paper (hep-ex 2508.05058). This is not a content issue but a file-integrity problem. I recommend that the editor ask the authors to resubmit a correct, readable PDF before the manuscript is considered further. Even after a clean submission, the abstract-level claims lack statistical support and system-disclosure details; the next version should include full experimental tables, ablations, and sensitivity analyses."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the abstract makes a genuinely useful claim—first system-combination method for meeting recognition where outputs differ in both diarization and ASR, with relative tcpWER gains on CHiME-8 and NOTSOFAR-1. If true, that's a contribution worth knowing about. But the full text is a corrupted download: unreadable mojibake, and the running header reads arXiv:2508.05058v1 [hep-ex], not this paper. So none of the method, baselines, ablations, or even the equations are available for inspection. That's a verifiability failure, not evidence of misconduct.\n\nWhat's creditable: the abstract is specific—two official benchmarks, a standard metric, and numbers that beat SOTA. The extension of ROVER (word fusion) and DOVER (diarization fusion) to fuse both axes jointly is a natural next step and, as far as the abstract truthfully says, new. But that's all we have.\n\nWhere it's soft: the central technical risk—speaker alignment and segment grouping can relate hypotheses with different time boundaries and speaker labels only if errors are not correlated across systems—is real, but we can't even begin to assess it without the method. No error bars, significance tests, sensitivity analysis, or diversity metrics are provided in the abstract. And we can't rule out overlap between the fused systems and the SOTA baseline, so the comparison could be circular; not assessable.\n\nBottom line: this paper deserves serious referee time once it exists in a readable form. As submitted, the right editorial action is a desk reject with an invitation to resubmit a corrected PDF, ideally with code or data. I'd not send this version to referees, and I'd not cite it yet.","headline":"The idea is the right kind of idea, but the submission is unreadable: mojibake full text and a hep-ex header, so the claimed gains are unverifiable.","tokens_in":16376,"tokens_out":3052,"would_cite":false,"duration_ms":32560,"reading_group":"no","serious_thinker":"unclear","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"MOVER is the first approach to combine meeting recognition outputs that differ in both diarization and ASR, improving tcpWER by 9.55% and 8.51% over state-of-the-art single systems on two benchmarks.","keywords":["MOVER","meeting recognition","system combination","diarization","automatic speech recognition","tcpWER","CHiME-8 DASR","NOTSOFAR-1"],"falsifier":"Run MOVER on a development set in which one system's speaker labels are deliberately permuted on every recording; if the aligned output does not beat the better single system's tcpWER, the combination claim fails.","tokens_in":15407,"feed_emoji":"🎙️","tokens_out":5743,"duration_ms":62705,"temperature":0.7,"pith_summary":"MOVER (Meeting recognizer Output Voting Error Reduction) is a five-stage method for combining complete meeting recognition hypotheses — diarization plus transcripts — rather than combining only speaker labels or only words. Earlier combination tools DOVER and ROVER each handle one side of the problem; MOVER is positioned as the first to handle both at once, so the systems being combined can disagree on who spoke, when, and what was said. Tested on the CHiME-8 DASR task and the multi-channel track of NOTSOFAR-1, MOVER combines multiple diverse systems and improves the official tcpWER metric by 9.55% and 8.51% relative to the state-of-the-art single systems. If the result is correct, it gives meeting recognition a drop-in way to turn a collection of independent systems into a stronger one without retraining any of them.","feed_headline":"MOVER combines meeting recognizers, cutting errors up to 9.55%","feed_subtitle":"The first method to merge outputs that disagree on both speaker IDs and timing beats the best single systems on CHiME-8 and NOTSOFAR-1.","key_machinery":"The carrying mechanism is MOVER's five-stage combination pipeline: speaker alignment maps each system's speaker labels onto a common set of speakers; segment grouping finds which time intervals from different systems correspond to the same speech; word and timing combination merges the hypotheses into one word sequence with one timing. Earlier methods DOVER and ROVER combine only diarization outputs or only ASR outputs respectively; MOVER's pipeline is what lets both be combined at once, so errors that differ across systems can cancel by voting.","core_discovery":"The central claim is that hypotheses from meeting recognition systems that differ in both diarization and ASR can be combined into one hypothesis that beats every individual system on tcpWER. The method aligns speakers across outputs, groups the corresponding segments, and combines words and timings by voting, producing a single consensus transcript with one set of speaker labels and time boundaries. On the two evaluated tasks, the combined hypothesis achieves relative tcpWER improvements of 9.55% (CHiME-8 DASR) and 8.51% (NOTSOFAR-1 multi-channel) over the best individual systems.","pith_inferences":["If the gain mechanism is error diversity, MOVER should improve more as the combined systems become more different; combining near-identical systems should add little or nothing. This can be tested directly by controlling the overlap between system outputs.","The same alignment-plus-voting structure could be lifted to other output formats that carry speaker labels and time intervals, such as multimodal diarization that uses video or motion cues.","A finer analysis separating alignment errors from word errors would show where the remaining tcpWER comes from and whether confidence-weighted voting pushes the gains further."],"forward_implications":["Combining several meeting recognizers with MOVER yields a single output that scores better on tcpWER than any of the individual systems.","The combination works without retraining or modifying the component systems, so it can be applied to off-the-shelf recognizers.","The method accepts systems with different speaker labelings and segment boundaries, removing a prior restriction in system-combination work.","The gains replicate across two distinct tasks, showing the approach is not tuned to a single benchmark."],"supporting_citations":[],"fun_headline_variants":["MOVER fuses diverse meeting systems, trimming tcpWER by 9.55%","Combine meeting recognizers with differing speakers and timings via MOVER","MOVER: first to merge diarization and ASR outputs, improving accuracy","MOVER unifies multiple meeting systems for lower word error","Combine multiple meeting recognizers to cut tcpWER by 9.55%"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"MOVER's gains depend on its speaker-alignment and segment-grouping stages reliably matching hypotheses that disagree in both time boundaries and speaker labels, so that the combination removes more errors than the alignment process introduces.","fun_headline_variants_meta":{"raw":{"variants":["MOVER fuses diverse meeting systems, trimming tcpWER by 9.55%","Combine meeting recognizers with differing speakers and timings via MOVER","MOVER: first to merge diarization and ASR outputs, improving accuracy","MOVER unifies multiple meeting systems for lower word error","Combine multiple meeting recognizers to cut tcpWER by 9.55%"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000685,"raw_usage":{"total_tokens":2919,"prompt_tokens":691,"completion_tokens":2228,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":435,"completion_tokens_details":{"reasoning_tokens":2140}},"tokens_in":435,"tokens_out":2228,"duration_ms":16195,"temperature":1.0,"reasoning_tokens":2140,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T23:33:51.289295+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run MOVER on a development set in which one system's speaker labels are deliberately permuted on every recording; if the aligned output does not beat the better single system's tcpWER, the combination claim fails.","supporting_citations":[],"review_version":1}