{"id":"5cfcfe86-286c-4e00-b688-cc511d2799d4","arxiv_id":"2601.04876","paper_version":2,"verdict":"REJECT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":5,"one_line_summary":"A long-audio benchmark claims severe performance collapse in 16 audio LLMs, but unspecified metric weights and internal inconsistencies make the headline numbers uninterpretable.","lead":"ChronosAudio is a new benchmark with 36,000 audio questions meant to test whether AI models can handle long recordings. The authors report that all 16 tested models fail badly on longer audio, but the paper's own scoring formulas are missing key weights and its data totals are inconsistent.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Table 2 scores are not reproducible: metric coefficients (β1,β2,γ1,γ2,w1,w2,δ1,δ2) are never specified, the Appendix substitutes an LLM-judge average for Eqs. (6)–(9), and Summary ground truth is a placeholder—so the >90% collapse and 50% ceiling are unverifiable.","rationale":"I read the paper as a benchmark proposal whose contribution is not the dataset alone but three quantitative findings about ALLM long-context behavior. The load-bearing condition is that the numbers in Table 2 are trustworthy and reproducible. The paper fails this condition in two independent ways: equations define scores with free parameters that are never set, and the appendix describes a different scoring protocol. The Summary ground-truth label 'No specific summary answer of this content' in the case study concretely shows the advertised Eq. (9) metric cannot be applied to at least those examples. I considered other concerns—e.g., the internal duration arithmetic (6,000 clips at the stated length ranges sum to far more than 237 hours), the 'first benchmark' claim conflicting with the cited BLAB, and the attention-dilution evidence being qualitative and single-model—but those are secondary or partly framing issues. The scoring ambiguity is the most load-bearing because it undermines every quantitative headline. This matches the reader's weakest_assumption, so I mark agree. I do not propose changing the REJECT verdict; the concern supports it.","tokens_in":24591,"tokens_out":5406,"duration_ms":57119,"concrete_test":"Run an independent reproducibility audit: take a random 200-example subset per task, obtain raw outputs from 3–4 of the listed models, and compute scores with the released evaluation script. Determine whether the script uses Eqs. (6)–(9) with explicit coefficient values or the LLM-judge average; if coefficients are missing or the judge prompt lacks the task formula, Table 2 cannot be regenerated. Also supply real gold summaries for at least 100 Summary examples and recompute Summary scores; if the scores change materially, the Summary column and the claims built on it are invalid.","verdict_should_be":"UNCHANGED","load_bearing_attack":"All three headline findings—the >90% collapse (Abstract; §4.2 Obs.❶), the attention-dilution explanation (§4.3), and the 50% mitigation ceiling (§4.4)—depend entirely on the scores in Table 2. Those scores are not well-defined. Eqs. (6)–(9) define composite metrics for Transcription, Multi-Speaker, Comprehension, and Summary as weighted combinations of mWER, BERTScore, speaker F1, exact match, hallucination penalty, coverage and factuality, with balancing coefficients β1,β2,γ1,γ2,w1,w2,δ1,δ2. The paper never assigns values to any of these coefficients, and the appendix does not give them. Appendix 3 then describes a different scoring procedure: three LLM judges (DeepSeek-V3.1-Terminus, Qwen3-VL-235B-A22B-Thinking, Kimi-K2-Thinking) each give an integer 1–100 and the final score is the arithmetic mean over the three judges. It is never explained how this judge-based score relates to Eqs. (6)–(9), so Table 2 cannot be reproduced under either reading. The Summary task is particularly problematic: the Appendix case study shows all three Summary examples (Short/Middle/Long) labeled 'No specific summary answer of this content.' With no gold summary, Eq. (9) has no defined G; if a judge was instead comparing against the transcript, that is a different metric than the one advertised. Since §4.2's 'staggering degradation of over 90%' is computed from these uninterpretable numbers (e.g., closed-source Transcription 42.66→3.86), the central empirical claim is not established.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces ChronosAudio, a proposed multi-task benchmark for evaluating audio large language models (ALLMs) on long-form audio. It contains 36,000 test instances over six tasks (Dictation, Localization, Transcription, Multi-Speaker, Comprehension, Summary), stratified into short, middle, and long durations, and reports evaluations of 16 open- and closed-source models. The central claims are: (1) ALLMs show a 'precipitous' performance collapse of over 90% when moving from short to long audio; (2) this collapse is caused by 'Structural Attention Dilution', a diffusion of attention in later sequence positions; and (3) existing mitigation strategies such as sparse attention recover only about 50% of short-context proficiency. The paper also proposes mitigation experiments with sparse and sliding-window attention.","tokens_in":25186,"tokens_out":5047,"duration_ms":55292,"significance":"If the benchmark and its measurements were valid, this would be a useful contribution to an important and underexplored problem: how ALLMs behave on minute-scale and document-scale audio. The paper has strengths: it covers 16 models, spans six task types, stratifies by duration, and releases code/data. The appendix includes raw prompt/response examples and a transparent ensemble-LLM-judge protocol, which is commendable. However, the central empirical findings are not currently supported: the composite metrics in Eqs. (6)–(9) are under-specified, the Summary ground truth appears to be a placeholder, the data volume is internally inconsistent, and the attention-dilution explanation rests on a qualitative visualization of a single model. These issues compromise the interpretability of every headline number in Table 2, and therefore the significance of the paper's conclusions.","major_comments":[{"comment":"The composite scores for Transcription, Multi-Speaker, Comprehension, and Summary are not computable as written. The balancing coefficients β1, β2, γ1, γ2, w1, w2, δ1, δ2 are never assigned values, and no constraints such as β1+β2=1 are stated. Appendix 3 then describes a completely different protocol: three LLM judges each give an integer 1–100 and the final score is the arithmetic mean. The relationship between this judge-based score and Eqs. (6)–(9) is never explained. Since all downstream observations—including the >90% collapse, the attention-dilution claim, and the 50% recovery ceiling—rely on the resulting Table 2 scores, the central empirical claims are not reproducible.","section":"§3.3–3.4, Eqs. (6)–(9)"},{"comment":"The reported scale of the benchmark is internally inconsistent. Section 3.1 states 6,000 distinct audio clips with an average length of 322 seconds. That product is 6,000 × 322 s ≈ 537 hours, but the abstract claims 'over 200 hours' and Table 1 reports '237h'. The paper must reconcile these numbers; as it stands, the reader cannot tell how much audio the benchmark actually contains, which affects every claim about long-form coverage and model workload.","section":"§3.1 and Table 1"},{"comment":"The ground-truth labels for the Summary task are shown as 'No specific summary answer of this content' for the Short, Middle, and Long examples. Equation (9) requires a ground-truth key-point set K(G) to compute coverage and factuality, but no such summary reference is provided. Moreover, the LLM-judge prompt in Figure 6 supplies the 'Audio Transcript (Ground Truth)', not a reference summary. Thus the Summary scores in Table 2—including the open-source 40.96→14.14 and closed-source 75.00→59.22 comparisons—are not valid measurements of summarization quality as defined by Eq. (9).","section":"Appendix, Summary Task case study"},{"comment":"The 'Structural Attention Dilution' explanation is not established. The evidence consists of qualitative attention heatmaps for a single model, Qwen2-Audio-7B, at three 19-token windows marked First, Middle, and Last. No quantitative definition of 'dilution' is given, no entropy or diagonal-concentration metric is reported, no comparison is made to models that do not collapse, and no link is shown between the heatmap pattern and the task scores in Table 2. As presented, this is an illustration, not a causal explanation.","section":"§4.3 and Figure 4"},{"comment":"The '50% recovery' ceiling is based on a single example: Qwen2.5-Omni-3B Transcription, where Sparse Attention gives 25.20 versus the short-form baseline of 50.10. Table 3 reports only three open-source models, and despite the claim in Table 2 that scores are 'averaged over 5 experimental rounds', no variance or significance is reported. Generalizing from one model–task pair to a universal 'restorative ceiling' is unsupported.","section":"§4.4, Obs.❻"}],"minor_comments":[{"comment":"The text 'with a remarkable decline about 18 points, from 3.37 to 21.30 ↑' is garbled: the numbers appear in the wrong order and the arrow direction is inconsistent with a decline. Please correct.","section":"§4.2, Obs.❶"},{"comment":"The symbol I(·) is introduced in §3.4 as an 'information extraction function' but is then used in Eq. (8) as an indicator function. This ambiguity makes the comprehension metric hard to interpret.","section":"Eq. (8)"},{"comment":"The case-study headings for the Summary task are mislabeled as 'Comprehension Task (Middle)' and 'Comprehension Task (Long)', even though the prompts and labels are for summarization. This is confusing and should be fixed.","section":"Appendix, Summary case study headers"},{"comment":"References 'Li et al. 2025a' and 'Li et al. 2025b' appear to be the same paper (ISA-Bench). Please merge or distinguish them.","section":"References"},{"comment":"The formatting of Table 1 is broken: benchmark names are concatenated with durations (e.g., '400h14s'), making the comparison difficult to parse. A clean table is needed.","section":"Table 1"}],"recommendation":"reject","confidential_remarks":"The manuscript has a useful idea and a great deal of evaluation work, but the current form does not support its headline claims. The missing coefficients, the placeholder Summary labels, and the unvalidated judge protocol are not presentation issues—they make Table 2 uninterpretable. The attention and mitigation analyses are also too thin. I would not invite a resubmission without a complete re-specification of the metrics, corrected data accounting, and re-analysis of the central results. The overlap of the author list with a cited reference (Lin et al., 2025) is not itself a concern, but the naming and citation should be transparent."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear [Colleague],\n\nThe short version: this paper is a serious attempt at a long-audio benchmark, and the authors clearly put in the hours — 16 models, six tasks, three duration buckets, qualitative failure-mode case studies. The task taxonomy (perception, verbatim generation, reasoning) is sensible, and the qualitative examples of models refusing or looping on long inputs are genuinely illustrative. But the central empirical claims — the 90% collapse, the attention-dilution explanation, and the 50% mitigation ceiling — rest on Table 2, and Table 2 is not reproducible.\n\nThe problems are load-bearing, not cosmetic. Equations (6)–(9) define composite scores with balancing coefficients β1, β2, γ1, γ2, w1, w2, δ1, δ2, and the paper never assigns values to any of them. In the appendix, the authors describe a different procedure: three LLM judges score each answer 1–100 and the final score is their average. No bridge between those two protocols is given. You cannot tell what the numbers in Table 2 mean. The Summary task is worse: the case study labels three ground truths as 'No specific summary answer of this content.' There is no G to plug into Eq. (9). If the judge is comparing against the full transcript, that's a different metric than the one advertised. The limitations section mentions English-centric data and the lack of training-based fixes, but it doesn't acknowledge the metric problem.\n\nThe dataset statistics also don't add up. 6,000 clips × 322 seconds ≈ 537 hours, not the 'over 200 hours' in the abstract or the 237h in Table 1. And the 'first long-audio benchmark' claim is directly contradicted by BLAB, which appears in their own Table 1 with an 833h, 51-minute average.\n\nWhat's genuinely new is the benchmark artifact itself — 36,000 instances across six task types stratified by duration. If the authors fix the metrics, release a data card, correct the volume accounting, and recalibrate their novelty claims, this could be a useful resource. As it stands, the headline findings are unverifiable and the paper needs major revision before it is referee-ready. I would not cite it in its current form.","headline":"A serious long-audio benchmark effort undone by undefined metrics, inconsistent data totals, and placeholder summary ground truth — the headline findings are not reproducible.","tokens_in":25608,"tokens_out":6496,"would_cite":false,"duration_ms":60679,"reading_group":"maybe","serious_thinker":"no","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"ChronosAudio, a 36,000-question benchmark over 200+ hours of audio, measures a >90% performance collapse in audio LLMs when clips stretch from seconds to 10–20 minutes, traces it to attention dilution, and finds mitigations recover only hal","keywords":["long-audio benchmark","audio large language models","length generalization","attention dilution","lost in the middle","long-context evaluation","speech transcription","LLM-as-a-judge"],"falsifier":"Re-score a stratified sample of ChronosAudio instances with concretely specified coefficients and human-verified references for the Summary task, then re-measure the short-to-long degradation; if the >90% collapse shrinks to a modest drop (or vanishes) on that sample, the paper's headline finding fails. A second check: compute a diagonal-dominance index of attention maps across sequence positions; if the index does not decline with position in models that nonetheless show the collapse, the attention-dilution mechanism is not the cause.","tokens_in":24535,"feed_emoji":"🎧","tokens_out":4729,"duration_ms":50610,"temperature":0.7,"pith_summary":"This paper builds ChronosAudio, a benchmark of 36,000 test questions over 200 hours of audio, organized into six tasks and three duration bands (short, middle, long), to measure whether audio large language models understand anything beyond short clips. The central claim is that they do not: moving from short clips to 10–20 minute recordings cuts performance by over 90% on some tasks, with most of the loss already visible at 5–10 minutes. The paper attributes this collapse to structural attention dilution—attention maps lose their sharp local focus as the sequence progresses—and shows that sparse or sliding-window attention recovers at most about 50% of short-form skill. A sympathetic reader would care because long-form audio (meetings, calls, lectures) is exactly where these models would be useful, and the benchmark makes the failure measurable and comparable.","feed_headline":"Audio LLMs lose 90% of skill on long clips","feed_subtitle":"36,000-question benchmark over 200 hours of audio finds attention dilution and a 50% repair ceiling.","key_machinery":"The load-bearing instrument is ChronosAudio itself: 6,000 distinct audio clips (3,000 short at 30s–5min, 2,000 middle at 5–10min, 1,000 long at 10–20min) rendered into 36,000 test instances across six tasks (Dictation, Localization, Transcription, Multi-Speaker, Comprehension, Summary). The three findings are produced by length-stratified scoring on those tasks, visualization of self-attention weights across First/Middle/Last sequence segments to reveal structural attention dilution, and surgical modification of attention (Sparse Attention and Sliding Window Attention) to test the restorative ceiling. The named mechanism, structural attention dilution, is the loss of a sharp diagonal attenti","core_discovery":"The discovery the paper argues for is that current audio large language models, both open- and closed-source, have no reliable long-audio understanding. On ChronosAudio, open-source models fall from an average of 28.06 to 0.00 on long-form transcription, and closed-source models fall from 42.66 to 3.86; comprehension for open-source models collapses to 3.90. The paper further argues that the mechanism is structural attention dilution, visible as the loss of the diagonal attention pattern in later sequence positions, and that mitigation by sparse attention or sliding windows is capped at roughly 50% of the short-context score for fidelity-heavy tasks. The authors present this as evidence that","pith_inferences":["Because the metric coefficients in Equations (6)–(9) are never assigned, and the Summary ground truth in the appendix reads 'No specific summary answer of this content,' a re-scoring with explicit coefficients and human-validated summaries is a necessary check before the quantitative scale of the collapse is taken at face value.","If attention dilution is causal rather than merely correlational, then interventions that force local attention—chunk-wise cross-attention, audio-native positional encodings, or token downsampling—should be compared against the 50% ceiling the paper reports; a method that breaks the ceiling would validate the mechanism.","The benchmark's duration stratification enables a simple testable prediction: model performance on any fidelity task should be a monotone decreasing function of duration, so plotting score versus duration on a held-out sample would let the community check the collapse curve outside the 16 models tested.","The paper's English-only, clean-audio scope leaves open whether the collapse is a general audio-length effect or partly an artifact of the test distribution; extending the same protocol to noisy and multilingual long audio would separate those factors."],"forward_implications":["Long-form dictation, localization, transcription, and multi-speaker tasks are effectively unsolved: most models score near zero on 10–20 minute audio.","The lost-in-the-middle effect starts earlier in audio LLMs than in text LLMs: the steepest drop is from short to middle duration, suggesting the effective high-fidelity context window is under 5 minutes.","Attention dilution gives a concrete target for future work: models need mechanisms that preserve temporal locality across long sequences, not just longer context windows.","Sparse attention can nearly eliminate the drop on retrieval-oriented tasks (93% recovery in dictation) but leaves transcription at about 50% of short-form proficiency, so retrieval and verbatim fidelity need different fixes.","Closed-source models retain substantial high-level reasoning on long audio (Summary around 59) while open-source models collapse to about 14, indicating a gap between perception and reasoning abilities."],"fun_headline_variants":["Audio LLMs lose 90% accuracy on long clips despite fixes","ChronosAudio benchmark exposes 90% long-audio collapse","Long audio crashes ALLMs: attention dilutes, fixes cap at 50%","Why audio LLMs fail on long clips: attention drift and 50% ceiling","New benchmark shows audio LLMs can't handle long audio"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The numbers carry the argument: every composite score in Equations (6)–(9) depends on balancing coefficients (β1, β2, γ1, γ2, w1, w2, δ1, δ2) that the paper never assigns, and the Summary ground truth in the appendix is the placeholder 'No specific summary answer of this content'; if those coefficients are arbitrary or those labels are invalid, the 90% collapse, attention-dilution, and 50%-ceiling findings are uninterpretable.","fun_headline_variants_meta":{"raw":{"variants":["Audio LLMs lose 90% accuracy on long clips despite fixes","ChronosAudio benchmark exposes 90% long-audio collapse","Long audio crashes ALLMs: attention dilutes, fixes cap at 50%","Why audio LLMs fail on long clips: attention drift and 50% ceiling","New benchmark shows audio LLMs can't handle long audio"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000667,"raw_usage":{"total_tokens":2889,"prompt_tokens":760,"completion_tokens":2129,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":504,"completion_tokens_details":{"reasoning_tokens":2034}},"tokens_in":504,"tokens_out":2129,"duration_ms":14969,"temperature":1.0,"reasoning_tokens":2034,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T11:50:27.984798+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-score a stratified sample of ChronosAudio instances with concretely specified coefficients and human-verified references for the Summary task, then re-measure the short-to-long degradation; if the >90% collapse shrinks to a modest drop (or vanishes) on that sample, the paper's headline finding fails. A second check: compute a diagonal-dominance index of attention maps across sequence positions; if the index does not decline with position in models that nonetheless show the collapse, the attention-dilution mechanism is not the cause.","supporting_citations":[],"review_version":1}