{"id":"5095a8d1-61b4-4223-a8c3-ffe22ba7841e","arxiv_id":"2607.09880","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"CLIR-Bench shows generalist and time-series LLMs struggle to ground clinical answers in sparse irregular ICU evidence, with top accuracy near 50% and weak causal evidence use.","lead":"CLIR-Bench is a 6,600-question ICU benchmark that tests whether AI models can answer clinical questions from sparse, irregular vital-sign and lab streams, not just from medical knowledge. It matters because safety-critical clinical QA needs answers that are grounded in the right timestamps, and current models largely fail that test.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified beyond the reader's already-flagged gold-label clinical validity risk.","rationale":"I read the strongest claim as an empirical diagnostic about model behavior on this evidence-auditable benchmark, not as a claim that every gold label is prospectively clinically optimal. Table 2, Figs. 5–8, and Findings 1–5 cohere: models often answer without locating gold evidence, do not flip under causal edits, and are unstable under order/timestamp perturbations. Those results are strongest on deterministic temporal tasks with rule-based labels, which the construction pipeline (task schemas → rule-derived answers → cross-check → human audit) is designed to support. The softest clinical operationalization is indeed IID/MED (retrospective controlled labels, §3.2.2), plus missing IAA and error bars—exactly the reader's weakest assumption. That justifies CONDITIONAL rather than ACCEPT, but does not supply a stronger internal inconsistency or a reason to reject the resource or the model-failure narrative. No separate load-bearing attack is needed; the concrete re-derivation check above is the right verification step and, if passed, leaves the reader's verdict intact.","tokens_in":20301,"tokens_out":552,"duration_ms":14719,"concrete_test":"Independently re-derive gold answers and evidence spans for a stratified sample of 50 non-decision items (e.g., 10 each from TG, ASR, TPR, CVR, IR) from raw MIMIC-IV stays using only the stated evidence rules in §3.3/A.3; if >10% disagree with released labels, the faithfulness/causal claims weaken; if agreement is high, the central empirical claim stands under the existing CONDITIONAL verdict.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim—that current generalist and time-series LLMs fail to reliably retrieve and reason over sparse asynchronous ICU evidence—is supported by multiple independent diagnostics (Full-TS vs QA-only, evidence-only/removed, causal edits, irregularity stress tests) that do not all collapse if a subset of decision-task labels are imperfect. Deterministic understanding/reasoning tasks (TG, ASR, TPR, MA, CVR, IR) have explicit timestamp rules and case studies that make the low faithful accuracy and ~1% causal flip rates hard to dismiss as pure pipeline artifacts. The reader's weakest assumption (retrospective decision labels for IID/MED and unreported IAA) is real and correctly limits ACCEPT, but it is not a new load-bearing crack that overturns Findings 1–4 or Table 2 for the benchmark as a whole.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper introduces CLIR-Bench, a 6,600-instance multiple-choice QA benchmark for irregular, sparse, asynchronous ICU time series built from de-identified MIMIC-IV stays. Through a four-stage pipeline (curation, task instantiation with deterministic evidence rules, LLM-assisted QA generation, human verification), it defines 11 tasks across understanding, reasoning, forecasting, and decision-making, each linked to timestamp-level evidence. Experiments under Full-TS, QA-only, Evidence-only, and Evidence-removed inputs, plus evidence-faithfulness, counterfactual causal edits, and irregularity stress tests, report that current generalist and time-series LLMs remain weak: best macro accuracy is 50.15% (GPT-5.4 mini), faithful accuracy lags answer accuracy, and causal evidence edits flip answers at rates below ~1.2%. The authors conclude that stronger evidence selection and native irregular time-series reasoning are needed.","tokens_in":20590,"tokens_out":1667,"duration_ms":21949,"significance":"If the gold labels and evidence spans are clinically and temporally well-specified, CLIR-Bench fills a clear gap: existing time-series QA suites largely assume regular sampling, while medical QA rarely stresses sparse asynchronous trajectories. The evidence-auditable design (rule-derived answers, Evidence-only/removed ablations, causal vs irrelevant edits) is a genuine methodological contribution and goes beyond accuracy-only leaderboards. Public data and code further strengthen reuse. The multi-diagnostic results (Table 2; Figs. 5–8) make a credible case that current models do not reliably ground answers in sparse ICU evidence, which is useful for both clinical AI and time-series LLM research. Credit is due for the structured task taxonomy, explicit cutoff handling for forecasting, and the stress-test suite that separates shortcut relief from robustness.","major_comments":[{"comment":"Sections 3.2.2 and 3.3 (and Appendix A.3): Immediate Intervention Decision (IID) and Monitoring/Escalation Decision (MED) are framed as decision-making evaluation, but the manuscript states they use retrospective clinical decision labels in a controlled setting rather than prospective clinician actions. This is load-bearing for the claim that the benchmark tests clinical decision-making. Please (i) specify exactly how gold intervention/escalation labels are derived from MIMIC-IV events, (ii) separate or reweight these two tasks when reporting the overall macro average if labels are proxy labels, and (iii) discuss the risk that models are scored for matching historical chart patterns rather than clinically justified decisions. Without this, Findings on Decision-Making and the 11-task overall score overstate clinical decision validity.","section":"§3.2.2 Temporal decision-making; §3.3; Appendix A.3"},{"comment":"Section 3.3.3 and Appendix A.3 describe human verification of task suitability, option quality, evidence support, and label consistency, but report no inter-annotator agreement, acceptance/rejection rates, number of reviewers, or adjudication protocol. For an evidence-auditable benchmark whose central claim rests on gold evidence spans and rule-derived answers, IAA (or at least dual-review rates and disagreement examples) is necessary. Please add quantitative audit statistics and, if dual review was not done for all 6,600 items, the sampling fraction and how disagreements were resolved.","section":"§3.3.3 Human Verification; Appendix A.3"},{"comment":"Table 2 and §4.1–4.2: several reported accuracies (e.g., 41.67, 43.33, 58.33) are consistent with very small per-task evaluation sets (on the order of n≈60), while the abstract advertises 6,600 instances. The evaluation protocol does not clearly state whether Full-TS results use the full benchmark, a fixed subset, or model-dependent subsets, nor does it report confidence intervals or significance tests for macro averages and TS Lift. Please state exact n per task/model, how the eval split was chosen, and add uncertainty estimates so that the 50.15% best-score claim and cross-model rankings can be interpreted.","section":"Table 2; §4.1 Experimental Setup; §4.2 RQ1"},{"comment":"§3.3.2 and Appendix A.3: candidate questions are drafted with Qwen3.6-27B and rewritten with GPT-5.5, while Qwen3.6-27B (and related Qwen models) appear in the evaluated model pool (Table 2). This creates a construction–evaluation contamination risk for those families (shared phrasing priors, option style). Please either (i) exclude generator models from the main ranking, (ii) regenerate a held-out rewrite with a disjoint model and re-score, or (iii) quantify style/overlap bias. Also clarify whether gold answers ever depend on LLM judgment rather than the deterministic evidence rule alone.","section":"§3.3.2 QA Generation; Appendix A.3; Table 2"}],"minor_comments":[{"comment":"Title and abstract call the benchmark “multimodal,” but the inputs are serialized irregular time series plus text questions/options (no imaging or other modalities). Consider “time-series–language” or define multimodality explicitly in §1/§3.1.","section":"Title; Abstract; §3.1"},{"comment":"Table 1 comparison is useful, but several cited benchmarks have very different scopes; a short column on clinical vs general domain and on evidence-audit support would make the novelty claim sharper.","section":"Table 1"},{"comment":"Figure 1 task examples are helpful; ensure every abbreviation in Table 2 (TG, ASR, TPR, MA, TSS, TF, NIF, CVR, IR, IID, MED) is defined once in a single glossary near §3.2.2 for readers skimming results.","section":"Figure 1; Table 2; §3.2.2"},{"comment":"ACM reference block still has placeholder conference metadata (“Conference acronym ’XX’, Woodstock, NY, 2018”). Clean for camera-ready.","section":"Front matter / ACM Reference Format"},{"comment":"§4.6 latency comparison mixes backends and is correctly caveated as wall-clock; still, state hardware and batch size so the 41× claim is not over-read as pure model efficiency.","section":"§4.6 RQ6; Figure 9"},{"comment":"Appendix A.1 ethics note is appropriate; add a one-sentence limitation that MIMIC-IV demographics and practice patterns may not transfer to other ICUs when discussing decision tasks.","section":"Appendix A.1; Conclusion"}],"recommendation":"major_revision","confidential_remarks":"The central empirical story (poor evidence grounding on irregular ICU series) looks real and is backed by several independent diagnostics that survive imperfect decision labels. The main risk for the journal is overselling clinical decision-making validity and under-documenting audit/eval n. If the authors tighten IID/MED framing, add IAA and eval-size/CIs, and address generator–evaluatee overlap, this is a solid benchmark contribution. Model names (GPT-5.4 mini, DeepSeek-V4-flash, etc.) are unusual relative to today’s public catalog but consistent with a 2026 arXiv stamp; no action needed beyond ensuring reproducibility of API versions."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"This is a solid benchmark paper, not a methods breakthrough. What is new is CLIR-Bench: 6,600 multiple-choice QA items over sparse, asynchronous MIMIC-IV ICU trajectories, with timestamp-level evidence metadata and rule-derived gold answers across 11 tasks (understanding, reasoning, forecasting, decision). Table 1 is fair; prior time-series QA suites mostly assume regular grids and do not stress missingness or evidence faithfulness the way this one does.\n\nThey do the evaluation right. Full-TS vs QA-only, Evidence-only vs Evidence-removed, faithful accuracy (answer correct and evidence F1 ≥ 0.5), causal vs irrelevant edits, and irregularity stress tests (order shuffle, timestamp removal, etc.) all point the same way: best macro accuracy is only ~50% (GPT-5.4 mini), faithful accuracy is much lower, and causal flips stay under ~1.2%. That pattern is hard to dismiss as pure construction artifact on the deterministic tasks (TG, ASR, TPR, MA, CVR, IR). Code and data are released. Citations cover the right prior benches and irregular clinical series work.\n\nSoft spots, in proportion: the weakest link is gold-label clinical meaning for Immediate Intervention Decision and Monitoring/Escalation Decision, which use retrospective labels in a controlled setting rather than prospective clinician actions. Unreported inter-annotator agreement and missing error bars are real but secondary. Construction uses LLM drafting plus human audit; mild circularity risk exists, but the answer rules are cross-checked against timestamps, so it is not load-bearing for Findings 1–4. Cohort is 500 stays; fine for a first release.\n\nWho it is for: people building clinical time-series LLMs or evidence-selection methods. Not a general medical-AI paper. I would bring it to reading group if the group cares about temporal grounding or ICU data. It deserves peer review; a serious editor should send it out, with pressure to tighten decision-task justification, report IAA, and add uncertainty. I would cite the resource and the diagnostic protocol if I work on irregular clinical QA in the next year.","headline":"Useful evidence-auditable ICU irregular-series QA benchmark; models really do fail at sparse temporal grounding, with the main soft spot being clinical validity of the decision-task labels.","tokens_in":21205,"tokens_out":527,"would_cite":true,"duration_ms":5592,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Current models cannot reliably ground clinical answers in sparse, irregular ICU time series.","keywords":["irregular clinical time series","time series question answering","large language models","evidence faithfulness","ICU monitoring","multimodal QA","temporal reasoning"],"falsifier":"A model that, under full irregular ICU trajectories, simultaneously reaches high answer accuracy, high evidence F1, high causal flip rate when supporting observations are edited, and stable accuracy under evidence-preserving missingness and order perturbations would falsify the claim that current approaches fail at this form of reasoning.","tokens_in":21176,"feed_emoji":"⏱️","tokens_out":793,"duration_ms":7266,"temperature":0.7,"pith_summary":"CLIR-Bench is a 6,600-question benchmark built from de-identified ICU records to test whether models can answer clinical questions by finding and using sparse, irregular, asynchronous measurements rather than general medical knowledge. The authors construct each item with explicit timestamped evidence and a deterministic answer rule, then organize the set into four capability dimensions and eleven tasks spanning understanding, reasoning, forecasting, and decision-making. Across closed-source, open-source, and time-series LLMs, the best macro accuracy under full time-series input is only about 50 percent, faithful accuracy is far lower, full trajectories often fail to help or even hurt, and editing the causal evidence almost never flips the answer. The paper therefore claims that existing generalist models do not yet perform evidence-grounded reasoning over irregular clinical time series and that stronger selection and native irregular modeling are required.","feed_headline":"Models fail to ground ICU answers in sparse time series","feed_subtitle":"Best score is 50 percent; correct answers rarely track the timestamps that determine them","key_machinery":"Evidence-auditable QA construction: each of the 6,600 multiple-choice items is tied to explicit temporal evidence and a task-specific deterministic answer rule, enabling accuracy, faithfulness, sufficiency, necessity, and counterfactual edit diagnostics.","core_discovery":"Existing generalist and time-series language models cannot reliably retrieve and reason over sparse, asynchronous ICU evidence for clinical question answering: under full irregular context the strongest model reaches only 50.15 percent macro accuracy, correct answers frequently lack matching evidence, and causal edits to the supporting observations flip answers at rates below roughly 1.2 percent across model families.","pith_inferences":["Benchmarks that keep gold evidence editable will become the default way to audit whether medical LLMs are using patient data or clinical priors.","The same evidence-auditable pipeline could be applied to other sparse event streams (wearables, industrial sensors) where answers depend on a few irregular observations.","If causal flip rates remain near zero after training on this data, the bottleneck may be architectural rather than purely data-scale."],"forward_implications":["Answer accuracy alone is an insufficient metric for clinical time-series QA; evidence faithfulness and causal sensitivity must be reported.","Simply serializing longer irregular trajectories into text often adds distraction rather than signal, so evidence selection becomes a first-class modeling problem.","Native irregular time-series interfaces can cut latency by an order of magnitude while matching text-serialization accuracy near chance, motivating hybrid designs.","Future clinical QA systems need explicit mechanisms for locating sparse supporting timestamps before they can be trusted for monitoring or escalation decisions."],"fun_headline_variants":["Models top out at 50% on irregular ICU time-series QA","Generalist models miss sparse evidence for clinical answers","Correct ICU answers often lack matching temporal evidence","Sparse asynchronous ICU series stump current QA models","Models rarely ground ICU answers in supporting timestamps"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"The deterministic answer rules and human-audited labels truly capture clinically meaningful temporal reasoning, especially for the retrospective intervention and monitoring decision tasks.","fun_headline_variants_meta":{"raw":{"variants":["Models top out at 50% on irregular ICU time-series QA","Generalist models miss sparse evidence for clinical answers","Correct ICU answers often lack matching temporal evidence","Sparse asynchronous ICU series stump current QA models","Models rarely ground ICU answers in supporting timestamps"]},"model":"grok-4.5","effort":"low","cost_usd":0.005192,"raw_usage":{"total_tokens":1392,"prompt_tokens":742,"num_sources_used":0,"completion_tokens":56,"cost_in_usd_ticks":51920000,"prompt_tokens_details":{"text_tokens":742,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":594,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":742,"tokens_out":56,"duration_ms":6396,"temperature":1.0,"reasoning_tokens":594,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-14T14:51:35.138096+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"A model that, under full irregular ICU trajectories, simultaneously reaches high answer accuracy, high evidence F1, high causal flip rate when supporting observations are edited, and stable accuracy under evidence-preserving missingness and order perturbations would falsify the claim that current approaches fail at this form of reasoning.","supporting_citations":[],"review_version":1}