Pith. sign in

REVIEW 4 major objections 6 minor 48 references

Current models cannot reliably ground clinical answers in sparse, irregular ICU time series.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-14 14:51 UTC pith:AFFEJIDX

load-bearing objection Useful evidence-auditable ICU irregular-series QA benchmark; models really do fail at sparse temporal grounding, with the main soft spot being clinical validity of the decision-task labels. the 4 major comments →

arxiv 2607.09880 v1 pith:AFFEJIDX submitted 2026-07-10 cs.CL cs.AI

CLIR-Bench: Benchmarking Multimodal Question Answering over Irregular Clinical Time Series

classification cs.CL cs.AI
keywords irregular clinical time seriestime series question answeringlarge language modelsevidence faithfulnessICU monitoringmultimodal QAtemporal reasoning
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

CLIR-Bench is a 6,600-question benchmark built from de-identified ICU records to test whether models can answer clinical questions by finding and using sparse, irregular, asynchronous measurements rather than general medical knowledge. The authors construct each item with explicit timestamped evidence and a deterministic answer rule, then organize the set into four capability dimensions and eleven tasks spanning understanding, reasoning, forecasting, and decision-making. Across closed-source, open-source, and time-series LLMs, the best macro accuracy under full time-series input is only about 50 percent, faithful accuracy is far lower, full trajectories often fail to help or even hurt, and editing the causal evidence almost never flips the answer. The paper therefore claims that existing generalist models do not yet perform evidence-grounded reasoning over irregular clinical time series and that stronger selection and native irregular modeling are required.

Core claim

Existing generalist and time-series language models cannot reliably retrieve and reason over sparse, asynchronous ICU evidence for clinical question answering: under full irregular context the strongest model reaches only 50.15 percent macro accuracy, correct answers frequently lack matching evidence, and causal edits to the supporting observations flip answers at rates below roughly 1.2 percent across model families.

What carries the argument

Evidence-auditable QA construction: each of the 6,600 multiple-choice items is tied to explicit temporal evidence and a task-specific deterministic answer rule, enabling accuracy, faithfulness, sufficiency, necessity, and counterfactual edit diagnostics.

Load-bearing premise

The deterministic answer rules and human-audited labels truly capture clinically meaningful temporal reasoning, especially for the retrospective intervention and monitoring decision tasks.

What would settle it

A model that, under full irregular ICU trajectories, simultaneously reaches high answer accuracy, high evidence F1, high causal flip rate when supporting observations are edited, and stable accuracy under evidence-preserving missingness and order perturbations would falsify the claim that current approaches fail at this form of reasoning.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Answer accuracy alone is an insufficient metric for clinical time-series QA; evidence faithfulness and causal sensitivity must be reported.
  • Simply serializing longer irregular trajectories into text often adds distraction rather than signal, so evidence selection becomes a first-class modeling problem.
  • Native irregular time-series interfaces can cut latency by an order of magnitude while matching text-serialization accuracy near chance, motivating hybrid designs.
  • Future clinical QA systems need explicit mechanisms for locating sparse supporting timestamps before they can be trusted for monitoring or escalation decisions.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Benchmarks that keep gold evidence editable will become the default way to audit whether medical LLMs are using patient data or clinical priors.
  • The same evidence-auditable pipeline could be applied to other sparse event streams (wearables, industrial sensors) where answers depend on a few irregular observations.
  • If causal flip rates remain near zero after training on this data, the bottleneck may be architectural rather than purely data-scale.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper introduces CLIR-Bench, a 6,600-instance multiple-choice QA benchmark for irregular, sparse, asynchronous ICU time series built from de-identified MIMIC-IV stays. Through a four-stage pipeline (curation, task instantiation with deterministic evidence rules, LLM-assisted QA generation, human verification), it defines 11 tasks across understanding, reasoning, forecasting, and decision-making, each linked to timestamp-level evidence. Experiments under Full-TS, QA-only, Evidence-only, and Evidence-removed inputs, plus evidence-faithfulness, counterfactual causal edits, and irregularity stress tests, report that current generalist and time-series LLMs remain weak: best macro accuracy is 50.15% (GPT-5.4 mini), faithful accuracy lags answer accuracy, and causal evidence edits flip answers at rates below ~1.2%. The authors conclude that stronger evidence selection and native irregular time-series reasoning are needed.

Significance. If the gold labels and evidence spans are clinically and temporally well-specified, CLIR-Bench fills a clear gap: existing time-series QA suites largely assume regular sampling, while medical QA rarely stresses sparse asynchronous trajectories. The evidence-auditable design (rule-derived answers, Evidence-only/removed ablations, causal vs irrelevant edits) is a genuine methodological contribution and goes beyond accuracy-only leaderboards. Public data and code further strengthen reuse. The multi-diagnostic results (Table 2; Figs. 5–8) make a credible case that current models do not reliably ground answers in sparse ICU evidence, which is useful for both clinical AI and time-series LLM research. Credit is due for the structured task taxonomy, explicit cutoff handling for forecasting, and the stress-test suite that separates shortcut relief from robustness.

major comments (4)
  1. [§3.2.2 Temporal decision-making; §3.3; Appendix A.3] Sections 3.2.2 and 3.3 (and Appendix A.3): Immediate Intervention Decision (IID) and Monitoring/Escalation Decision (MED) are framed as decision-making evaluation, but the manuscript states they use retrospective clinical decision labels in a controlled setting rather than prospective clinician actions. This is load-bearing for the claim that the benchmark tests clinical decision-making. Please (i) specify exactly how gold intervention/escalation labels are derived from MIMIC-IV events, (ii) separate or reweight these two tasks when reporting the overall macro average if labels are proxy labels, and (iii) discuss the risk that models are scored for matching historical chart patterns rather than clinically justified decisions. Without this, Findings on Decision-Making and the 11-task overall score overstate clinical decision validity.
  2. [§3.3.3 Human Verification; Appendix A.3] Section 3.3.3 and Appendix A.3 describe human verification of task suitability, option quality, evidence support, and label consistency, but report no inter-annotator agreement, acceptance/rejection rates, number of reviewers, or adjudication protocol. For an evidence-auditable benchmark whose central claim rests on gold evidence spans and rule-derived answers, IAA (or at least dual-review rates and disagreement examples) is necessary. Please add quantitative audit statistics and, if dual review was not done for all 6,600 items, the sampling fraction and how disagreements were resolved.
  3. [Table 2; §4.1 Experimental Setup; §4.2 RQ1] Table 2 and §4.1–4.2: several reported accuracies (e.g., 41.67, 43.33, 58.33) are consistent with very small per-task evaluation sets (on the order of n≈60), while the abstract advertises 6,600 instances. The evaluation protocol does not clearly state whether Full-TS results use the full benchmark, a fixed subset, or model-dependent subsets, nor does it report confidence intervals or significance tests for macro averages and TS Lift. Please state exact n per task/model, how the eval split was chosen, and add uncertainty estimates so that the 50.15% best-score claim and cross-model rankings can be interpreted.
  4. [§3.3.2 QA Generation; Appendix A.3; Table 2] §3.3.2 and Appendix A.3: candidate questions are drafted with Qwen3.6-27B and rewritten with GPT-5.5, while Qwen3.6-27B (and related Qwen models) appear in the evaluated model pool (Table 2). This creates a construction–evaluation contamination risk for those families (shared phrasing priors, option style). Please either (i) exclude generator models from the main ranking, (ii) regenerate a held-out rewrite with a disjoint model and re-score, or (iii) quantify style/overlap bias. Also clarify whether gold answers ever depend on LLM judgment rather than the deterministic evidence rule alone.
minor comments (6)
  1. [Title; Abstract; §3.1] Title and abstract call the benchmark “multimodal,” but the inputs are serialized irregular time series plus text questions/options (no imaging or other modalities). Consider “time-series–language” or define multimodality explicitly in §1/§3.1.
  2. [Table 1] Table 1 comparison is useful, but several cited benchmarks have very different scopes; a short column on clinical vs general domain and on evidence-audit support would make the novelty claim sharper.
  3. [Figure 1; Table 2; §3.2.2] Figure 1 task examples are helpful; ensure every abbreviation in Table 2 (TG, ASR, TPR, MA, TSS, TF, NIF, CVR, IR, IID, MED) is defined once in a single glossary near §3.2.2 for readers skimming results.
  4. [Front matter / ACM Reference Format] ACM reference block still has placeholder conference metadata (“Conference acronym ’XX’, Woodstock, NY, 2018”). Clean for camera-ready.
  5. [§4.6 RQ6; Figure 9] §4.6 latency comparison mixes backends and is correctly caveated as wall-clock; still, state hardware and batch size so the 41× claim is not over-read as pure model efficiency.
  6. [Appendix A.1; Conclusion] Appendix A.1 ethics note is appropriate; add a one-sentence limitation that MIMIC-IV demographics and practice patterns may not transfer to other ICUs when discussing decision tasks.

Circularity Check

0 steps flagged

No significant circularity: empirical benchmark with rule-derived gold labels, not a derivation that redefines its target from fitted inputs.

full rationale

CLIR-Bench is an empirical benchmark paper, not a first-principles derivation. Its load-bearing chain is: (1) extract irregular ICU trajectories from MIMIC-IV; (2) instantiate task schemas with explicit evidence rules and answer formats; (3) draft multiple-choice items with an LLM but determine gold answers by deterministic rules cross-checked against timestamped records; (4) human-audit; (5) measure model accuracy, evidence F1/faithfulness, sufficiency/necessity, causal flip rates, and irregularity stress tests against those fixed labels. None of the six circularity patterns apply. Gold answers are not defined by model outputs or fitted parameters (Section 3.3.2: “The correct answer, however, is determined by the task-specific evidence rule and cross-checked against the original timestamped records rather than relying on the generated text alone”). Findings 1–6 report measured performance gaps (e.g., best Full-TS macro accuracy 50.15%, causal flip rates <1.2%), not predictions forced by construction. Self-citations are ordinary related-work references, not uniqueness theorems or ansatzes that force the central claim. Mild construction-loop risk (LLM drafts questions later used to evaluate LLMs) does not make labels circular, because labels remain rule- and timestamp-grounded. Score 0 is therefore appropriate.

Axiom & Free-Parameter Ledger

4 free parameters · 4 axioms · 3 invented entities

The load-bearing claim is empirical model failure on evidence-grounded irregular clinical QA. It rests on MIMIC-IV as a valid ICU source, on hand-designed task schemas and deterministic answer rules as proxies for clinical temporal reasoning, on LLM-assisted generation plus human audit as quality control, and on a few diagnostic thresholds (e.g., evidence F1 ≥ 0.5). No physical free constants; free choices are benchmark-design knobs.

free parameters (4)
  • faithful_evidence_F1_threshold
    Faithful accuracy counts a prediction only if answer is correct and evidence F1 ≥ 0.5 (Section 4.4); the 0.5 cutoff is a design choice that directly shapes Finding 3.
  • forecasting_cutoff_time
    TF/NIF separate model-visible history from future label windows via a cutoff (Sections 3.2.2, 3.3.2); cutoff placement affects difficulty and leakage control.
  • cohort_size_500_ICU_stays
    Curated cohort size and stay selection (Section 3.2.1) determine coverage of sparsity and duration distributions in Figures 2–3.
  • four_way_multiple_choice_format
    All tasks are forced into four options with 25% random baseline (Section 4.1); format choice constrains what “reasoning” can express and how distractors are built.
axioms (4)
  • domain assumption MIMIC-IV de-identified ICU trajectories are a suitable substrate for evaluating irregular clinical time-series QA.
    Section 3.2.1 builds the entire benchmark from MIMIC-IV; ethics appendix notes population/practice biases may limit generalization.
  • ad hoc to paper Task-specific deterministic evidence rules define the correct answer independently of model-generated text.
    Sections 3.3.1–3.3.2 and Appendix A.3: gold labels come from schemas/rules cross-checked to timestamps, not free-form clinical adjudication.
  • ad hoc to paper Retrospective intervention/monitoring labels in IID and MED are valid proxies for decision-making evaluation in a controlled benchmark.
    Section 3.2.2 frames decision tasks as connecting irregular evidence to retrospective clinical decision labels, not prospective trials.
  • domain assumption Serialized irregular trajectories plus natural-language questions are a fair multimodal interface for comparing generalist and time-series LLMs.
    Problem formulation (Section 3.1) and input settings (Section 4.1) assume text serialization is a legitimate evaluation mode even while RQ6 shows large latency cost.
invented entities (3)
  • CLIR-Bench independent evidence
    purpose: Provide an evidence-auditable irregular clinical time-series QA suite with 6,600 instances and diagnostic input variants.
    The benchmark is the paper’s primary constructed object; independent existence is the public Hugging Face release, not an external physical entity.
  • Four capability dimensions and 11 clinical QA tasks (TG, ASR, TPR, MA, TSS, CVR, IR, TF, NIF, IID, MED) no independent evidence
    purpose: Structure evaluation of understanding, reasoning, forecasting, and decision-making under irregularity.
    Task taxonomy is author-defined (Figure 1, Section 3.2.2); clinical meaningfulness depends on schema design rather than prior standardized clinical exams.
  • Evidence-auditable QA instance with timestamp-level evidence metadata no independent evidence
    purpose: Enable faithfulness, sufficiency/necessity, and counterfactual evidence-edit metrics beyond answer accuracy.
    Core methodological construct of the pipeline (Abstract; Sections 3.3, 4.3–4.5; Appendix A.3).

pith-pipeline@v1.1.0-grok45 · 24352 in / 3665 out tokens · 38257 ms · 2026-07-14T14:51:35.138096+00:00 · methodology

0 comments
read the original abstract

Clinical time series are central to patient monitoring, risk assessment, and clinical decision support. However, they are often sparse, irregularly sampled, and asynchronous, making it difficult for models to identify the temporal evidence required for clinical Question Answering (QA). Existing benchmarks primarily focus on regularly sampled time-series QA or medical QA over static data, and therefore rarely assess whether models can faithfully ground their answers in irregular temporal observations. To fill this gap, we introduce CLIR-Bench, a benchmark for irregular clinical time series QA constructed from de-identified ICU records through a principled four-stage pipeline. CLIR-Bench contains 6,600 QA instances spanning 11 clinical variables, organized into four capability dimensions and 11 tasks. Each question is linked to explicit temporal evidence and task-specific answer derivation rules, enabling evaluation of both answer accuracy and evidence use. Experiments show that existing generalist models struggle to retrieve and reason over sparse clinical evidence, highlighting the need for stronger irregular time-series reasoning methods. Our code and data are available at https://huggingface.co/datasets/winall/CLIR-Bench.

Figures

Figures reproduced from arXiv: 2607.09880 by Ethan B. Liu, Frank Nie, Jindong Han, Loe Yan, Wei Fan, Yuan Zhu.

Figure 1
Figure 1. Figure 1: Task categories of the CLIR-Bench. CLIR-Bench is developed through a principled four-stage pipeline: (1) time series data extraction from de-identified ICU records, (2) task instantiation to encode analytical goals in executable and veri￾fiable form, (3) QA pairs generation with explicit temporal evidence, and (4) human verification to ensure the quality of generated sam￾ples. Each QA instance is linked to… view at source ↗
Figure 2
Figure 2. Figure 2: Duration distribution of time-series samples. [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Observation sparsity distribution in CLIR-Bench. The sparsity degree denotes the fraction of missing variable￾time cells in QA instance. variable at a clinically defined anchor time, such as the value of DBP when SBP reaches its lowest point. Trend Pattern Recogni￾tion (TPR) requires the model to identify the temporal pattern of a variable within a specified window. Missingness Awareness (MA) evaluates whe… view at source ↗
Figure 4
Figure 4. Figure 4: Overview of the construction pipeline. Starting from irregular time-series records, we first instantiate task-specific [PITH_FULL_IMAGE:figures/full_fig_p005_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Comparison between QA-only and Full-TS inputs. [PITH_FULL_IMAGE:figures/full_fig_p005_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Evidence-use diagnostics across model families. (a) Strictness ladder from answer accuracy to evidence-faithful [PITH_FULL_IMAGE:figures/full_fig_p007_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Counterfactual evidence-edit evaluation. Each [PITH_FULL_IMAGE:figures/full_fig_p007_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Irregularity stress-test summary. (a) Model-family mean accuracy change under each perturbation, computed as [PITH_FULL_IMAGE:figures/full_fig_p008_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: Efficiency analysis. corresponds to a 41× median latency increase for serialized text inputs. The same pattern appears at the model level. Qwen3.5-4B and Qwen3.5-9B reach 26.36–26.44% accuracy, with median latency between 3.65 and 5.18 seconds. Qwen3.6-27B is much slower, with 40.33 seconds median latency, but does not provide a clear accuracy gain. In contrast, native time-series-input models are much fas… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

48 extracted references · 1 canonical work pages

  1. [1]

    Taha Aksu, Gerald Woo, Juncheng Liu, Xu Liu, Chenghao Liu, Silvio Savarese, Caiming Xiong, and Doyen Sahoo. 2024. GIFT-Eval: A Benchmark For General Time Series Forecasting Model Evaluation. arXiv:2410.10393 [cs.LG] https: //arxiv.org/abs/2410.10393

  2. [2]

    Anthony Bagnall, Hoang Anh Dau, Jason Lines, Michael Flynn, James Large, Aaron Bostrom, Paul Southam, and Eamonn Keogh. 2018. The UEA multivariate time series classification archive, 2018. arXiv:1811.00075 [cs.LG] https://arxiv. org/abs/1811.00075

  3. [3]

    Ane Blázquez-García, Angel Conde, Usue Mori, and Jose A. Lozano. 2020. A review on outlier/anomaly detection in time series data. arXiv:2002.04236 [cs.LG] https://arxiv.org/abs/2002.04236

  4. [4]

    Hung Bui, Harikrishna Warrier, and Yogesh Gupta. 2024. Benchmarking with MIMIC-IV, an irregular, spare clinical time series dataset. arXiv:2401.15290 [cs.LG] https://arxiv.org/abs/2401.15290

  5. [5]

    Yifu Cai, Arjun Choudhry, Mononito Goswami, and Artur Dubrawski. 2024. TimeSeriesExam: A time series understanding exam. arXiv:2410.14752 [cs.AI] https://arxiv.org/abs/2410.14752

  6. [6]

    Nimeesha Chan, Felix Parker, William Bennett, Tianyi Wu, Mung Yao Jia, James Fackler, and Kimia Ghobadi. 2024. MedTsLLM: Leveraging LLMs for Multimodal Medical Time Series Analysis. arXiv:2408.07773 [cs.LG] https://arxiv.org/abs/ 2408.07773

  7. [7]

    Jialin Chen, Aosong Feng, Ziyu Zhao, Juan Garza, Gaukhar Nurbek, Cheng Qin, Ali Maatouk, Leandros Tassiulas, Yifeng Gao, and Rex Ying. 2026. MTBench: A Multimodal Time Series Benchmark for Temporal Reasoning and Question Answering. arXiv:2503.16858 [cs.CL] https://arxiv.org/abs/2503.16858

  8. [8]

    Hsing-Huan Chung, Shijun Li, Yoav Wald, Xing Han, Suchi Saria, and Joydeep Ghosh. 2026. MILM: Large Language Models for Multimodal Irregular Time Series with Informative Sampling. arXiv:2605.13711 [cs.LG] https://arxiv.org/ abs/2605.13711

  9. [9]

    Hejie Cui, Alyssa Unell, Bowen Chen, Jason Alan Fries, Emily Alsentzer, Sanmi Koyejo, and Nigam Shah. 2025. TIMER: Temporal Instruction Modeling and Evaluation for Longitudinal Clinical Records. arXiv:2503.04176 [cs.AI] https: //arxiv.org/abs/2503.04176

  10. [10]

    Hoang Anh Dau, Anthony Bagnall, Kaveh Kamgar, Chin-Chia Michael Yeh, Yan Zhu, Shaghayegh Gharghabi, Chotirat Ann Ratanamahatana, and Eamonn Keogh. 2019. The UCR Time Series Archive. arXiv:1810.07758 [cs.LG] https: //arxiv.org/abs/1810.07758

  11. [11]

    Bodong Du, Bowen Liu, Yang Yu, Xinpeng Ding, Zhiheng Wu, Shuning Wang, Shuo Nie, Naiming Liu, Qifeng Chen, Yangqiu Song, and Xiaomeng Li. 2026. MedHorizon: Towards Long-context Medical Video Understanding in the Wild. arXiv:2605.06537 [cs.CV] https://arxiv.org/abs/2605.06537

  12. [12]

    Fleming, Alejandro Lozano, William J

    Scott L. Fleming, Alejandro Lozano, William J. Haberkorn, Jenelle A. Jindal, Eduardo P. Reis, Rahul Thapa, Louis Blankemeier, Julian Z. Genkins, Ethan Steinberg, Ashwin Nayak, Birju S. Patel, Chia-Chun Chiang, Alison Callahan, Zepeng Huo, Sergios Gatidis, Scott J. Adams, Oluseyi Fayanju, Shreya J. Shah, Thomas Savage, Ethan Goh, Akshay S. Chaudhari, Nima ...

  13. [13]

    Tong Guan, Zijie Meng, Dianqi Li, Shiyu Wang, Chao-Han Huck Yang, Qing- song Wen, Zuozhu Liu, Sabato Marco Siniscalchi, Ming Jin, and Shirui Pan. 2026. TimeOmni-1: Incentivizing Complex Reasoning with Time Series in Large Lan- guage Models. arXiv:2509.24803 [cs.LG] https://arxiv.org/abs/2509.24803

  14. [14]

    Yutao Hu, Tianbin Li, Quanfeng Lu, Wenqi Shao, Junjun He, Yu Qiao, and Ping Luo

  15. [15]

    arXiv:2402.09181 [eess.IV] https://arxiv.org/abs/2402.09181

    OmniMedVQA: A New Large-Scale Comprehensive Evaluation Benchmark for Medical LVLM. arXiv:2402.09181 [eess.IV] https://arxiv.org/abs/2402.09181

  16. [16]

    Hassan Ismail Fawaz, Germain Forestier, Jonathan Weber, Lhassane Idoumghar, and Pierre-Alain Muller. 2019. Deep learning for time series classification: a review.Data Mining and Knowledge Discovery33, 4 (March 2019), 917–963. doi:10.1007/s10618-019-00619-1

  17. [17]

    Yushan Jiang, Zijie Pan, Xikun Zhang, Sahil Garg, Anderson Schneider, Yuriy Nevmyvaka, and Dongjin Song. 2024. Empowering Time Series Analysis with Large Language Models: A Survey. arXiv:2402.03182 [cs.LG] https://arxiv.org/ abs/2402.03182

  18. [18]

    Zhang, Xiaoming Shi, Pin-Yu Chen, Yuxuan Liang, Yuan-Fang Li, Shirui Pan, and Qingsong Wen

    Ming Jin, Shiyu Wang, Lintao Ma, Zhixuan Chu, James Y. Zhang, Xiaoming Shi, Pin-Yu Chen, Yuxuan Liang, Yuan-Fang Li, Shirui Pan, and Qingsong Wen. 2024. Time-LLM: Time Series Forecasting by Reprogramming Large Language Models. arXiv:2310.01728 [cs.LG] https://arxiv.org/abs/2310.01728

  19. [19]

    Baoyu Jing, Sanhorn Chen, Lecheng Zheng, Boyu Liu, Zihao Li, Jiaru Zou, Tianxin Wei, Zhining Liu, Zhichen Zeng, Ruizhong Qiu, Xiao Lin, Yuchen Yan, Dongqi Fu, Jingchao Ni, Jingrui He, and Hanghang Tong. 2026. TSAQA: Time Series Analysis Question And Answering Benchmark. arXiv:2601.23204 [cs.AI] https: //arxiv.org/abs/2601.23204

  20. [20]

    Alistair Johnson, Lucas Bulgarelli, Tom Pollard, Brian Gow, Benjamin Moody, Steven Horng, Leo Anthony Celi, and Roger Mark. 2024. MIMIC-IV.PhysioNet (Oct. 2024). doi:10.13026/kpb9-mt58 Version 3.1

  21. [21]

    Yaxuan Kong, Yiyuan Yang, Yoontae Hwang, Wenjie Du, Stefan Zohren, Zhangyang Wang, Ming Jin, and Qingsong Wen. 2025. Time-MQA: Time Series Multi-Task Question Answering with Context Enhancement. arXiv:2503.01875 [cs.CL] https://arxiv.org/abs/2503.01875

  22. [22]

    Maya Kruse, Shiyue Hu, Nicholas Derby, Yifu Wu, Samantha Stonbraker, Bing- sheng Yao, Dakuo Wang, Elizabeth Goldberg, and Yanjun Gao. 2025. Large Lan- guage Models with Temporal Reasoning for Longitudinal Clinical Summarization and Prediction. arXiv:2501.18724 [cs.CL] https://arxiv.org/abs/2501.18724

  23. [23]

    Hao Li, Bowen Deng, Chang Xu, Zhiyuan Feng, Viktor Schlegel, Yu-Hao Huang, Yizheng Sun, Jingyuan Sun, Kailai Yang, Yiyao Yu, and Jiang Bian

  24. [24]

    MIRA: Medical Time Series Foundation Model for Real-World Health Conference acronym ’XX, June 03–05, 2018, Woodstock, NY Trovato et al. Data. arXiv:2506.07584 [cs.LG] https://arxiv.org/abs/2506.07584

  25. [25]

    Bryan Lim and Stefan Zohren. 2021. Time-series forecasting with deep learning: a survey.Philosophical Transactions of the Royal Society A: Mathematical, Physical and Engineering Sciences379, 2194 (Feb. 2021), 20200209. doi:10.1098/rsta.2020. 0209

  26. [26]

    Bowen Liu, Haoyang Li, Shuning Wang, Shuo Nie, and Shanghang Zhang. 2025. Subgraph Aggregation for Out-of-Distribution Generalization on Graphs.Pro- ceedings of the AAAI Conference on Artificial Intelligence39, 18 (April 2025), 18763–18771. doi:10.1609/aaai.v39i18.34065

  27. [27]

    Bowen Liu, Li Yang, Shanshan Song, Mingyu Tang, Zhifang Gao, Qifeng Chen, Yangqiu Song, Huimin Chen, and Xiaomeng Li. 2026. Divide-then-Diagnose: Weaving Clinician-Inspired Contexts for Ultra-Long Capsule Endoscopy Videos. arXiv preprint arXiv:2604.21814(2026)

  28. [28]

    Yong Liu, Guo Qin, Xiangdong Huang, Jianmin Wang, and Mingsheng Long

  29. [29]

    arXiv:2402.02370 [cs.LG] https://arxiv.org/abs/2402.02370

    AutoTimes: Autoregressive Time Series Forecasters via Large Language Models. arXiv:2402.02370 [cs.LG] https://arxiv.org/abs/2402.02370

  30. [30]

    Shuo Nie, Hexuan Deng, Chao Wang, Ruiyu Fang, Xuebo Liu, Shuangyong Song, Yu Li, Min Zhang, and Xuelong Li. 2026. Stop Rewarding Hallucinated Steps: Faithfulness-Aware Step-Level Reinforcement Learning for Small Reasoning Models. arXiv:2602.05897 [cs.CL] https://arxiv.org/abs/2602.05897

  31. [31]

    Omiye, Haiwen Gui, Shawheen J

    Jesutofunmi A. Omiye, Haiwen Gui, Shawheen J. Rezaei, James Zou, and Roxana Daneshjou. 2024. Large Language Models in Medicine: The Potentials and Pitfalls: A Narrative Review.Annals of Internal Medicine177, 2 (Feb. 2024), 210–220. doi:10.7326/m23-2772

  32. [32]

    Shrey Pandit, Jiawei Xu, Junyuan Hong, Zhangyang Wang, Tianlong Chen, Kaidi Xu, and Ying Ding. 2025. MedHallu: A Comprehensive Benchmark for Detecting Medical Hallucinations in Large Language Models. arXiv:2502.14302 [cs.CL] https://arxiv.org/abs/2502.14302

  33. [33]

    Jensen, Zhenli Sheng, and Bin Yang

    Xiangfei Qiu, Jilin Hu, Lekui Zhou, Xingjian Wu, Junyang Du, Buang Zhang, Chenjuan Guo, Aoying Zhou, Christian S. Jensen, Zhenli Sheng, and Bin Yang

  34. [34]

    arXiv:2403.20150 [cs.LG] https://arxiv.org/abs/2403.20150

    TFB: Towards Comprehensive and Fair Benchmarking of Time Series Forecasting Methods. arXiv:2403.20150 [cs.LG] https://arxiv.org/abs/2403.20150

  35. [35]

    Jensen, and Bin Yang

    Xiangfei Qiu, Zhe Li, Wanghui Qiu, Shiyan Hu, Lekui Zhou, Xingjian Wu, Zhengyu Li, Chenjuan Guo, Aoying Zhou, Zhenli Sheng, Jilin Hu, Christian S. Jensen, and Bin Yang. 2025. TAB: Unified Benchmarking of Time Series Anomaly Detection Methods. arXiv:2506.18046 [cs.LG] https://arxiv.org/abs/2506.18046

  36. [36]

    Qwen Team. 2026. Qwen3.6-Plus: Towards Real World Agents. https://qwen.ai/ blog?id=qwen3.6

  37. [37]

    Satya Narayan Shukla and Benjamin M. Marlin. 2021. Multi-Time Attention Networks for Irregularly Sampled Time Series. arXiv:2101.10318 [cs.LG] https: //arxiv.org/abs/2101.10318

  38. [38]

    Qwen Team. 2026. Qwen3.5: Accelerating Productivity with Native Multimodal Agents. https://qwen.ai/blog?id=qwen3.5

  39. [39]

    Sindhu Tipirneni and Chandan K. Reddy. 2022. Self-Supervised Trans- former for Sparse and Irregularly Sampled Multivariate Clinical Time-Series. arXiv:2107.14293 [cs.LG] https://arxiv.org/abs/2107.14293

  40. [40]

    Yilin Wang, Peixuan Lei, Jie Song, Yuzhe Hao, Tao Chen, Yuxuan Zhang, Lei Jia, Yuanxiang Li, and Zhongyu Wei. 2025. ITFormer: Bridging Time Series and Natural Language for Multi-Modal QA with Large-Scale Multitask Dataset. arXiv:2506.20093 [cs.CL] https://arxiv.org/abs/2506.20093

  41. [41]

    Stephan Xie, Ben Cohen, Mononito Goswami, Junhong Shen, Emaad Khwaja, Chenghao Liu, David Asker, Othmane Abou-Amal, and Ameet Talwalkar. 2026. ARFBench: Benchmarking Time Series Question Answering Ability for Software Incident Response. arXiv:2604.21199 [cs.LG] https://arxiv.org/abs/2604.21199

  42. [42]

    Zhe Xie, Zeyan Li, Xiao He, Longlong Xu, Xidao Wen, Tieying Zhang, Jianjun Chen, Rui Shi, and Dan Pei. 2025. ChatTS: Aligning Time Series with LLMs via Synthetic Data for Enhanced Understanding and Reasoning.Proceedings of the VLDB Endowment18, 8 (April 2025), 2385–2398. doi:10.14778/3742728.3742735

  43. [43]

    Lawrence K. Q. Yan, Qian Niu, Ming Li, Yichao Zhang, Caitlyn Heqi Yin, Cheng Fei, Benji Peng, Ziqian Bi, Pohsun Feng, Keyu Chen, Tianyang Wang, Yunze Wang, Silin Chen, Ming Liu, Junyu Liu, Xinyuan Song, Riyang Bao, Zekun Jiang, and Ziyuan Qin. 2025. Large Language Model Benchmarks in Medical Tasks. arXiv:2410.21348 [cs.CL] https://arxiv.org/abs/2410.21348

  44. [44]

    Yao Yin, Zhenyu Xiao, Musheng Li, Yiwen Liu, Sutong Nan, Yiting He, Ruiqi Wang, Zhenwei Zhang, Qingmin Liao, and Yuantao Gu. 2026. MMTS-BENCH: A Comprehensive Benchmark for Time Series Understanding and Reasoning. arXiv:2602.08588 [cs.DB] https://arxiv.org/abs/2602.08588

  45. [45]

    Fangxu Yu, Xingang Guo, Lingzhi Yuan, Haoqiang Kang, Hongyu Zhao, Lianhui Qin, Furong Huang, Bin Hu, and Tianyi Zhou. 2026. TSRBench: A Comprehensive Multi-task Multi-modal Time Series Reasoning Benchmark for Generalist Models. arXiv:2601.18744 [cs.AI] https://arxiv.org/abs/2601.18744

  46. [46]

    Fangxu Yu, Hongyu Zhao, and Tianyi Zhou. 2025. TS-Reasoner: Aligning Time Series Foundation Models with LLM Reasoning. arXiv:2510.03519 [cs.CL] https: //arxiv.org/abs/2510.03519

  47. [47]

    Gupta, and Jingbo Shang

    Xiyuan Zhang, Ranak Roy Chowdhury, Rajesh K. Gupta, and Jingbo Shang. 2024. Large Language Models for Time Series: A Survey. arXiv:2402.01801 [cs.LG] https://arxiv.org/abs/2402.01801

  48. [48]

    Feixiang Zheng, Yu Wu, Cecilia Mascolo, and Ting Dang. 2026. Rethinking Large Language Models For Irregular Time Series Classification In Critical Care. arXiv:2601.16516 [cs.LG] https://arxiv.org/abs/2601.16516 CLIR-Bench: Benchmarking Multimodal Question Answering over Irregular Clinical Time Series Conference acronym ’XX, June 03–05, 2018, Woodstock, NY...