REVIEW 4 major objections 6 minor 48 references
Current models cannot reliably ground clinical answers in sparse, irregular ICU time series.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-14 14:51 UTC pith:AFFEJIDX
load-bearing objection Useful evidence-auditable ICU irregular-series QA benchmark; models really do fail at sparse temporal grounding, with the main soft spot being clinical validity of the decision-task labels. the 4 major comments →
CLIR-Bench: Benchmarking Multimodal Question Answering over Irregular Clinical Time Series
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Existing generalist and time-series language models cannot reliably retrieve and reason over sparse, asynchronous ICU evidence for clinical question answering: under full irregular context the strongest model reaches only 50.15 percent macro accuracy, correct answers frequently lack matching evidence, and causal edits to the supporting observations flip answers at rates below roughly 1.2 percent across model families.
What carries the argument
Evidence-auditable QA construction: each of the 6,600 multiple-choice items is tied to explicit temporal evidence and a task-specific deterministic answer rule, enabling accuracy, faithfulness, sufficiency, necessity, and counterfactual edit diagnostics.
Load-bearing premise
The deterministic answer rules and human-audited labels truly capture clinically meaningful temporal reasoning, especially for the retrospective intervention and monitoring decision tasks.
What would settle it
A model that, under full irregular ICU trajectories, simultaneously reaches high answer accuracy, high evidence F1, high causal flip rate when supporting observations are edited, and stable accuracy under evidence-preserving missingness and order perturbations would falsify the claim that current approaches fail at this form of reasoning.
If this is right
- Answer accuracy alone is an insufficient metric for clinical time-series QA; evidence faithfulness and causal sensitivity must be reported.
- Simply serializing longer irregular trajectories into text often adds distraction rather than signal, so evidence selection becomes a first-class modeling problem.
- Native irregular time-series interfaces can cut latency by an order of magnitude while matching text-serialization accuracy near chance, motivating hybrid designs.
- Future clinical QA systems need explicit mechanisms for locating sparse supporting timestamps before they can be trusted for monitoring or escalation decisions.
Where Pith is reading between the lines
- Benchmarks that keep gold evidence editable will become the default way to audit whether medical LLMs are using patient data or clinical priors.
- The same evidence-auditable pipeline could be applied to other sparse event streams (wearables, industrial sensors) where answers depend on a few irregular observations.
- If causal flip rates remain near zero after training on this data, the bottleneck may be architectural rather than purely data-scale.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces CLIR-Bench, a 6,600-instance multiple-choice QA benchmark for irregular, sparse, asynchronous ICU time series built from de-identified MIMIC-IV stays. Through a four-stage pipeline (curation, task instantiation with deterministic evidence rules, LLM-assisted QA generation, human verification), it defines 11 tasks across understanding, reasoning, forecasting, and decision-making, each linked to timestamp-level evidence. Experiments under Full-TS, QA-only, Evidence-only, and Evidence-removed inputs, plus evidence-faithfulness, counterfactual causal edits, and irregularity stress tests, report that current generalist and time-series LLMs remain weak: best macro accuracy is 50.15% (GPT-5.4 mini), faithful accuracy lags answer accuracy, and causal evidence edits flip answers at rates below ~1.2%. The authors conclude that stronger evidence selection and native irregular time-series reasoning are needed.
Significance. If the gold labels and evidence spans are clinically and temporally well-specified, CLIR-Bench fills a clear gap: existing time-series QA suites largely assume regular sampling, while medical QA rarely stresses sparse asynchronous trajectories. The evidence-auditable design (rule-derived answers, Evidence-only/removed ablations, causal vs irrelevant edits) is a genuine methodological contribution and goes beyond accuracy-only leaderboards. Public data and code further strengthen reuse. The multi-diagnostic results (Table 2; Figs. 5–8) make a credible case that current models do not reliably ground answers in sparse ICU evidence, which is useful for both clinical AI and time-series LLM research. Credit is due for the structured task taxonomy, explicit cutoff handling for forecasting, and the stress-test suite that separates shortcut relief from robustness.
major comments (4)
- [§3.2.2 Temporal decision-making; §3.3; Appendix A.3] Sections 3.2.2 and 3.3 (and Appendix A.3): Immediate Intervention Decision (IID) and Monitoring/Escalation Decision (MED) are framed as decision-making evaluation, but the manuscript states they use retrospective clinical decision labels in a controlled setting rather than prospective clinician actions. This is load-bearing for the claim that the benchmark tests clinical decision-making. Please (i) specify exactly how gold intervention/escalation labels are derived from MIMIC-IV events, (ii) separate or reweight these two tasks when reporting the overall macro average if labels are proxy labels, and (iii) discuss the risk that models are scored for matching historical chart patterns rather than clinically justified decisions. Without this, Findings on Decision-Making and the 11-task overall score overstate clinical decision validity.
- [§3.3.3 Human Verification; Appendix A.3] Section 3.3.3 and Appendix A.3 describe human verification of task suitability, option quality, evidence support, and label consistency, but report no inter-annotator agreement, acceptance/rejection rates, number of reviewers, or adjudication protocol. For an evidence-auditable benchmark whose central claim rests on gold evidence spans and rule-derived answers, IAA (or at least dual-review rates and disagreement examples) is necessary. Please add quantitative audit statistics and, if dual review was not done for all 6,600 items, the sampling fraction and how disagreements were resolved.
- [Table 2; §4.1 Experimental Setup; §4.2 RQ1] Table 2 and §4.1–4.2: several reported accuracies (e.g., 41.67, 43.33, 58.33) are consistent with very small per-task evaluation sets (on the order of n≈60), while the abstract advertises 6,600 instances. The evaluation protocol does not clearly state whether Full-TS results use the full benchmark, a fixed subset, or model-dependent subsets, nor does it report confidence intervals or significance tests for macro averages and TS Lift. Please state exact n per task/model, how the eval split was chosen, and add uncertainty estimates so that the 50.15% best-score claim and cross-model rankings can be interpreted.
- [§3.3.2 QA Generation; Appendix A.3; Table 2] §3.3.2 and Appendix A.3: candidate questions are drafted with Qwen3.6-27B and rewritten with GPT-5.5, while Qwen3.6-27B (and related Qwen models) appear in the evaluated model pool (Table 2). This creates a construction–evaluation contamination risk for those families (shared phrasing priors, option style). Please either (i) exclude generator models from the main ranking, (ii) regenerate a held-out rewrite with a disjoint model and re-score, or (iii) quantify style/overlap bias. Also clarify whether gold answers ever depend on LLM judgment rather than the deterministic evidence rule alone.
minor comments (6)
- [Title; Abstract; §3.1] Title and abstract call the benchmark “multimodal,” but the inputs are serialized irregular time series plus text questions/options (no imaging or other modalities). Consider “time-series–language” or define multimodality explicitly in §1/§3.1.
- [Table 1] Table 1 comparison is useful, but several cited benchmarks have very different scopes; a short column on clinical vs general domain and on evidence-audit support would make the novelty claim sharper.
- [Figure 1; Table 2; §3.2.2] Figure 1 task examples are helpful; ensure every abbreviation in Table 2 (TG, ASR, TPR, MA, TSS, TF, NIF, CVR, IR, IID, MED) is defined once in a single glossary near §3.2.2 for readers skimming results.
- [Front matter / ACM Reference Format] ACM reference block still has placeholder conference metadata (“Conference acronym ’XX’, Woodstock, NY, 2018”). Clean for camera-ready.
- [§4.6 RQ6; Figure 9] §4.6 latency comparison mixes backends and is correctly caveated as wall-clock; still, state hardware and batch size so the 41× claim is not over-read as pure model efficiency.
- [Appendix A.1; Conclusion] Appendix A.1 ethics note is appropriate; add a one-sentence limitation that MIMIC-IV demographics and practice patterns may not transfer to other ICUs when discussing decision tasks.
Circularity Check
No significant circularity: empirical benchmark with rule-derived gold labels, not a derivation that redefines its target from fitted inputs.
full rationale
CLIR-Bench is an empirical benchmark paper, not a first-principles derivation. Its load-bearing chain is: (1) extract irregular ICU trajectories from MIMIC-IV; (2) instantiate task schemas with explicit evidence rules and answer formats; (3) draft multiple-choice items with an LLM but determine gold answers by deterministic rules cross-checked against timestamped records; (4) human-audit; (5) measure model accuracy, evidence F1/faithfulness, sufficiency/necessity, causal flip rates, and irregularity stress tests against those fixed labels. None of the six circularity patterns apply. Gold answers are not defined by model outputs or fitted parameters (Section 3.3.2: “The correct answer, however, is determined by the task-specific evidence rule and cross-checked against the original timestamped records rather than relying on the generated text alone”). Findings 1–6 report measured performance gaps (e.g., best Full-TS macro accuracy 50.15%, causal flip rates <1.2%), not predictions forced by construction. Self-citations are ordinary related-work references, not uniqueness theorems or ansatzes that force the central claim. Mild construction-loop risk (LLM drafts questions later used to evaluate LLMs) does not make labels circular, because labels remain rule- and timestamp-grounded. Score 0 is therefore appropriate.
Axiom & Free-Parameter Ledger
free parameters (4)
- faithful_evidence_F1_threshold
- forecasting_cutoff_time
- cohort_size_500_ICU_stays
- four_way_multiple_choice_format
axioms (4)
- domain assumption MIMIC-IV de-identified ICU trajectories are a suitable substrate for evaluating irregular clinical time-series QA.
- ad hoc to paper Task-specific deterministic evidence rules define the correct answer independently of model-generated text.
- ad hoc to paper Retrospective intervention/monitoring labels in IID and MED are valid proxies for decision-making evaluation in a controlled benchmark.
- domain assumption Serialized irregular trajectories plus natural-language questions are a fair multimodal interface for comparing generalist and time-series LLMs.
invented entities (3)
-
CLIR-Bench
independent evidence
-
Four capability dimensions and 11 clinical QA tasks (TG, ASR, TPR, MA, TSS, CVR, IR, TF, NIF, IID, MED)
no independent evidence
-
Evidence-auditable QA instance with timestamp-level evidence metadata
no independent evidence
read the original abstract
Clinical time series are central to patient monitoring, risk assessment, and clinical decision support. However, they are often sparse, irregularly sampled, and asynchronous, making it difficult for models to identify the temporal evidence required for clinical Question Answering (QA). Existing benchmarks primarily focus on regularly sampled time-series QA or medical QA over static data, and therefore rarely assess whether models can faithfully ground their answers in irregular temporal observations. To fill this gap, we introduce CLIR-Bench, a benchmark for irregular clinical time series QA constructed from de-identified ICU records through a principled four-stage pipeline. CLIR-Bench contains 6,600 QA instances spanning 11 clinical variables, organized into four capability dimensions and 11 tasks. Each question is linked to explicit temporal evidence and task-specific answer derivation rules, enabling evaluation of both answer accuracy and evidence use. Experiments show that existing generalist models struggle to retrieve and reason over sparse clinical evidence, highlighting the need for stronger irregular time-series reasoning methods. Our code and data are available at https://huggingface.co/datasets/winall/CLIR-Bench.
Figures
Reference graph
Works this paper leans on
-
[1]
Taha Aksu, Gerald Woo, Juncheng Liu, Xu Liu, Chenghao Liu, Silvio Savarese, Caiming Xiong, and Doyen Sahoo. 2024. GIFT-Eval: A Benchmark For General Time Series Forecasting Model Evaluation. arXiv:2410.10393 [cs.LG] https: //arxiv.org/abs/2410.10393
Pith/arXiv arXiv 2024
-
[2]
Anthony Bagnall, Hoang Anh Dau, Jason Lines, Michael Flynn, James Large, Aaron Bostrom, Paul Southam, and Eamonn Keogh. 2018. The UEA multivariate time series classification archive, 2018. arXiv:1811.00075 [cs.LG] https://arxiv. org/abs/1811.00075
Pith/arXiv arXiv 2018
-
[3]
Ane Blázquez-García, Angel Conde, Usue Mori, and Jose A. Lozano. 2020. A review on outlier/anomaly detection in time series data. arXiv:2002.04236 [cs.LG] https://arxiv.org/abs/2002.04236
Pith/arXiv arXiv 2020
-
[4]
Hung Bui, Harikrishna Warrier, and Yogesh Gupta. 2024. Benchmarking with MIMIC-IV, an irregular, spare clinical time series dataset. arXiv:2401.15290 [cs.LG] https://arxiv.org/abs/2401.15290
Pith/arXiv arXiv 2024
-
[5]
Yifu Cai, Arjun Choudhry, Mononito Goswami, and Artur Dubrawski. 2024. TimeSeriesExam: A time series understanding exam. arXiv:2410.14752 [cs.AI] https://arxiv.org/abs/2410.14752
Pith/arXiv arXiv 2024
-
[6]
Nimeesha Chan, Felix Parker, William Bennett, Tianyi Wu, Mung Yao Jia, James Fackler, and Kimia Ghobadi. 2024. MedTsLLM: Leveraging LLMs for Multimodal Medical Time Series Analysis. arXiv:2408.07773 [cs.LG] https://arxiv.org/abs/ 2408.07773
Pith/arXiv arXiv 2024
-
[7]
Jialin Chen, Aosong Feng, Ziyu Zhao, Juan Garza, Gaukhar Nurbek, Cheng Qin, Ali Maatouk, Leandros Tassiulas, Yifeng Gao, and Rex Ying. 2026. MTBench: A Multimodal Time Series Benchmark for Temporal Reasoning and Question Answering. arXiv:2503.16858 [cs.CL] https://arxiv.org/abs/2503.16858
arXiv 2026
-
[8]
Hsing-Huan Chung, Shijun Li, Yoav Wald, Xing Han, Suchi Saria, and Joydeep Ghosh. 2026. MILM: Large Language Models for Multimodal Irregular Time Series with Informative Sampling. arXiv:2605.13711 [cs.LG] https://arxiv.org/ abs/2605.13711
Pith/arXiv arXiv 2026
-
[9]
Hejie Cui, Alyssa Unell, Bowen Chen, Jason Alan Fries, Emily Alsentzer, Sanmi Koyejo, and Nigam Shah. 2025. TIMER: Temporal Instruction Modeling and Evaluation for Longitudinal Clinical Records. arXiv:2503.04176 [cs.AI] https: //arxiv.org/abs/2503.04176
Pith/arXiv arXiv 2025
-
[10]
Hoang Anh Dau, Anthony Bagnall, Kaveh Kamgar, Chin-Chia Michael Yeh, Yan Zhu, Shaghayegh Gharghabi, Chotirat Ann Ratanamahatana, and Eamonn Keogh. 2019. The UCR Time Series Archive. arXiv:1810.07758 [cs.LG] https: //arxiv.org/abs/1810.07758
Pith/arXiv arXiv 2019
-
[11]
Bodong Du, Bowen Liu, Yang Yu, Xinpeng Ding, Zhiheng Wu, Shuning Wang, Shuo Nie, Naiming Liu, Qifeng Chen, Yangqiu Song, and Xiaomeng Li. 2026. MedHorizon: Towards Long-context Medical Video Understanding in the Wild. arXiv:2605.06537 [cs.CV] https://arxiv.org/abs/2605.06537
Pith/arXiv arXiv 2026
-
[12]
Fleming, Alejandro Lozano, William J
Scott L. Fleming, Alejandro Lozano, William J. Haberkorn, Jenelle A. Jindal, Eduardo P. Reis, Rahul Thapa, Louis Blankemeier, Julian Z. Genkins, Ethan Steinberg, Ashwin Nayak, Birju S. Patel, Chia-Chun Chiang, Alison Callahan, Zepeng Huo, Sergios Gatidis, Scott J. Adams, Oluseyi Fayanju, Shreya J. Shah, Thomas Savage, Ethan Goh, Akshay S. Chaudhari, Nima ...
Pith/arXiv arXiv 2023
-
[13]
Tong Guan, Zijie Meng, Dianqi Li, Shiyu Wang, Chao-Han Huck Yang, Qing- song Wen, Zuozhu Liu, Sabato Marco Siniscalchi, Ming Jin, and Shirui Pan. 2026. TimeOmni-1: Incentivizing Complex Reasoning with Time Series in Large Lan- guage Models. arXiv:2509.24803 [cs.LG] https://arxiv.org/abs/2509.24803
Pith/arXiv arXiv 2026
-
[14]
Yutao Hu, Tianbin Li, Quanfeng Lu, Wenqi Shao, Junjun He, Yu Qiao, and Ping Luo
-
[15]
arXiv:2402.09181 [eess.IV] https://arxiv.org/abs/2402.09181
OmniMedVQA: A New Large-Scale Comprehensive Evaluation Benchmark for Medical LVLM. arXiv:2402.09181 [eess.IV] https://arxiv.org/abs/2402.09181
-
[16]
Hassan Ismail Fawaz, Germain Forestier, Jonathan Weber, Lhassane Idoumghar, and Pierre-Alain Muller. 2019. Deep learning for time series classification: a review.Data Mining and Knowledge Discovery33, 4 (March 2019), 917–963. doi:10.1007/s10618-019-00619-1
-
[17]
Yushan Jiang, Zijie Pan, Xikun Zhang, Sahil Garg, Anderson Schneider, Yuriy Nevmyvaka, and Dongjin Song. 2024. Empowering Time Series Analysis with Large Language Models: A Survey. arXiv:2402.03182 [cs.LG] https://arxiv.org/ abs/2402.03182
Pith/arXiv arXiv 2024
-
[18]
Zhang, Xiaoming Shi, Pin-Yu Chen, Yuxuan Liang, Yuan-Fang Li, Shirui Pan, and Qingsong Wen
Ming Jin, Shiyu Wang, Lintao Ma, Zhixuan Chu, James Y. Zhang, Xiaoming Shi, Pin-Yu Chen, Yuxuan Liang, Yuan-Fang Li, Shirui Pan, and Qingsong Wen. 2024. Time-LLM: Time Series Forecasting by Reprogramming Large Language Models. arXiv:2310.01728 [cs.LG] https://arxiv.org/abs/2310.01728
Pith/arXiv arXiv 2024
-
[19]
Baoyu Jing, Sanhorn Chen, Lecheng Zheng, Boyu Liu, Zihao Li, Jiaru Zou, Tianxin Wei, Zhining Liu, Zhichen Zeng, Ruizhong Qiu, Xiao Lin, Yuchen Yan, Dongqi Fu, Jingchao Ni, Jingrui He, and Hanghang Tong. 2026. TSAQA: Time Series Analysis Question And Answering Benchmark. arXiv:2601.23204 [cs.AI] https: //arxiv.org/abs/2601.23204
Pith/arXiv arXiv 2026
-
[20]
Alistair Johnson, Lucas Bulgarelli, Tom Pollard, Brian Gow, Benjamin Moody, Steven Horng, Leo Anthony Celi, and Roger Mark. 2024. MIMIC-IV.PhysioNet (Oct. 2024). doi:10.13026/kpb9-mt58 Version 3.1
-
[21]
Yaxuan Kong, Yiyuan Yang, Yoontae Hwang, Wenjie Du, Stefan Zohren, Zhangyang Wang, Ming Jin, and Qingsong Wen. 2025. Time-MQA: Time Series Multi-Task Question Answering with Context Enhancement. arXiv:2503.01875 [cs.CL] https://arxiv.org/abs/2503.01875
Pith/arXiv arXiv 2025
-
[22]
Maya Kruse, Shiyue Hu, Nicholas Derby, Yifu Wu, Samantha Stonbraker, Bing- sheng Yao, Dakuo Wang, Elizabeth Goldberg, and Yanjun Gao. 2025. Large Lan- guage Models with Temporal Reasoning for Longitudinal Clinical Summarization and Prediction. arXiv:2501.18724 [cs.CL] https://arxiv.org/abs/2501.18724
Pith/arXiv arXiv 2025
-
[23]
Hao Li, Bowen Deng, Chang Xu, Zhiyuan Feng, Viktor Schlegel, Yu-Hao Huang, Yizheng Sun, Jingyuan Sun, Kailai Yang, Yiyao Yu, and Jiang Bian
-
[24]
MIRA: Medical Time Series Foundation Model for Real-World Health Conference acronym ’XX, June 03–05, 2018, Woodstock, NY Trovato et al. Data. arXiv:2506.07584 [cs.LG] https://arxiv.org/abs/2506.07584
arXiv 2018
-
[25]
Bryan Lim and Stefan Zohren. 2021. Time-series forecasting with deep learning: a survey.Philosophical Transactions of the Royal Society A: Mathematical, Physical and Engineering Sciences379, 2194 (Feb. 2021), 20200209. doi:10.1098/rsta.2020. 0209
-
[26]
Bowen Liu, Haoyang Li, Shuning Wang, Shuo Nie, and Shanghang Zhang. 2025. Subgraph Aggregation for Out-of-Distribution Generalization on Graphs.Pro- ceedings of the AAAI Conference on Artificial Intelligence39, 18 (April 2025), 18763–18771. doi:10.1609/aaai.v39i18.34065
-
[27]
Bowen Liu, Li Yang, Shanshan Song, Mingyu Tang, Zhifang Gao, Qifeng Chen, Yangqiu Song, Huimin Chen, and Xiaomeng Li. 2026. Divide-then-Diagnose: Weaving Clinician-Inspired Contexts for Ultra-Long Capsule Endoscopy Videos. arXiv preprint arXiv:2604.21814(2026)
Pith/arXiv arXiv 2026
-
[28]
Yong Liu, Guo Qin, Xiangdong Huang, Jianmin Wang, and Mingsheng Long
-
[29]
arXiv:2402.02370 [cs.LG] https://arxiv.org/abs/2402.02370
AutoTimes: Autoregressive Time Series Forecasters via Large Language Models. arXiv:2402.02370 [cs.LG] https://arxiv.org/abs/2402.02370
-
[30]
Shuo Nie, Hexuan Deng, Chao Wang, Ruiyu Fang, Xuebo Liu, Shuangyong Song, Yu Li, Min Zhang, and Xuelong Li. 2026. Stop Rewarding Hallucinated Steps: Faithfulness-Aware Step-Level Reinforcement Learning for Small Reasoning Models. arXiv:2602.05897 [cs.CL] https://arxiv.org/abs/2602.05897
Pith/arXiv arXiv 2026
-
[31]
Jesutofunmi A. Omiye, Haiwen Gui, Shawheen J. Rezaei, James Zou, and Roxana Daneshjou. 2024. Large Language Models in Medicine: The Potentials and Pitfalls: A Narrative Review.Annals of Internal Medicine177, 2 (Feb. 2024), 210–220. doi:10.7326/m23-2772
doi:10.7326/m23-2772 2024
-
[32]
Shrey Pandit, Jiawei Xu, Junyuan Hong, Zhangyang Wang, Tianlong Chen, Kaidi Xu, and Ying Ding. 2025. MedHallu: A Comprehensive Benchmark for Detecting Medical Hallucinations in Large Language Models. arXiv:2502.14302 [cs.CL] https://arxiv.org/abs/2502.14302
Pith/arXiv arXiv 2025
-
[33]
Jensen, Zhenli Sheng, and Bin Yang
Xiangfei Qiu, Jilin Hu, Lekui Zhou, Xingjian Wu, Junyang Du, Buang Zhang, Chenjuan Guo, Aoying Zhou, Christian S. Jensen, Zhenli Sheng, and Bin Yang
-
[34]
arXiv:2403.20150 [cs.LG] https://arxiv.org/abs/2403.20150
TFB: Towards Comprehensive and Fair Benchmarking of Time Series Forecasting Methods. arXiv:2403.20150 [cs.LG] https://arxiv.org/abs/2403.20150
-
[35]
Xiangfei Qiu, Zhe Li, Wanghui Qiu, Shiyan Hu, Lekui Zhou, Xingjian Wu, Zhengyu Li, Chenjuan Guo, Aoying Zhou, Zhenli Sheng, Jilin Hu, Christian S. Jensen, and Bin Yang. 2025. TAB: Unified Benchmarking of Time Series Anomaly Detection Methods. arXiv:2506.18046 [cs.LG] https://arxiv.org/abs/2506.18046
Pith/arXiv arXiv 2025
-
[36]
Qwen Team. 2026. Qwen3.6-Plus: Towards Real World Agents. https://qwen.ai/ blog?id=qwen3.6
2026
-
[37]
Satya Narayan Shukla and Benjamin M. Marlin. 2021. Multi-Time Attention Networks for Irregularly Sampled Time Series. arXiv:2101.10318 [cs.LG] https: //arxiv.org/abs/2101.10318
Pith/arXiv arXiv 2021
-
[38]
Qwen Team. 2026. Qwen3.5: Accelerating Productivity with Native Multimodal Agents. https://qwen.ai/blog?id=qwen3.5
2026
-
[39]
Sindhu Tipirneni and Chandan K. Reddy. 2022. Self-Supervised Trans- former for Sparse and Irregularly Sampled Multivariate Clinical Time-Series. arXiv:2107.14293 [cs.LG] https://arxiv.org/abs/2107.14293
Pith/arXiv arXiv 2022
-
[40]
Yilin Wang, Peixuan Lei, Jie Song, Yuzhe Hao, Tao Chen, Yuxuan Zhang, Lei Jia, Yuanxiang Li, and Zhongyu Wei. 2025. ITFormer: Bridging Time Series and Natural Language for Multi-Modal QA with Large-Scale Multitask Dataset. arXiv:2506.20093 [cs.CL] https://arxiv.org/abs/2506.20093
Pith/arXiv arXiv 2025
-
[41]
Stephan Xie, Ben Cohen, Mononito Goswami, Junhong Shen, Emaad Khwaja, Chenghao Liu, David Asker, Othmane Abou-Amal, and Ameet Talwalkar. 2026. ARFBench: Benchmarking Time Series Question Answering Ability for Software Incident Response. arXiv:2604.21199 [cs.LG] https://arxiv.org/abs/2604.21199
Pith/arXiv arXiv 2026
-
[42]
Zhe Xie, Zeyan Li, Xiao He, Longlong Xu, Xidao Wen, Tieying Zhang, Jianjun Chen, Rui Shi, and Dan Pei. 2025. ChatTS: Aligning Time Series with LLMs via Synthetic Data for Enhanced Understanding and Reasoning.Proceedings of the VLDB Endowment18, 8 (April 2025), 2385–2398. doi:10.14778/3742728.3742735
-
[43]
Lawrence K. Q. Yan, Qian Niu, Ming Li, Yichao Zhang, Caitlyn Heqi Yin, Cheng Fei, Benji Peng, Ziqian Bi, Pohsun Feng, Keyu Chen, Tianyang Wang, Yunze Wang, Silin Chen, Ming Liu, Junyu Liu, Xinyuan Song, Riyang Bao, Zekun Jiang, and Ziyuan Qin. 2025. Large Language Model Benchmarks in Medical Tasks. arXiv:2410.21348 [cs.CL] https://arxiv.org/abs/2410.21348
arXiv 2025
-
[44]
Yao Yin, Zhenyu Xiao, Musheng Li, Yiwen Liu, Sutong Nan, Yiting He, Ruiqi Wang, Zhenwei Zhang, Qingmin Liao, and Yuantao Gu. 2026. MMTS-BENCH: A Comprehensive Benchmark for Time Series Understanding and Reasoning. arXiv:2602.08588 [cs.DB] https://arxiv.org/abs/2602.08588
arXiv 2026
-
[45]
Fangxu Yu, Xingang Guo, Lingzhi Yuan, Haoqiang Kang, Hongyu Zhao, Lianhui Qin, Furong Huang, Bin Hu, and Tianyi Zhou. 2026. TSRBench: A Comprehensive Multi-task Multi-modal Time Series Reasoning Benchmark for Generalist Models. arXiv:2601.18744 [cs.AI] https://arxiv.org/abs/2601.18744
Pith/arXiv arXiv 2026
-
[46]
Fangxu Yu, Hongyu Zhao, and Tianyi Zhou. 2025. TS-Reasoner: Aligning Time Series Foundation Models with LLM Reasoning. arXiv:2510.03519 [cs.CL] https: //arxiv.org/abs/2510.03519
arXiv 2025
-
[47]
Xiyuan Zhang, Ranak Roy Chowdhury, Rajesh K. Gupta, and Jingbo Shang. 2024. Large Language Models for Time Series: A Survey. arXiv:2402.01801 [cs.LG] https://arxiv.org/abs/2402.01801
Pith/arXiv arXiv 2024
-
[48]
Feixiang Zheng, Yu Wu, Cecilia Mascolo, and Ting Dang. 2026. Rethinking Large Language Models For Irregular Time Series Classification In Critical Care. arXiv:2601.16516 [cs.LG] https://arxiv.org/abs/2601.16516 CLIR-Bench: Benchmarking Multimodal Question Answering over Irregular Clinical Time Series Conference acronym ’XX, June 03–05, 2018, Woodstock, NY...
arXiv 2026
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.