{"id":"b3d49960-b2da-4796-9905-b899e5979921","arxiv_id":"2507.13655","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":2.0,"correctness_risk":"high","formal_verification":"none","parameter_count":7,"one_line_summary":"CU-ICU applies LoRA, AdaLoRA, and (IA)3 to FLAN-T5 for ICU sepsis detection, mortality prediction, and note generation, claiming efficiency gains that are not backed by reported baselines.","lead":"This paper proposes CU-ICU, a way to adapt the FLAN-T5 language model to ICU tasks using parameter-efficient fine-tuning. It reports accuracy gains on sepsis detection, mortality prediction, and clinical note generation, but does not show the standard fine-tuning baseline that the headline improvements are measured against.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The claimed 15%/20% gains over standard fine-tuning cannot be evaluated because no full-fine-tuning baseline appears in Tables 1 or 2; the relative improvement is the central claim and is unsupported.","rationale":"I aligned with the reader's REJECT verdict. The strongest claim is relative outperformance over standard fine-tuning; the weakest assumption is that such baselines exist. The paper's own tables make this assumption unverifiable: every row is a PEFT method and no non-PEFT baseline appears. Inspecting §4.5, the baseline statement is a single sentence with no table, no numbers, no hyperparameters, and no dataset identifiers. Because '15% increase' and '20% enhancement' are explicitly relative claims, the absence of denominators makes those headline numbers unfalsifiable from the manuscript. I considered whether the absolute scores or PEFT comparisons might still support a weaker claim, and they might, but that weaker claim is not the one made in the abstract. The nBERTScore validity is a secondary concern; without a clinically validated metric, interpretability gains are also unestablished, but the missing baseline is sufficient. The mathematical formulation of the PEFT methods (Eqs. 3-8) is internally consistent, so the issue is not derivational but empirical; however, the empirical support for the central comparative assertion is absent. I recommend no change to the REJECT verdict: the concern reinforces the reader's weakest_assumption rather than altering it.","tokens_in":9510,"tokens_out":3095,"duration_ms":37653,"concrete_test":"Re-run the same ICU tasks with three controls: (1) full fine-tuning of all FLAN-T5 parameters, (2) frozen FLAN-T5 with only the 16-shot prompts, and (3) the reported PEFT configurations, all on identical data splits, seeds, and evaluation scripts. Report accuracy/nBERTScore per control in a new table and then recompute the claimed 15% and 20% improvements. If full fine-tuning is not run or the gains are not reproduced, the central claim fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is comparative: CU-ICU 'consistently improves predictive accuracy and interpretability over standard fine-tuning methods' (Abstract), with up to 15% sepsis accuracy and 20% nBERTScore gains (§5). The only empirical support for this is Section 4.5, which says standard fine-tuning baselines were used, but Tables 1 and 2 report exclusively PEFT variants (LoRA, AdaLoRA, (IA)3). No full-fine-tuning accuracy, nBERTScore, training budget, or standard-deviation row is shown. Because the baseline numbers are absent, the 15%/20% relative improvements stated in §5 cannot be recomputed or verified from the reported data; the central comparative assertion is therefore unsupported even if the absolute scores (e.g., 85.6% sepsis, 32.1 nBERTScore) are internally consistent. A secondary but related problem is that the 'Avg' column in Table 2 averages sepsis accuracy and mortality accuracy against nBERTScore with no defined weighting, so the reported 'avg performance' is not meaningful. This is a load-bearing gap: the contribution of CU-ICU over ordinary fine-tuning is the entire advertised benefit.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents CU-ICU, a framework that adapts FLAN-T5 to ICU tasks via sparse parameter-efficient fine-tuning (LoRA, AdaLoRA, and (IA)3). The tasks are sepsis detection, mortality prediction, and clinical note generation. The abstract and Section 5 claim that CU-ICU consistently improves predictive accuracy and interpretability over standard fine-tuning, with up to 15% higher sepsis accuracy and 20% higher nBERTScore while updating fewer than 1% of parameters. The reported experiments show absolute scores for the three PEFT variants (e.g., 85.6% sepsis accuracy and 32.1 nBERTScore for the best (IA)3 configuration), but the comparison to standard fine-tuning is not documented.","tokens_in":9753,"tokens_out":7344,"duration_ms":77138,"significance":"The practical motivation is clear: efficient adaptation of instruction-tuned language models to data-scarce clinical domains is a relevant problem, and the paper systematically compares three PEFT methods under a unified text-to-text setup with standard deviations over five seeds. If the comparative claims were supported, the resource savings (0.5-6.2% trainable parameters) would be a useful contribution. However, the core contribution is the claimed superiority over standard fine-tuning, and that claim is not supported by the reported data. The paper also introduces no new method; its value hinges entirely on the empirical comparison, which is incomplete.","major_comments":[{"comment":"The central claim that CU-ICU 'consistently improves predictive accuracy and interpretability over standard fine-tuning methods' is unsupported because no standard fine-tuning baseline is reported. Section 4.5 states that such baselines were used, but Tables 1 and 2 contain only LoRA, AdaLoRA, and (IA)3 configurations, with no accuracy, nBERTScore, parameter counts, or training budgets for a fully fine-tuned model. Consequently, the 'approximately 15% increase in early sepsis detection accuracy' and '20% enhancement in generating clinically relevant notes' reported in Section 5 cannot be recomputed, and the comparison that defines the paper's contribution is unverifiable. Please add the missing baseline results (with standard deviations) or reframe the contributions as absolute scores without comparative claims.","section":"§4.5, §5, Tables 1-2"},{"comment":"The 'Avg' column averages sepsis accuracy, mortality accuracy, and note nBERTScore, which are not commensurable: the first two are bounded percentages while nBERTScore is a semantic similarity score on a different scale. No normalization or weighting is defined, so an average such as 66.0 has no meaningful interpretation. Remove the column or replace it with a clearly defined aggregate over comparable metrics.","section":"Table 2"},{"comment":"The datasets underlying all experiments are unnamed. 'Real-world ICU records' with unspecified sources, sample sizes, class distributions, and split arrangements make the absolute accuracies (e.g., 85.6% sepsis) impossible to reproduce or assess. Since few-shot performance is a key claim, please specify the datasets (with citations), the number of examples per task, how the 16-shot prompts were constructed and sampled, and the train/validation/test splits.","section":"§4.1-4.2"},{"comment":"The phrase 'unsupervised instruction-finetuned' is inaccurate for FLAN-T5, which was instruction-finetuned on supervised data (Chung et al., 2022). Because this phrase appears in the title and framing of the contribution, the terminology should be corrected or explicitly defined. The current wording mischaracterizes the base model and undermines the paper's conceptual framing.","section":"Abstract, §1, §5"}],"minor_comments":[{"comment":"The citation for (IA)3 is inconsistent: Section 2 cites it as [15] (Lester et al.), while Section 3.3 cites [8] (Guo et al.). Please cite the original (IA)3 paper consistently.","section":"§2, §3.3"},{"comment":"The formalization of AdaLoRA in Eq. (5) as ΔW = A diag(α) B omits the sum over rank components and the SVD-based triplet structure of the original method; if this is an intentional simplification, state that explicitly.","section":"§3.3"},{"comment":"The size of the FLAN-T5 base model (e.g., base, large, XL) is not specified, yet parameter percentages (0.5%-6.2%) and the nBERTScore values depend on it. Please state the model variant and the total parameter count.","section":"§4.2"},{"comment":"The explanation that (IA)3 excels because it 'modulate[s] attention weights adaptively' conflicts with the activation-scaling mechanism described in Section 3.3 and is not supported by any attention analysis. Please align the interpretation with the method or add supporting evidence.","section":"§5.1"},{"comment":"The example response from (IA)3 claims that the vital signs 'meet Sepsis-3 criteria,' but Sepsis-3 defines sepsis as organ dysfunction (SOFA score increase) rather than the SIRS-like combination of fever, tachycardia, hypotension, and leukocytosis. This sample highlights the need for clinical validation of the generated explanations, which is currently absent.","section":"Appendix A.1"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a preprint with signs of rushed composition, including reference inconsistencies and unspecified datasets. The missing standard fine-tuning baseline is a substantial gap because the paper's headline improvements are relative claims. I recommend major revision rather than outright rejection because the missing baseline could in principle be supplied, but the authors must also correct the aggregate metric and dataset reporting before the central claims can be evaluated."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a routine application of LoRA, AdaLoRA, and (IA)3 to FLAN-T5 on three ICU text tasks, wrapped in a prompt-engineering setup. The one thing that would make it interesting—consistent gains over full fine-tuning—is asserted but not shown. The standard fine-tuning baseline never appears in any table, so the headline 15% and 20% numbers are unsupported.\n\nWhat it does well: the T5 prompt formats for sepsis, mortality, and note generation are described clearly, and the hyperparameter table is transparent. The authors report parameter percentages and standard deviations over five seeds, which is more than many clinical NLP preprints bother to do. The limitations section is also frank about generalizability and missing multimodal data, even if it is a bit boilerplate.\n\nThe problems are in the evidence. Table 1 has only PEFT variants, no full fine-tuning. Table 2 is worse: the 'Avg' column averages accuracy percentages (0–100) with nBERTScore (likely 0–1 rescaled? unclear) into a single number with no defined weighting. That column is meaningless. The datasets are never named—the experimental section says 'real-world ICU records,' and the limitations say 'publicly available datasets,' but which ones? The reader cannot check whether the absolute numbers are plausible. There is also a citation numbering mess: (IA)3 is [8] in the approach and [15] in related work; AdaLoRA is both [26] and [28]; and [15] is actually the Lester prompt-tuning paper. Sloppy, but the bigger issue is the missing baseline. The abstract promises '15% increase in sepsis detection accuracy' over standard fine-tuning; Section 5 repeats it, but the only basis is a sentence in Section 4.5 saying such baselines were compared. That is not enough.\n\nIs the paper incoherent? No. The equations are standard, the training setup is plausible, and I don't see signs of fabrication—just an incomplete evaluation. Still, the central comparative claim is a load-bearing gap, and the contribution is incremental. This is an engineering report that needs more work, not a paper that should take up referee time in its current form. My recommendation: desk reject, with an invitation to resubmit if the authors add a proper full-fine-tuning baseline, name the datasets, report per-task standard deviations, and fix the averaging and citations. I would not bring it to a reading group, and I wouldn't cite it.","headline":"A straightforward PEFT application whose central claim rests on a baseline that never appears in the paper.","tokens_in":10290,"tokens_out":3276,"would_cite":false,"duration_ms":34796,"reading_group":"no","serious_thinker":"yes","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that sparse parameter-efficient fine-tuning of FLAN-T5, updating fewer than 1% of parameters, outperforms standard fine-tuning on ICU sepsis detection, mortality prediction, and clinical note generation using only 16-shot…","keywords":["parameter-efficient fine-tuning","FLAN-T5","intensive care unit","sepsis detection","mortality prediction","clinical note generation","few-shot prompting","(IA)3"],"falsifier":"Run the same three ICU tasks with the same data and random seeds under standard full fine-tuning of FLAN-T5 and report its sepsis accuracy, mortality accuracy, and note nBERTScore; if that baseline reaches or exceeds 85.6%, 80.2%, and 32.1 respectively, the claimed advantage of CU-ICU's sparse updates over standard fine-tuning disappears.","tokens_in":9272,"feed_emoji":"🏥","tokens_out":8779,"duration_ms":86726,"temperature":0.7,"pith_summary":"CU-ICU is a recipe for taking an already instruction-tuned language model and pushing it into the intensive care unit without retraining the whole network. The paper claims that freezing the base model and updating only a sparse set of adapter parameters—via LoRA, AdaLoRA, or (IA)3—beats standard fine-tuning on sepsis detection, mortality prediction, and clinical note generation, all from as few as sixteen labeled examples. If true, this would make specialized clinical language models practical in hospitals where labeled data and compute are scarce. The headline numbers are 85.6% sepsis accuracy, 80.2% mortality accuracy, and a 32.1 note-quality score with under 1% of parameters updated.","feed_headline":"Under 1% of parameters yields 85.6% sepsis accuracy","feed_subtitle":"FLAN-T5 adapted with LoRA, AdaLoRA, and (IA)3 needs only 16 examples to lift ICU predictions.","key_machinery":"The central object is a sparse parameter delta, $\\Delta\\theta$, added to frozen base weights, $\\theta_0$, so the adapted model is $\\theta = \\theta_0 + \\Delta\\theta$. The paper instantiates this delta with three mechanisms: LoRA, which writes the update as a product of two low-rank matrices; AdaLoRA, which adds per-component importance weights and prunes rank during training; and (IA)3, which multiplies transformer activations by learned element-wise scaling vectors $\\gamma$. The (IA)3 variant carries the argument: it gives the best reported numbers while updating fewer than 1% of parameters. All three are driven by 16-shot prompts, so the framework's bet is that sparse updates plus instruction-finetuned priors are enough to absorb the ICU domain shift.","core_discovery":"The paper's central claim is that domain adaptation to ICU data does not require updating the large model at all. Starting from FLAN-T5, CU-ICU freezes the pretrained weights and learns only small delta parameters, then evaluates every task through the same text-to-text interface: a clinical prompt in, a label or note out. Across the three tasks, the authors report that the sparsest configuration—(IA)3, which learns element-wise scaling vectors for transformer activations—performs best: 85.6% accuracy for early sepsis detection, 80.2% for mortality prediction, and a 32.1 note nBERTScore for clinical note generation. They further report an average improvement of roughly 15% over standard fine-tuning on sepsis accuracy and 20% on clinically relevant note quality, which is the empirical basis for calling the framework both accurate and interpretable.","pith_inferences":["An implication the paper leaves implicit is that nothing in the prompt design is ICU-specific, so the same sparse-adaptation recipe could be pointed at other data-scarce clinical text tasks such as radiology reports, discharge summaries, or triage notes.","Because the standard fine-tuning baseline is described but never appears in the tables, the 15% and 20% headline gains are best read as claims about an unreported comparison; a decisive extension would be to publish the full fine-tuning run with identical data and seeds.","The paper treats nBERTScore as a proxy for clinically relevant explanations, but reports no clinician evaluation; a testable extension is to have ICU clinicians rate generated notes and check whether their judgments track the score.","The title's 'unsupervised' label is not doing work in the experiments: FLAN-T5 is instruction-finetuned with supervision, and the adaptation stage uses labeled 16-shot examples, so the method's actual regime is few-shot supervised adaptation."],"forward_implications":["If the reported comparisons hold, ICU decision support can be built by adapting a general instruction-finetuned model with a handful of labeled examples and under 1% of parameters, making deployment feasible where annotated data and compute are scarce.","Because all three tasks run through the same text-to-text interface, one backbone can serve sepsis detection, mortality prediction, and note generation, simplifying clinical software maintenance.","Since (IA)3 modifies only activation scaling vectors, the adapted model stays close to the original FLAN-T5, which should make it easier to inspect what the domain adaptation changed.","The reported average gains of about 15% on sepsis accuracy and 20% on note quality are the direct evidence for preferring sparse parameter-efficient fine-tuning over standard fine-tuning in this setting."],"supporting_citations":[{"why":"Supplies the T5 text-to-text architecture that CU-ICU uses to unify all ICU tasks into prompt-to-text mappings.","marker":"[22]"},{"why":"Supplies FLAN-T5, the instruction-finetuned backbone whose few-shot priors the paper adapts.","marker":"[6]"},{"why":"Supplies LoRA, the low-rank update mechanism that is one of the three PEFT variants evaluated.","marker":"[10]"},{"why":"Supplies AdaLoRA, the adaptive-rank update mechanism that is the second PEFT variant evaluated.","marker":"[26]"},{"why":"Supplies (IA)3, the activation-scaling method that yields the paper's best reported results.","marker":"[8]"},{"why":"Supplies the note nBERTScore metric used to measure clinical note generation quality.","marker":"[17]"}],"fun_headline_variants":["Freeze T5, add sparse deltas: 85.6% sepsis accuracy","CU-ICU tunes <1% of T5 params, beats full fine-tuning","85.6% sepsis accuracy from T5 with under 1% parameter updates"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The headline gains are computed against a standard fine-tuning baseline that the paper describes in its methods but never reports in its results tables, so the 15% and 20% improvements stand or fall on that baseline having actually been run.","fun_headline_variants_meta":{"raw":{"variants":["Freeze T5, add sparse deltas: 85.6% sepsis accuracy","CU-ICU tunes <1% of T5 params, beats full fine-tuning","85.6% sepsis accuracy from T5 with under 1% parameter updates"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000434,"raw_usage":{"total_tokens":2192,"prompt_tokens":910,"completion_tokens":1282,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":526,"completion_tokens_details":{"reasoning_tokens":1213}},"tokens_in":526,"tokens_out":1282,"duration_ms":12524,"temperature":1.0,"reasoning_tokens":1213,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T16:19:04.688969+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same three ICU tasks with the same data and random seeds under standard full fine-tuning of FLAN-T5 and report its sepsis accuracy, mortality accuracy, and note nBERTScore; if that baseline reaches or exceeds 85.6%, 80.2%, and 32.1 respectively, the claimed advantage of CU-ICU's sparse updates over standard fine-tuning disappears.","supporting_citations":[{"cited_title":"Scaling instruction-finetuned language models","cited_arxiv_id":null,"evidence_quote":"Supplies FLAN-T5, the instruction-finetuned backbone whose few-shot priors the paper adapts."},{"cited_title":"Lora: Low-rank adaptation of large language models","cited_arxiv_id":null,"evidence_quote":"Supplies LoRA, the low-rank update mechanism that is one of the three PEFT variants evaluated."},{"cited_title":"Adalora: Adaptive low-rank adaptation for efficient fine-tuning of large language models","cited_arxiv_id":null,"evidence_quote":"Supplies AdaLoRA, the adaptive-rank update mechanism that is the second PEFT variant evaluated."},{"cited_title":"Parameter-efficient transfer learning with adaptive attention","cited_arxiv_id":null,"evidence_quote":"Supplies (IA)3, the activation-scaling method that yields the paper's best reported results."},{"cited_title":"nbertscore: Evaluating clinical note generation with semantic and clinical similarity","cited_arxiv_id":null,"evidence_quote":"Supplies the note nBERTScore metric used to measure clinical note generation quality."}],"review_version":1}