{"id":"60bbc5b4-d183-47b2-8162-ff8aff00fb57","arxiv_id":"2506.22805","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"FLAME models outcome risk as a sum of an unknown duration-response curve over individual exposure episodes, and its cardiac surgery application estimates that sustained hypotension raises AKI risk more per minute than fragmented hypotension, with wide credible intervals.","lead":"FLAME is a new statistical model that estimates how the risk of an outcome grows with the duration of each exposure episode, instead of only total exposure time. Applied to low blood pressure during heart surgery, it suggests a single 60-minute episode may carry more kidney injury risk than 60 one-minute episodes, though the evidence is not conclusive.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The estimated duration-response and the f(60) vs 60f(1) contrast rest entirely on the untested additive exchangeable-episode assumption in Eq. (1); no comparison to total-duration/count or phase-specific models is provided.","rationale":"The most load-bearing requirement for the central claim is not the spline estimator (which the simulations support) but the structural assumption that each episode's contribution depends only on its duration and adds on the logit scale. The paper states this model in Eq. (1), acknowledges in Section 3.1 that a single RAF over the whole surgery is a simplification, and defers phase-specific RAFs to future work. Because the simulations generate data from exactly the same additive model, they provide no evidence about this assumption. The application's headline contrast f(60)-60f(1) is a direct function of additivity: if long episodes correlate with high-risk surgical phases, or if numerous short episodes interact, the contrast is not the duration-response the paper claims. A comparison to total-duration/count models or phase-specific models is therefore the missing piece that would justify the scientific conclusion. The paper has real strengths—clear writing, reproducible R package, and a sensible semiparametric framework—so the issue warrants a conditional, not a rejection, verdict. The reader's conditional verdict already captures this, so no verdict adjustment is needed.","tokens_in":11148,"tokens_out":8307,"duration_ms":104639,"concrete_test":"Refit the JHH data with (i) a model replacing the episode sum with a smooth term for total hypotensive duration plus a smooth or linear term for number of episodes, e.g., logit(p_i)=X_iβ+h(T_i)+s(J_i); (ii) a phase-specific FLAME, logit(p_i)=X_iβ+sum_j f_{phase(i,j)}(z_ij), with phases defined by CPB status or surgery tercile. Compare LOO-CV ELPD and the posterior of f(60)-60f(1) across models. If the total-duration/count model fits comparably or better, or if the phase-specific model shifts the contrast by more than the width of the original 95% credible interval, then Eq. (1)'s exchangeability assumption is not supported and the headline contrast cannot be interpreted as duration-dependent risk.","verdict_should_be":"UNCHANGED","load_bearing_attack":"FLAME's central claim—that it recovers the true duration-response and that the application contrast f(60) versus 60f(1) is a valid measure of duration-dependent risk—depends on Eq. (1)'s assumption that episodes contribute additively on the logit scale with no dependence on timing, order, episode count, or surgical phase. The paper explicitly adopts a single RAF over the entire surgery (Section 3.1) and defers phase-specific RAFs to future work. The simulations in Section 4 cannot validate this assumption because they generate outcomes from the same additive model (logit(p_i)=X_iβ+sum_j f(z_ij)); they only verify that the spline estimator recovers f when the model is true. The real-data analysis never compares FLAME with a model using total hypotensive time (smooth) plus number of episodes, nor with phase- or order-dependent models. If long episodes occur disproportionately during cardiopulmonary bypass or late in surgery, or if many short episodes interact or sensitize, the estimated f and especially f(60)-60f(1) are biased summaries of duration-specific risk. The credible interval for the difference already includes 0, so the 38% claim is not statistically significant; but the deeper problem is that even the point estimate is only interpretable under the exchangeability/additivity assumption, which is unexamined.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes FLAME, a Bayesian semiparametric model for outcomes in which each exposure episode contributes a term f(duration) to the linear predictor, with f estimated by penalized B-splines under a nonnegativity constraint. The authors demonstrate through simulations that the spline estimator recovers linear, piecewise-linear, logarithmic, and sigmoid risk accumulation functions, and they apply the model to 1,188 cardiac surgery patients to estimate how the duration of intraoperative hypotension episodes relates to postoperative acute kidney injury. The application reports that a single 60-minute hypotensive episode is associated with an AKI probability of about 0.32 versus 0.23 for 60 one-minute episodes, a difference of 0.09 with a 95% credible interval that includes zero. The paper also announces an R package for the method.","tokens_in":11454,"tokens_out":10016,"duration_ms":96374,"significance":"If the additive episode-level model is correct, FLAME is a useful and conceptually simple addition to the toolkit for high-resolution exposure data, since it avoids collapsing episodes into total duration and directly targets contrasts such as f(60) versus 60f(1). The strengths are the systematic simulation study, the transparent Bayesian implementation in Stan, the sensitivity analysis with respect to the number of spline bases, and the provision of software. However, the application's headline contrast is not statistically significant, is estimated from very sparse long-duration data, and is only interpretable under an exchangeability/additivity assumption that the paper does not test; these issues currently prevent the paper's conclusions from being accepted at face value.","major_comments":[{"comment":"The central claim that FLAME recovers the duration-response and that f(60) versus 60f(1) is a valid contrast depends entirely on the assumption that episodes contribute additively and exchangeably on the logit scale, ignoring timing, order, episode-count interactions, and surgical phase; the model uses a single RAF over the entire surgery, as acknowledged in Section 6. The simulations in Section 4 generate outcomes from exactly this model, so they cannot validate the assumption, and the real-data analysis in Section 5 never compares FLAME with a model using total hypotensive time plus number of episodes or with phase- or order-dependent specifications. If long episodes occur disproportionately during cardiopulmonary bypass or late in surgery, or if repeated episodes sensitize the kidney, the estimated f and the headline contrast are biased summaries of duration-specific risk; please add sensitivity analyses that relax the additivity/exchangeability assumption or clearly label the application as conditional on this untested assumption.","section":"Section 3.1, Eq. (1)"},{"comment":"The application's headline finding is not statistically significant: the 95% credible interval for the probability difference between one 60-minute episode and 60 one-minute episodes is [0, 0.21] for K=30 and [-0.01, 0.20] for K=40, both of which include zero or have a boundary at zero. The relevant exposure range is extremely sparse, since only 2% of episodes last 20 minutes or longer and only 23 patients have an episode of at least 60 minutes (Table 2), so the estimated acceleration beyond 20 minutes and the value f(60) are driven by very few observations. The abstract's '38% increase' and 'reveals' overstate the evidence; please report the posterior probability that f(60) exceeds 60f(1), discuss the implications of the sparse support, and temper the conclusion accordingly.","section":"Section 5.2, Table 3"},{"comment":"The simulation study demonstrates that the spline estimator recovers f when the data are generated from the same additive episode model, but it does not test robustness to any plausible misspecification of Eq. (1), such as phase-specific effects, order dependence, or an additional effect of episode count. The simulations also only cover durations up to 30 minutes, so the estimator's behavior at f(60), the key application contrast, is never evaluated. Consequently, the statement in Section 6 that simulations 'confirm FLAME's ability to recover the true RAF' is too strong; at minimum, the claim should be limited to the correctly specified case, and ideally the authors would add a misspecification scenario, such as generating from a phase-specific or total-duration-plus-count model and examining bias in the contrast.","section":"Section 4 and Table 1"},{"comment":"Because the half-normal priors enforce f(z) ≥ 0, the posterior of the contrast f(60) − 60f(1) is supported on [0, ∞), so the interval [0, 0.21] is a one-sided interval whose lower endpoint is a boundary value; it does not by itself convey the strength of evidence for a positive contrast. The paper should report the posterior probability P(f(60) > 60f(1)) and should also present an unconstrained fit, or otherwise justify the nonnegativity constraint, so that readers can assess how much of the estimated effect is due to the constraint.","section":"Section 3.2, Table 3"}],"minor_comments":[{"comment":"The abstract and the body report inconsistent numbers: the abstract states probabilities 0.24 and 0.33 and a 38% increase, while the body reports 0.23 and 0.32 and a 37% increase; harmonize these values.","section":"Abstract and Section 5.2"},{"comment":"The text refers to the third simulation row as 'exponential,' but the four simulation shapes are linear, piecewise linear, logarithm, and sigmoid; correct this wording.","section":"Section 4.1"},{"comment":"The notation f(z)=zγ is ambiguous: if γ is a multiplier, write γz, and if it is an exponent, the displayed identity sum_j z_j^γ = (sum_j z_j)γ is algebraically incorrect.","section":"Section 3.1"},{"comment":"The R package is called 'flameRisk' in the abstract and 'flame' in the body; make the name consistent and provide a repository or link.","section":"Abstract and Section 5.2"},{"comment":"No MCMC convergence diagnostics, such as R-hat or effective sample size, are reported; add them for both the simulation and application analyses.","section":"Sections 4 and 5"},{"comment":"The acronym 'FOSR' is introduced without being defined; spell out the term at first use.","section":"Section 2"},{"comment":"The predicted probabilities in Table 3 do not state the covariate values at which they are evaluated; specify whether these are marginal predictions or predictions at mean covariates.","section":"Table 3"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a reasonable methods paper with a useful and transparent model, but the abstract and conclusions overstate the application result, which is not statistically significant and depends on an untested exchangeability/additivity assumption. The reader's stress-test concern about Eq. (1) lands: the simulations only validate the estimator under the same additive model used to generate the data. If the authors add sensitivity analyses or substantially weaken the causal and clinical claims, the paper could become acceptable; in its current form, the central application claim is not supported by the evidence presented."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick read. The modeling idea is real: write risk as a sum of an unknown function over each exposure episode, estimate it with Bayesian P-splines, then form contrasts like f(60) vs 60 f(1). I don't know prior work that frames duration-dependent accumulation at episode level this way, and the linear total-duration model falling out as a special case is good calibration. The simulation work is solid as estimation evidence: data generated from the same additive model, and the estimator recovers linear, piecewise-linear, log, and sigmoid shapes with ISE shrinking as n grows. That is not circular; it validates the computational machinery. The R package and the honest reporting of CI coverage are pluses.\n\nThe soft spots are real but mostly in the application. The paper admits the key contrast has a 95% CI including zero (0.09, CI [0, 0.21] in Table 3), that only 2% of episodes exceed 20 minutes, and that episodes over 75 minutes were excluded. The abstract's \"38% increase\" reads as a finding when it is an exploratory point estimate over a very thin tail. The deeper issue is the one the stress-test note names: the RAF and its contrasts are only interpretable under the additive exchangeable-episode assumption in Eq. (1). Timing, order, episode count, and surgical phase are all ignored, and no comparison to total-duration or phase-specific models is offered. The simulations cannot address this because they generate outcomes from the same additive structure; they only verify the estimator, not the model. The paper does flag phase-specific RAFs as future work, which is good, but the application section does not treat the assumption as a limitation on interpretation.\n\nNet assessment: the methodological contribution is worth having, and the writing and simulation reporting are clear and fair. The flaws are addressable in revision: tone down the applied claim, add comparison models, and frame the clinical result as illustrative. I'd send this to peer review and ask for those changes.\n\nWho this is for: biostatisticians and epidemiologists working with high-resolution episodic exposures who need a model that distinguishes duration patterns from total dose. It deserves a serious referee despite my skepticism about the applied headline.","headline":"A genuinely useful episode-level GAM for duration-dependent risk, with honest simulations, but the applied headline contrast is overstated and the underlying exchangeability assumption is unexamined.","tokens_in":11981,"tokens_out":1651,"would_cite":true,"duration_ms":21225,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62G08","62F15","62P10"],"pacs":[],"model":"deepseek-v4-flash","headline":"A single 60-minute hypotensive episode is associated with roughly 0.09 higher AKI probability than sixty 1-minute episodes in a cardiac-surgery cohort.","keywords":["risk accumulation function","episodic exposure","duration-dependent risk","Bayesian P-splines","intraoperative hypotension","acute kidney injury","semiparametric regression"],"falsifier":"Re-fit FLAME on the same cohort with an added term for surgical phase or episode order; if the estimated risk curve differs by phase, or if the 0.09 contrast between one 60-minute episode and sixty 1-minute episodes shrinks to zero, the key claim is not supported.","tokens_in":10964,"feed_emoji":"🩺","tokens_out":10965,"duration_ms":98420,"temperature":0.7,"pith_summary":"The paper introduces FLAME, a semiparametric model that treats each episode of exposure as contributing risk through an unknown function of its duration, with contributions summed on the linear predictor scale. It claims this episode-level structure recovers the true duration-response curve in simulations and, applied to 1,188 cardiac-surgery patients, shows AKI risk accumulates roughly linearly for hypotensive episodes under 20 minutes and faster than linearly afterward. The central contrast is that a single continuous 60-minute episode is associated with AKI probability 0.32 versus 0.23 for 60 one-minute episodes, a difference of 0.09 on the probability scale at identical total exposure time. If true, this gives clinicians a duration-based target for intraoperative blood-pressure management and a general template for episodic exposures in other fields.","feed_headline":"Longer hypotensive stretches raise AKI risk faster than short ones","feed_subtitle":"When total hypotension time is the same, how it is packaged changes predicted AKI risk.","key_machinery":"The central object is the risk accumulation function $f(z)$: the amount of logit-scale risk contributed by a single exposure episode of duration $z$. It is modeled with cubic B-spline bases, equally spaced knots, a second-order difference penalty, and a half-Cauchy prior on the smoothing variance, with the constraint $f(0) \\approx 0$ encoded through a tight prior on the first spline coefficient. The sum over episodes makes total duration a special case and turns the scientific question 'is one long episode riskier than many short ones?' into a posterior comparison between $f(60)$ and $60 f(1)$.","core_discovery":"FLAME claims that the risk contribution of an exposure process is additive over episodes: $g(\\mathbb{E}[y_i]) = X_i\\beta + \\sum_j f(z_{ij})$, where $z_{ij}$ is the duration of episode $j$ and $f$ is a smooth risk accumulation function estimated with Bayesian penalized B-splines. Under this model, total duration as a single covariate is the special linear case $f(z) = z^\\gamma$. The paper reports that the estimated $f$ is approximately linear below 20 minutes and superlinear above, so the same total hypotension time carries different risk depending on how it is packaged: the posterior mean AKI probability rises from 0.21 with no hypotension to 0.23 for sixty one-minute episodes and to 0.32 for one sixty-minute episode, a difference of 0.09 on the probability scale with 95% credible interval [0, 0.21].","pith_inferences":["If the additive duration-only structure holds, a natural extension is a phase-specific FLAME with a separate risk accumulation function before, during, and after cardiopulmonary bypass; material differences would refine the clinical trigger.","The same framework could test dose packaging in other settings, such as whether short bursts of air pollution separated by clean periods carry less risk than one sustained exposure of equal average concentration.","Because the paper excludes the eight patients with episodes over 75 minutes, the superlinear tail beyond one hour rests on sparse data; a larger cohort with longer episodes would test whether the bend at 20 minutes continues."],"forward_implications":["If risk accumulation is superlinear beyond 20 minutes, then intraoperative intervention targets should be stated as a maximum continuous duration of hypotension, not just a total-minutes budget.","Because total duration is nested as the linear special case, a clearly nonlinear estimated $f$ is evidence against the usual scalar-summary model.","The posterior contrast $f(60) - 60 f(1)$ provides a single interpretable quantity for communicating risk, and its credible interval shows the strength of evidence for the episodic effect.","Similar estimated AKI probabilities across $K = 20$, $30$, and $40$ B-spline bases indicate the headline conclusion is not an artifact of the spline basis count."],"supporting_citations":[{"why":"Supplies the cardiac-surgery cohort, high-resolution MAP recordings, and the AKI outcome data used in the application.","marker":"Goeddel et al., 2024"},{"why":"Provides the B-spline basis with difference penalties that FLAME uses to model the risk accumulation function.","marker":"Eilers & Marx, 1996"},{"why":"Defines the Bayesian P-spline framework and the random-walk prior on spline coefficients adopted for estimation.","marker":"Lang & Brezger, 2004"},{"why":"Motivates the half-Cauchy prior on the smoothing variance used in the Bayesian implementation.","marker":"Gelman, 2006"},{"why":"Defines the KDIGO criteria used to classify acute kidney injury, the outcome in the application.","marker":"Khwaja, 2012"},{"why":"Establishes the clinical relevance of intraoperative hypotension and the MAP threshold underlying the exposure definition.","marker":"Scott & APSF Hemodynamic Instability Writing Group, 2024"}],"fun_headline_variants":["Same hypotension time, different AKI risk: duration packaging matters","AKI risk rises 38% when hypotension is one stretch instead of many","Duration of each low-blood-pressure episode drives AKI risk, not just total time","Same hypotension dose, different AKI risk: continuous episodes are worse","How hypotension is packaged over time alters AKI risk even with same total duration"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The model assumes that every episode of low blood pressure adds its risk separately, with only its length mattering, so the timing, order, and depth of episodes have no effect.","fun_headline_variants_meta":{"raw":{"variants":["Same hypotension time, different AKI risk: duration packaging matters","AKI risk rises 38% when hypotension is one stretch instead of many","Duration of each low-blood-pressure episode drives AKI risk, not just total time","Same hypotension dose, different AKI risk: continuous episodes are worse","How hypotension is packaged over time alters AKI risk even with same total duration"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000992,"raw_usage":{"total_tokens":4216,"prompt_tokens":973,"completion_tokens":3243,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":589,"completion_tokens_details":{"reasoning_tokens":3146}},"tokens_in":589,"tokens_out":3243,"duration_ms":23071,"temperature":1.0,"reasoning_tokens":3146,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T21:57:35.982511+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-fit FLAME on the same cohort with an added term for surgical phase or episode order; if the estimated risk curve differs by phase, or if the 0.09 contrast between one 60-minute episode and sixty 1-minute episodes shrinks to zero, the key claim is not supported.","supporting_citations":[{"cited_title":", Hernandez, M","cited_arxiv_id":null,"evidence_quote":"Supplies the cardiac-surgery cohort, high-resolution MAP recordings, and the AKI outcome data used in the application."},{"cited_title":"\\ Brezger, A","cited_arxiv_id":null,"evidence_quote":"Defines the Bayesian P-spline framework and the random-walk prior on spline coefficients adopted for estimation."},{"cited_title":"APACrefauthors \\ 2012","cited_arxiv_id":null,"evidence_quote":"Defines the KDIGO criteria used to classify acute kidney injury, the outcome in the application."},{"cited_title":"\\ APSF Hemodynamic Instability Writing Group","cited_arxiv_id":null,"evidence_quote":"Establishes the clinical relevance of intraoperative hypotension and the MAP threshold underlying the exposure definition."}],"review_version":1}