{"id":"275e4276-7237-4b03-a368-681e3d05685e","arxiv_id":"2507.22943","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A process combining NLP-aided annotation and multi-wave adaptive sampling cuts chart review time by roughly half and would have skipped 77% of charts in a validation study of an intentional self-harm algorithm.","lead":"This paper proposes a chart review process that combines natural language processing to highlight relevant text and multi-wave adaptive sampling with a stopping rule to reduce the number of charts human reviewers must read.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The NLP time-savings claim—the abstract's 40% reduction—rests on a small, annotator-confounded comparison, and the paper itself reports 48%, so the headline efficiency gain is not yet established.","rationale":"I read the paper as a methods/feasibility contribution whose central claim is that the proposed process expedites chart validation. The multi-wave adaptive sampling result is credible: the stopping rule was pre-specified, the futility decision after 120/530 charts was borne out by the full-sample PPV (0.6034 vs 0.6343, both upper bounds below 0.75), and the confidence intervals, though wider at stopping, support the same qualitative conclusion. The weaker pillar is the NLP time-saving claim. The reported experiment is small (40 charts), the two annotator conditions are not crossed within annotator, no inferential statistics are given, and the abstract's 40% contradicts the Discussion's 48%. This is not an accusation of misconduct; it is a straightforward internal-validity problem that can be checked from the posted data. The reader's chosen weakest assumption—generalizing PPV from claims-positive patients with MGB contact to those without—is real and is honestly disclosed, but it concerns the generalizability of the case-study estimates rather than the internal validity of the expediting claim. I therefore partially agree with the reader: the same overall conditional verdict is appropriate, but I would move the time-reduction comparison to the front as the most load-bearing concern.","tokens_in":11930,"tokens_out":5273,"duration_ms":64484,"concrete_test":"Download the OSF de-identified annotation/timing data and reconstruct each chart's two review times with annotator identifiers, then fit a paired model (chart random effect, annotator fixed effect) or a Wilcoxon signed-rank test on within-chart differences between NLP and no-NLP conditions. If the adjusted median reduction is not statistically significant, or if it is not close to the reported 40%/48%, the abstract's headline should be revised and the NLP-aided component should be reported as exploratory.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim has two load-bearing pillars: (i) NLP assistance reduces per-chart review time, and (ii) the multi-wave stopping rule avoids 77% of reviews. Pillar (ii) is well illustrated by the case study: futility was declared after 12 batches (120/530 charts) and the final full-sample PPV credible interval remains below the 0.75 threshold. Pillar (i), however, is not empirically secured. The comparison in the Results section is based on 40 charts in which \"each annotator additionally reviewed a sample of 20 charts without NLP highlighted terms that had previously been reviewed by the other annotator with NLP highlighted terms.\" Because the same chart is never reviewed with and without NLP by the same annotator, annotator speed is perfectly confounded with NLP condition across the paired comparisons. No paired test, confidence interval, or mixed model is reported, and the only evidence that the annotators are exchangeable is a statement that training times were \"similar.\" In addition, the abstract reports a 40% reduction while the Discussion reports 48% for the same 6.0 vs 11.4 minute medians. This internal inconsistency signals that the headline statistic has not been carefully verified. Because the abstract's first empirical result—\"reduced the time spent on review per chart by 40%\"—directly supports the expedited-review claim, this is the most load-bearing weakness. The PPV extrapolation for patients without MGB contact is a genuine limitation, but it is explicitly acknowledged and does not threaten the process-level feasibility claim as directly.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a two-part process to expedite validation of code-based algorithms in large database studies: (1) NLP highlighting of relevant text in EHR notes to speed up manual chart review, and (2) multi-wave adaptive sampling with a Bayesian stopping rule to reduce the number of charts reviewed. The method is illustrated in a case study validating a claims-based ICD-10-CM screening algorithm for intentional self-harm among 62,129 patients with obesity in Mass General Brigham (MGB) linked claims-EHR data. The authors report that NLP assistance reduced per-chart review time by 40% (abstract) or 48% (discussion) and that a predefined futility stopping rule, met after 120 of 530 charts, would have prevented review of 77% of charts with limited compromise to the PPV estimate.","tokens_in":12191,"tokens_out":7704,"duration_ms":91543,"significance":"If the efficiency gains are real, the proposed process could substantially lower the cost of chart validation and make routine algorithm validation more feasible, which would strengthen quantitative bias analyses in database studies. The paper has notable strengths: the proposed workflow is clearly specified; the full-sample comparison of PPV at stopping versus after all 530 charts supports the soundness of the stopping-rule decision; the open-source CORA tool, protocol, analytic code, and de-identified annotation inputs are made available; and the authors explicitly acknowledge key limitations such as the untestable assumption about claims-positive patients without MGB contact. The main gap is that the NLP time-savings claim, which is one of the two pillars of the expedited-review argument, is not statistically established and is internally inconsistent between the abstract and the discussion.","major_comments":[{"comment":"The NLP time-reduction claim is not supported by a valid statistical comparison. The design is described as follows: each annotator reviewed 20 charts without NLP highlights that had previously been reviewed by the other annotator with NLP highlights, giving 40 charts with recorded timing in both conditions. Because each chart is reviewed by a different annotator under the two conditions, annotator identity is perfectly confounded with NLP condition within each chart. No paired test, confidence interval, or mixed model with annotator fixed effects is reported, and the statement that the two annotators had 'similar timing' during training is not a formal test of exchangeability. Furthermore, the abstract's 40% reduction appears to be obtained by comparing the overall NLP median of 7.0 minutes with the no-NLP subset median of 11.4 minutes, whereas the Results report within-subset medians of 6.0 minutes (NLP) versus 11.4 minutes (no NLP), which give a 48% reduction, the number used in the Discussion. This is an apples-to-oranges comparison and an internal inconsistency. The authors should present a paired analysis that separates annotator effects from NLP effects, report an effect estimate with a confidence interval or a posterior interval, and correct the abstract's percentage so that it matches the analysis actually performed.","section":"Chart Review Process, step 5; Results"},{"comment":"The PPV estimate used for the case study is based on claims-positive patients with MGB healthcare contact near the outcome date (Figure 2, boxes f/g), and the authors assume that the proportion of true positives is the same among claims-positive patients without MGB contact (boxes d/e). This assumption is explicitly acknowledged as necessary, but it is load-bearing for the generalizability of the case-study PPV and for the stopping-rule decision, since the reviewed sample excludes the d/e subgroup entirely. The paper provides no sensitivity analysis or external benchmark to bound the possible bias. The authors should either present a sensitivity analysis under plausible scenarios (e.g., differential PPV in d/e) or clearly restrict the PPV claim to the subpopulation with linked EHR contact, and temper the conclusion that the algorithm's PPV is approximately 0.63.","section":"Figure 2; Discussion, Limitations"}],"minor_comments":[{"comment":"Please clarify which median (7.0 vs 6.0 minutes) corresponds to which set of charts, and explicitly derive the percentage reduction reported in the abstract (40%) from the results; as written, the abstract number cannot be reproduced from the within-subset comparison.","section":"Results"},{"comment":"The statement 'reduced the amount of time to annotate charts by 48%' should be reconciled with the abstract's 40%; if the correct within-subset estimate is 48%, the abstract should be corrected, and if the intended comparison uses the overall median, the denominator and sample should be stated explicitly.","section":"Discussion"},{"comment":"The claim that the two annotators had 'similar timing in terms of review' during training should be supported by quantitative summaries (medians, ranges, or a formal test), since this statement is used to argue that annotator speed does not confound the NLP comparison.","section":"Chart Review Process, step 5"},{"comment":"The report of Cohen's kappa equal to 100% after 30 charts would be more informative with the raw number of agreements and a confidence interval; a perfect kappa in a small sample is not strong evidence of interchangeability.","section":"Results"},{"comment":"The sensitivity estimate changes from 0.93 at the stopping point to 0.82 after the full sample, with both confidence intervals very wide; please state explicitly whether this difference is within the expected sampling variability, given that the stopping rule was based only on PPV.","section":"Table 2"},{"comment":"Reference 10 (Wright et al., JAMIA 2011) does not appear to describe multi-wave adaptive sampling and may be a citation error; please verify that this reference supports the statement about multi-wave sampling methods.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper is a potentially useful methods contribution, and the stopping-rule demonstration is strengthened by the full-sample comparison. The central obstacle is the NLP time-savings claim: the design is confounded, no statistical analysis is reported, and the abstract and discussion give different percentages. This issue is fixable if raw timing data per annotator and per chart are available; if not, the abstract should be substantially softened. The PPV extrapolation assumption is explicitly disclosed, but a sensitivity analysis would strengthen the case study. I would send back for major revision rather than reject, as the core methodological framework is sound and the main problem is localized to the time-savings evidence and its reporting."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: the adaptive sampling half of this paper is well done and worth referee time; the NLP time-savings claim is not yet established. The case study's 77% chart reduction follows from a clearly prespecified futility rule, and the full-sample follow-up shows the early estimate was not misleading. That is a genuine, useful demonstration.\n\nWhat's new: the integration of NLP highlighting with multi-wave adaptive sampling into a practical workflow, with protocol, code, and de-identified annotation data on OSF. The paper is transparent about the strata assumptions, including the untestable claim that PPV among claims-positive patients without MGB contact equals that among patients with contact. That limitation is real but acknowledged and does not sink the process-level feasibility result.\n\nThe soft spots are in the NLP time comparison. The abstract's 40% reduction and the Discussion's 48% for the same medians (6.0 vs 11.4 minutes) are inconsistent. More importantly, the comparison is based on 40 charts where each annotator reviewed a chart with NLP and a different chart without. Annotator speed is confounded with NLP condition; no paired analysis, mixed model, or confidence interval is reported. The fact that training times were 'similar' is weak evidence of exchangeability. So the headline efficiency gain is not empirically secure. This should be fixed before the paper is used to justify routine adoption. The stress-test note is right about this, and I agree with the conditional verdict.\n\nThe PPV generalization assumption is the weakest conceptual point, but the authors call it out; I wouldn't hold that against the paper beyond noting it. The remaining analysis, including inverse-weighted sensitivity/specificity/NPV, is clearly described and the confidence intervals are wide as expected with small samples.\n\nWho this is for: applied pharmacoepidemiologists and others doing chart validation studies, plus methodologists working on adaptive sampling. A serious referee should engage. The NLP time claim needs reanalysis or softened language, but the adaptive sampling component and the workflow description are solid. I'd encourage the editor to send it out.","headline":"Adaptive sampling demonstration is solid; the NLP time-savings claim needs reanalysis before it can be used.","tokens_in":12847,"tokens_out":1587,"would_cite":true,"duration_ms":18313,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that pairing NLP-assisted annotation with multi-wave adaptive sampling and Bayesian stopping rules can cut per-chart review time by about 40% and avoid reviewing most charts in a validation study, as demonstrated in a…","keywords":["chart review","natural language processing","multi-wave adaptive sampling","Bayesian stopping rule","positive predictive value","validation study","intentional self-harm","claims database"],"falsifier":"One could settle the central claim by linking out-of-network records or state death data for the unreviewed claims-positive patients with no health-system contact and computing their true-positive rate; if it differs materially from the 0.60 PPV seen in reviewed charts, the generalizable PPV and any bias adjustment built on it fail. A second check would annotate a random sample of the claims-negative, EHR-negative group to test the assumption that they are all true negatives.","tokens_in":11719,"feed_emoji":"⏱️","tokens_out":6565,"duration_ms":76235,"temperature":0.7,"pith_summary":"This paper proposes and demonstrates an expedited chart validation process for code-based algorithms used in large healthcare databases. The process combines two efficiency mechanisms: natural language processing (NLP) that highlights relevant portions of electronic health record notes, reducing the time a human annotator spends per chart, and multi-wave adaptive sampling with a predefined Bayesian stopping rule, which halts review once the primary performance metric is estimated with enough precision. In the case study, validating a claims-based screening algorithm for intentional self-harm in patients with obesity, the authors report that NLP-assisted review took about 40% less time per chart, and that applying the stopping rule would have prevented review of 77% of the 530 sampled charts. The estimated positive predictive value at the stopping point was 0.60 with a 95% credible interval of (0.47, 0.72), essentially unchanged after full review, supporting the claim that precision loss is limited. If the process holds up, validation studies could become routine enough to support quantitative bias analyses in database research.","feed_headline":"Chart validation gets 40% faster with NLP plus adaptive sampling","feed_subtitle":"A Bayesian stopping rule would skip 77% of charts with little precision loss, the self-harm case shows.","key_machinery":"The process rests on three linked components. First, study participants are stratified by whether they are claims-positive for the outcome, whether their EHR notes contain a relevant UMLS concept unique identifier (CUI) found by MetaMap, and whether state death records indicate suicide; patients negative on both claims and EHR are treated as true negatives, and death-record suicides are treated as true positives without annotation. Second, an open-source NLP-aided annotation tool called CORA loads linked claims and notes in chronological order and highlights text matching the self-harm CUI list, directing annotator attention and cutting review time. Third, sequential batches of ten charts, five claims-positive and five claims-negative, are annotated, and a $\\mathrm{Beta}(1,1)$ prior is updated with the observed binomial counts to form a posterior distribution for the positive predictive value; the stopping rule stops for success when the lower bound of the 95% credible interval exceeds 0.75 and for futility when the upper bound falls below 0.75. Inverse sampling weights convert the stratified sample into estimates of negative predictive value, sensitivity, and specificity for the full cohort.","core_discovery":"On its own terms, the paper establishes that an NLP-assisted, adaptively sampled chart review can validate a claims-based outcome algorithm with substantially less human effort than full manual review. The central empirical claims are that annotators using the NLP-highlighting interface took a median 6.0 minutes per chart versus 11.4 minutes without it, a roughly 40 to 48 percent reduction; that cumulative Bayesian estimates of positive predictive value met a predefined futility stopping rule after 12 batches of 10 charts, only 23% of the full sample; and that the inverse-weighted sensitivity, specificity, and negative predictive value at that stopping point were consistent with the values after all 530 charts were reviewed, although sensitivity intervals remained wide. The paper also characterizes the reasons behind false-positive and false-negative annotations, showing that many false positives arose from mental-status assessments recording no suicidal thoughts and from distant past history, and it argues that these gains make routine validation of algorithms for database studies more feasible.","pith_inferences":["The efficiency gain is likely to carry over to other outcomes only when the keyword or CUI list has high recall; an incomplete list would create silent false negatives in the unreviewed 'EHR-negative' stratum, so the process should be accompanied by a sensitivity analysis that annotates a sample of that stratum.","The NLP-aided versus unaided timing comparison used 40 charts reviewed by different annotators, so part of the measured time difference could reflect between-annotator speed; a within-annotator randomized crossover would isolate the NLP effect.","Because the algorithm being validated is fixed during the process, any future adaptive tuning of the code list between waves would need validation on a fresh sample to avoid overfitting to the reviewed charts.","The same machinery could be extended to estimate performance separately by exposure or comparator group, which would make the resulting measurement characteristics more directly usable in exposure-specific quantitative bias analyses."],"forward_implications":["If the method is applied to other code-based algorithms, validation studies that currently require hundreds of hours of chart review could be completed after reviewing only a fraction of the planned charts.","The reported PPV of 0.60 for the intentional self-harm screening algorithm can feed quantitative bias analyses for database studies of self-harm in patients with obesity instead of relying on unvalidated code.","For algorithms that are clearly high-performing or clearly poor, the stopping rule yields the largest savings; gains shrink for algorithms whose performance hovers near the threshold.","The false-positive profile gives concrete code-refinement targets, such as excluding notes that only document absence of suicidal thoughts and distinguishing distant past history from a current event.","The framework is tool-agnostic, so newer NLP approaches, including large language models, can be substituted for the MetaMap-based highlighting without changing the sampling or stopping machinery."],"supporting_citations":[{"why":"Supplies the Bayesian adaptive validation design that motivates multi-wave sampling and stopping rules.","marker":"[9]"},{"why":"Establishes the efficiency of the disproportionate stratified multi-wave sampling strategy used in the case study.","marker":"[17]"},{"why":"MetaMap is used to identify self-harm and suicide concept unique identifiers for strata creation and NLP highlighting.","marker":"[16]"},{"why":"Prior demonstration that NLP can improve efficiency of manual chart abstraction for research.","marker":"[8]"},{"why":"Systematic review defining screening versus outcome algorithms for suicide and providing the PPV benchmark the case study compares against.","marker":"[11]"},{"why":"The Sentinel CIDA query package is used to extract the obesity cohort and the claims-based outcomes.","marker":"[15]"},{"why":"Describes the linked claims-EHR data enterprise that supplies the case study's data.","marker":"[14]"}],"fun_headline_variants":["NLP and adaptive sampling cut chart review time by 40%","Adaptive sampling skips 77% of charts in validation study","Self-harm algorithm validated with 77% fewer charts via NLP","Chart review speedup: NLP plus Bayesian stopping rule"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that claims-positive patients with no contact at the linked health system around the outcome date have the same true-positive rate as those who were seen there, even though the former group cannot be reviewed; if their documentation differs, the reported PPV is biased.","fun_headline_variants_meta":{"raw":{"variants":["NLP and adaptive sampling cut chart review time by 40%","Adaptive sampling skips 77% of charts in validation study","Self-harm algorithm validated with 77% fewer charts via NLP","Chart review speedup: NLP plus Bayesian stopping rule"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000545,"raw_usage":{"total_tokens":2640,"prompt_tokens":1010,"completion_tokens":1630,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":626,"completion_tokens_details":{"reasoning_tokens":1559}},"tokens_in":626,"tokens_out":1630,"duration_ms":11449,"temperature":1.0,"reasoning_tokens":1559,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T14:17:57.904872+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"One could settle the central claim by linking out-of-network records or state death data for the unreviewed claims-positive patients with no health-system contact and computing their true-positive rate; if it differs materially from the 0.60 PPV seen in reviewed charts, the generalizable PPV and any bias adjustment built on it fail. A second check would annotate a random sample of the claims-negative, EHR-negative group to test the assumption that they are all true negatives.","supporting_citations":[{"cited_title":"Adaptive Validation Design: A Bayesian Ap- proach to Validation Substudy Design With Prospective Data Collection - PubMed","cited_arxiv_id":null,"evidence_quote":"Supplies the Bayesian adaptive validation design that motivates multi-wave sampling and stopping rules."},{"cited_title":"Adaptive multi-wave sampling for efficient chart validation","cited_arxiv_id":"2503.06308","evidence_quote":"Establishes the efficiency of the disproportionate stratified multi-wave sampling strategy used in the case study."},{"cited_title":"Using natural language processing to improve ef- ficiency of manual chart abstraction in research: the case of breast cancer recurrence - Pub- Med","cited_arxiv_id":null,"evidence_quote":"Prior demonstration that NLP can improve efficiency of manual chart abstraction for research."},{"cited_title":"A systematic re- view of validated suicide outcome classification in observational studies - PubMed","cited_arxiv_id":null,"evidence_quote":"Systematic review defining screening versus outcome algorithms for suicide and providing the PPV benchmark the case study compares against."},{"cited_title":"Accessed December 17, 2024, Browse Sentinel Documentation / Sentinel Routine Querying Tool Documentation - Sentinel Version Control System","cited_arxiv_id":null,"evidence_quote":"The Sentinel CIDA query package is used to extract the obesity cohort and the claims-based outcomes."},{"cited_title":"The FDA Sentinel Real World Evidence Data Enterprise (RWE‐DE)","cited_arxiv_id":null,"evidence_quote":"Describes the linked claims-EHR data enterprise that supplies the case study's data."}],"review_version":1}