{"id":"635fe1e2-64ae-4502-9bcc-556d92f734ed","arxiv_id":"2506.08231","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The VALID framework combines human abstraction benchmarks, automated consistency checks, and replication analyses to evaluate LLM-extracted EHR data quality.","lead":"This paper proposes a three-part framework, called VALID, for checking whether data extracted from electronic health records by large language models is accurate and reliable. It matters because more companies and regulators are relying on AI-extracted patient data from EHRs to make oncology research and treatment decisions.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central claim rests on an unverified human-abstraction reference standard; no noise-sensitivity check shows the three-pillar evaluation distinguishes model error from label error.","rationale":"The reader's weakest assumption is also the load-bearing point: the framework's performance metrics and relative comparisons are only meaningful if the human-abstraction reference standard has known and sufficient quality. I find the paper self-consistent but not empirically demonstrated. The Discussion is transparent about the limitation, and Table 1 shows awareness that adjudication can miss shared human/model errors. However, the paper does not provide a procedure to quantify reference-label noise or a sensitivity check showing that prioritization and fitness-for-purpose conclusions are stable. This is a correctness risk, not a disagreement with consensus, because every numeric output of the framework is mediated by the reference labels. The proposed test would settle the question directly: if the same pipeline run on labels with plausible human error rates yields different variable rankings or different fitness-for-purpose decisions, then the central claim is overstated as stated. Because this concern matches the reader's and the appropriate response is still a conditional acceptance requiring a demonstration, the verdict remains UNCHANGED.","tokens_in":10868,"tokens_out":4380,"duration_ms":55337,"concrete_test":"Construct a test set of EHR notes with a known adjudicated gold standard, then generate corrupted reference labels at error rates from 0% to 10%. Run the full VALID pipeline on each corrupted version: variable-level metrics, LLM-minus-human relative differences, the automated verification checks, and one replication analysis. Compare the resulting variable-improvement rankings and the binary fitness-for-purpose decision with the gold-standard run. If the rankings or decision flip at error rates at or below the abstractor disagreement rates shown in Table 2, the framework's load-bearing assumption about reference-standard quality fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The VALID framework's central claim is that its three pillars reliably identify variables most needing improvement, latent errors, and dataset fitness-for-purpose. This requires the reference standard used for variable-level benchmarking and internal replication to be accurate enough that measured differences reflect model behavior rather than label noise. The Discussion acknowledges this: 'effective evaluation of model performance against human abstraction is only meaningful when the reference data is of known and sufficient quality.' Table 1's reference-standard methods do not fully solve the problem. Duplicate abstraction uses a second abstractor's labels as if they were truth, so reference noise remains human noise. Double and triple adjudication resolve disagreements but cannot detect systematic cases where the model and abstractors agree on the same wrong value; Table 1 itself notes that double adjudication metrics 'may be underestimated in the case where the model and abstractor agree on the wrong answer.' Because the framework prioritizes variables using relative LLM-minus-human performance (Table 2), reference-label error can change which variables are flagged and whether an error is attributed to the model or to documentation ambiguity. The paper proposes no calibration step, sensitivity analysis, or empirical case study that would show the framework's outputs are stable under realistic reference-label noise. Without such a demonstration, the framework is an internally plausible proposal, but its central reliability claim is not established.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces VALID, a framework for evaluating the accuracy and reliability of clinical data extracted from electronic health records (EHRs) by large language models (LLMs) and other machine learning models. The framework has three pillars: variable-level performance metrics benchmarked against expert human abstraction, automated verification checks for internal consistency and plausibility, and replication/benchmarking analyses against internal or external reference datasets. It also describes how these components can be applied for bias assessment by stratifying metrics across demographic and clinical subgroups. The authors position the paper as filling a gap in existing RWD and AI quality-assurance guidance, and they draw on their experience at Flatiron Health. The manuscript is a proposal: it contains no empirical data, no case study, and no numerical results beyond illustrative examples in Tables 2 and 3. The discussion acknowledges key limitations, including the need for a high-quality reference standard and the resource constraints of adjudication.","tokens_in":11080,"tokens_out":4153,"duration_ms":55076,"significance":"If the framework were validated, it would provide a practical, integrated approach to quality assurance for LLM-extracted RWD in oncology, which is a timely and important problem. The paper's strengths are its clear articulation of three complementary pillars, the emphasis on benchmarking relative to human abstraction, the acknowledgment of bias assessment via stratification, and a transparent discussion of limitations. The authors honestly state that the framework is a proposal and that reference-data quality is a prerequisite. However, the central claim that this framework 'ensures reliability' or enables detection of latent errors is not evidenced by any empirical demonstration. The manuscript is best understood as a structured position statement rather than a validated method. As such, its significance depends on follow-up work that applies the framework to real datasets and tests its operating characteristics.","major_comments":[{"comment":"The central claim that the framework enables identification of variables most in need of improvement, systematic detection of latent errors, and confirmation of dataset fitness-for-purpose is not supported by any empirical data or case study; Tables 2 and 3 are explicitly illustrative. The reader cannot tell whether the three pillars jointly deliver the claimed benefits, how the outputs are interpreted in practice, or whether the workflow is feasible at scale. Please either add at least one worked example on a real or synthetic LLM-extracted dataset, with results from all three pillars, or revise the abstract and title to describe a proposed framework rather than one that 'ensures reliability.'","section":"Abstract; §2; Tables 2–3"},{"comment":"Duplicate abstraction defines Abstractor 2 as the reference standard for both the LLM and Abstractor 1, so the 'abstraction performance' column in Table 2 is effectively a measure of inter-abstractor agreement, not accuracy against an external truth. Double adjudication resolves disagreements but cannot detect systematic errors where the model and abstractor agree on the same wrong value, as the paper notes. The Discussion states that reference data must be 'of known and sufficient quality' but provides no method to verify that quality and no sensitivity analysis. Because variable prioritization is based on relative LLM-minus-human performance, reference-label noise could change the ranking of variables or the attribution of errors. Please add a calibration step, a small adjudicated gold-standard set, or a simulation under plausible label-noise rates to show that the framework's outputs are stable.","section":"§2, Table 1"},{"comment":"The verification checks in Table 3 are presented as qualitative flags ('may reflect', 'may be enriched') with no thresholds, no operating characteristics, and no criteria for deciding which conflicts require chart adjudication. Without a measure of sensitivity/specificity or at least a worked example of how flagged rates are compared across variables and cohorts, the claim that these checks 'systematically detect latent errors' is not substantiated. Please define how check results are aggregated and what action rules are triggered, or temper the claim to state that these are illustrative candidate checks.","section":"§3, Table 3"}],"minor_comments":[{"comment":"The abbreviation 'HER2' is expanded incorrectly as 'human estrogen receptor 2' in both table footnotes; the correct expansion is 'human epidermal growth factor receptor 2', as used in the main text.","section":"Table 2 footnote and Table 3 abbreviations"},{"comment":"The check 'Patients should not have both a positive and negative gBRCA 1 result' should be written as 'gBRCA1/2 result' for clarity and technical accuracy.","section":"Table 3, patient-level checks"},{"comment":"Figure 1 is referenced in the text as 'Figure 1' but the caption is missing and no legend describes the three components and their interconnections; please add a descriptive caption.","section":"Figure 1"},{"comment":"Reference 10 (the HELM project) lacks full bibliographic details (no authors, no arXiv number or report number), which hinders readers from locating the source.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is essentially a well-structured white paper from Flatiron Health, with heavy reliance on the authors' own prior work (references 11–25). That is not disqualifying, but it means the framework is presented as the consolidated practice of a single organization, and the lack of independent validation or a worked example limits its generality. The journal should consider whether a proposal-only contribution without empirical demonstration meets its scope for methods papers. If accepted, the authors should be encouraged to provide a supplementary case study or a synthetic-data demonstration to make the framework concrete."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a well-written, practical framework proposal from a group that has actually shipped ML-extracted oncology real-world data. The genuinely new piece is the relative performance metric: comparing LLM accuracy against expert human abstractor disagreement rates, so a low F1 on an ambiguous variable is read against how often two humans disagree. That is a real idea, and it makes the paper useful even before any empirical validation arrives.\n\nThe three-pillar structure is coherent, and the tables make it concrete. I also credit the authors for being explicit, in the Discussion, that the whole framework is meaningful only if the reference standard is of known and sufficient quality, and that double adjudication can underestimate errors when model and abstractor agree on the same wrong value. The bias-assessment section is a sensible extension rather than an afterthought. The citation pattern is heavy on the group's own prior work, but that work is directly relevant and the references are used to ground the proposal, not to claim external validation.\n\nThe soft spots are real but proportionate. There is no empirical demonstration: Tables 2 and 3 use illustrative numbers, not measured results. The central reliability claim—that the framework reliably identifies variables needing improvement, latent errors, and fitness-for-purpose—is plausible but unproven. The stress-test note is on point: the reference-standard problem is acknowledged but not addressed with a calibration step or sensitivity analysis. Duplicate abstraction uses a second abstractor's labels as truth; adjudication resolves disagreements but not cases where everyone agrees on the wrong value. Without a noise-sensitivity check or a case study, the relative metric could flag the wrong variables under realistic label noise. That said, the paper is transparent that it is a proposal, and the limitation section says the needed work is to build consensus and demonstrate.\n\nWho this is for: anyone building or evaluating LLM-based EHR extraction pipelines, especially in oncology, and regulators or industry groups looking for a structured QA approach. A serious referee should engage with it, ask for a demonstration (even a retrospective case study on one or two variables with a sensitivity analysis around reference-label noise), and then it could be a useful contribution to the RWD literature.","headline":"A clear, honest framework proposal for validating LLM-extracted EHR data, but the core reliability claim waits on a case study and on the reference-standard noise problem.","tokens_in":11624,"tokens_out":1853,"would_cite":true,"duration_ms":21608,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A three-pillar framework evaluates whether AI-extracted EHR data is accurate enough for oncology research.","keywords":["large language models","electronic health records","real-world data","data quality","oncology","accuracy assessment","replication analysis","bias assessment"],"falsifier":"Take an LLM-extracted oncology dataset that passes all three pillars, commission a second independent expert abstraction with full adjudication on a random sample, and compare variable-level precision against the framework's reported metrics; if the independently measured error rate is materially higher, the reference-standard assumption has failed and the fit-for-purpose conclusion would be overturned.","tokens_in":10694,"feed_emoji":"🩺","tokens_out":6538,"duration_ms":69983,"temperature":0.7,"pith_summary":"The paper argues that accuracy of LLM-extracted clinical data cannot be judged by conventional aggregate test-set metrics alone. It proposes the VALID framework, which combines variable-level performance benchmarking against expert human abstraction, automated verification checks for internal consistency and plausibility, and replication analyses that rerun established clinical findings on the extracted dataset. The central move is to measure LLM performance relative to human abstractor performance rather than only in absolute terms, so ambiguous or subjectively documented concepts are not mistaken for model failures. If correct, this gives data producers and regulators a practical, transparent standard for deciding when AI-extracted real-world oncology data is fit for research and regulatory use.","feed_headline":"Three-pillar framework vets AI-extracted EHR data for research","feed_subtitle":"Benchmarks LLMs against human abstractors, flags latent errors, and replicates known outcomes to judge fitness for purpose.","key_machinery":"The load-bearing mechanism is the relative performance difference between LLM and expert human abstraction, computed on a held-out test set against a common reference standard created by duplicate abstraction, double adjudication, or triple adjudication. This human benchmark anchors interpretation: a large negative gap flags a variable needing model development, a small gap on an ambiguous concept shows the task itself is hard, and a positive gap indicates the LLM may exceed human abstractors. Around this core, automated verification checks act as a full-data proxy for accuracy, and replication analyses test whether errors compound across variables in ways that bias cohort-level conclusions.","core_discovery":"The authors propose that a reliable LLM/ML-extracted oncology dataset is one that passes three linked tests: (1) per-variable recall, precision, F1, and completeness measured against expert human abstraction, with the gap between model and human metrics reported explicitly; (2) automated conformance, plausibility, and consistency checks that flag clinically illogical or internally contradictory records across the full dataset; and (3) replication of cohort distributions, treatment patterns, and outcomes against an internal human-abstracted reference or external benchmarks such as cancer-registry data. The framework treats end-to-end derived variables (e.g., triple-negative breast cancer status at metastatic diagnosis) as first-class evaluation targets, since compounding small errors across component variables can invalidate a research cohort even when each variable looks good. It extends to bias assessment by stratifying all three components by demographic and clinical subgroups, arguing that a divergence between model and human metrics across subgroups reveals differential model errors rather than case-mix differences.","pith_inferences":["Beyond the paper, the relative-performance logic suggests a concrete rework rule: invest in a variable only when the LLM-human gap exceeds the human-human disagreement rate for that variable, since smaller gaps are within the noise of expert annotation.","Beyond the paper, the verification checks could be embedded prospectively into the extraction pipeline as release gates, so a build that fails cohort-level plausibility checks is never delivered.","Beyond the paper, the same three pillars extend naturally to non-oncology indications and to multi-language EHR documents, where documentation variability will stress the human-reference assumption most.","Beyond the paper, the missing absolute thresholds could be derived empirically by calibrating variable-level LLM-human gaps against replication concordance across many datasets."],"forward_implications":["Datasets produced by LLM extraction can be certified as fit-for-purpose only after passing all three pillars, not merely by high F1 on a test set.","Reporting LLM-minus-human performance makes quality thresholds interpretable across variables and datasets, so development resources can target variables where the model lags human abstractors.","End-to-end metrics for derived variables such as biomarker status at diagnosis reveal error compounding that per-variable metrics miss.","Stratifying metrics, verification checks, and replication analyses by demographic subgroups turns the same machinery into a bias assessment tool.","Re-running these assessments after each model or pipeline change is necessary because LLM versions and prompts drift over time."],"supporting_citations":[{"why":"Supplies evidence that LLMs can produce unstable datapoints even with unchanged inputs, motivating the need for reliability checks.","marker":"[1]"},{"why":"Documents hallucination and other failure modes of large language models in biomedicine, motivating verification checks.","marker":"[2]"},{"why":"Provides regulatory guidance defining transparency and validation expectations for AI/ML tools, which this framework operationalizes.","marker":"[5]"},{"why":"Gives a good-practices checklist for machine learning methods in health economics that covers transparency but not LLM extraction accuracy.","marker":"[6]"},{"why":"Gives a checklist for assessing EHR-derived real-world data quality that this framework extends with accuracy methods.","marker":"[7]"},{"why":"Demonstrates reliable LLM/ML extraction of EHR variables and proposes an accuracy-assessment approach that this paper builds on.","marker":"[8]"},{"why":"Provides a method for evaluating clinical AI without ground-truth labels, motivating the need for reference-standard benchmarking.","marker":"[9]"},{"why":"Shows that ML-extracted EHR data can replicate established oncology real-world findings, grounding the replication pillar.","marker":"[13]"},{"why":"Is a prior research-centric evaluation framework for ML-extracted real-world data that this framework extends.","marker":"[25]"},{"why":"Is a prior description of data-quality dimensions and verification checks that the automated checks component adapts.","marker":"[26]"}],"fun_headline_variants":["VALID framework triple-checks LLM-extracted EHR data","LLM-extracted EHR data validated by human, logic, replication tests","Three-part framework vets AI-extracted oncology data reliability","Bias-aware triple validation for LLM-extracted EHR data","Framework ensures reliable LLM-extracted EHR data with three tests"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The framework assumes that expert human abstraction produces a reference standard of known and sufficient quality; the paper itself states that if this fails, every performance metric and downstream conclusion becomes unreliable.","fun_headline_variants_meta":{"raw":{"variants":["VALID framework triple-checks LLM-extracted EHR data","LLM-extracted EHR data validated by human, logic, replication tests","Three-part framework vets AI-extracted oncology data reliability","Bias-aware triple validation for LLM-extracted EHR data","Framework ensures reliable LLM-extracted EHR data with three tests"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000808,"raw_usage":{"total_tokens":3565,"prompt_tokens":981,"completion_tokens":2584,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":597,"completion_tokens_details":{"reasoning_tokens":2498}},"tokens_in":597,"tokens_out":2584,"duration_ms":22292,"temperature":1.0,"reasoning_tokens":2498,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T05:16:07.150864+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take an LLM-extracted oncology dataset that passes all three pillars, commission a second independent expert abstraction with full adjudication on a random sample, and compare variable-level precision against the framework's reported metrics; if the independently measured error rate is materially higher, the reference-standard assumption has failed and the fit-for-purpose conclusion would be overturned.","supporting_citations":[{"cited_title":"good enough","cited_arxiv_id":null,"evidence_quote":"Supplies evidence that LLMs can produce unstable datapoints even with unchanged inputs, motivating the need for reliability checks."},{"cited_title":"Use of Generative AI to Identify Helmet Status Among Patients With Micromobility-Related Injuries From Unstructured Clinical Notes","cited_arxiv_id":null,"evidence_quote":"Documents hallucination and other failure modes of large language models in biomedicine, motivating verification checks."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Gives a good-practices checklist for machine learning methods in health economics that covers transparency but not LLM extraction accuracy."},{"cited_title":"Machine Learning Methods in Health Economics and Outcomes Research—The PALISADE Checklist: A Good Practices Report of an ISPOR Task Force","cited_arxiv_id":null,"evidence_quote":"Gives a checklist for assessing EHR-derived real-world data quality that this framework extends with accuracy methods."},{"cited_title":"Assessing Real-World Data From Electronic Health Records for Health Technology Assessment: The SUITABILITY Checklist: A Good Practices Report of an ISPOR Task Force","cited_arxiv_id":null,"evidence_quote":"Demonstrates reliable LLM/ML extraction of EHR variables and proposes an accuracy-assessment approach that this paper builds on."},{"cited_title":"Implementing Accuracy, Completeness, and Traceability for Data Reliability","cited_arxiv_id":null,"evidence_quote":"Provides a method for evaluating clinical AI without ground-truth labels, motivating the need for reference-standard benchmarking."},{"cited_title":"Large Language Model Extraction of PD-L1 Biomarker Testing Details From Electronic Health Records","cited_arxiv_id":null,"evidence_quote":"Shows that ML-extracted EHR data can replicate established oncology real-world findings, grounding the replication pillar."},{"cited_title":"Real-World (Rw) CtDNA Testing Trends and Associated Outcomes in Patients (Pts) with Early Stage Breast Cancer (EBC)","cited_arxiv_id":null,"evidence_quote":"Is a prior research-centric evaluation framework for ML-extracted real-world data that this framework extends."},{"cited_title":"Considerations for the Use of Machine Learning Extracted Real-World Data to Support Evidence Generation: A Research-Centric Evaluation Framework","cited_arxiv_id":null,"evidence_quote":"Is a prior description of data-quality dimensions and verification checks that the automated checks component adapts."}],"review_version":1}