Pith. sign in

REVIEW 3 major objections 4 minor 29 references

Ensuring Reliability of Curated EHR-Derived Data: The Validation of Accuracy for LLM/ML-Extracted Information and Data (VALID) Framework

T0 review · 3 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read A three-pillar framework evaluates whether AI-extracted EHR data is accurate enough for oncology research.

desk verdict A clear, honest framework proposal for validating LLM-extracted EHR data, but the core reliability claim waits on a case study and on the reference-standard noise problem. read the letter →

arxiv 2506.08231 v1 pith:JM2RU62O submitted 2025-06-09 cs.LG cs.AIcs.PF

classification cs.LGcs.AIcs.PF
keywords largelanguagemodelselectronichealthrecordsreal-worlddataqualityoncologyaccuracyassessmentreplicationanalysisbias
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper argues that accuracy of LLM-extracted clinical data cannot be judged by conventional aggregate test-set metrics alone. It proposes the VALID framework, which combines variable-level performance benchmarking against expert human abstraction, automated verification checks for internal consistency and plausibility, and replication analyses that rerun established clinical findings on the extracted dataset. The central move is to measure LLM performance relative to human abstractor performance rather than only in absolute terms, so ambiguous or subjectively documented concepts are not mistaken for model failures. If correct, this gives data producers and regulators a practical, transparent standard for deciding when AI-extracted real-world oncology data is fit for research and regulatory use.

What carries the argument

The load-bearing mechanism is the relative performance difference between LLM and expert human abstraction, computed on a held-out test set against a common reference standard created by duplicate abstraction, double adjudication, or triple adjudication. This human benchmark anchors interpretation: a large negative gap flags a variable needing model development, a small gap on an ambiguous concept shows the task itself is hard, and a positive gap indicates the LLM may exceed human abstractors. Around this core, automated verification checks act as a full-data proxy for accuracy, and replication analyses test whether errors compound across variables in ways that bias cohort-level conclusions.

What would settle it

Take an LLM-extracted oncology dataset that passes all three pillars, commission a second independent expert abstraction with full adjudication on a random sample, and compare variable-level precision against the framework's reported metrics; if the independently measured error rate is materially higher, the reference-standard assumption has failed and the fit-for-purpose conclusion would be overturned.

Watch

Extended reading notes

Core claim

The authors propose that a reliable LLM/ML-extracted oncology dataset is one that passes three linked tests: (1) per-variable recall, precision, F1, and completeness measured against expert human abstraction, with the gap between model and human metrics reported explicitly; (2) automated conformance, plausibility, and consistency checks that flag clinically illogical or internally contradictory records across the full dataset; and (3) replication of cohort distributions, treatment patterns, and outcomes against an internal human-abstracted reference or external benchmarks such as cancer-registry data. The framework treats end-to-end derived variables (e.g., triple-negative breast cancer status at metastatic diagnosis) as first-class evaluation targets, since compounding small errors across component variables can invalidate a research cohort even when each variable looks good. It extends to bias assessment by stratifying all three components by demographic and clinical subgroups, arguing that a divergence between model and human metrics across subgroups reveals differential model errors rather than case-mix differences.

Load-bearing premise

The framework assumes that expert human abstraction produces a reference standard of known and sufficient quality; the paper itself states that if this fails, every performance metric and downstream conclusion becomes unreliable.

Editorial extensions

If this is right

  • Datasets produced by LLM extraction can be certified as fit-for-purpose only after passing all three pillars, not merely by high F1 on a test set.
  • Reporting LLM-minus-human performance makes quality thresholds interpretable across variables and datasets, so development resources can target variables where the model lags human abstractors.
  • End-to-end metrics for derived variables such as biomarker status at diagnosis reveal error compounding that per-variable metrics miss.
  • Stratifying metrics, verification checks, and replication analyses by demographic subgroups turns the same machinery into a bias assessment tool.
  • Re-running these assessments after each model or pipeline change is necessary because LLM versions and prompts drift over time.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper, the relative-performance logic suggests a concrete rework rule: invest in a variable only when the LLM-human gap exceeds the human-human disagreement rate for that variable, since smaller gaps are within the noise of expert annotation.
  • Beyond the paper, the verification checks could be embedded prospectively into the extraction pipeline as release gates, so a build that fails cohort-level plausibility checks is never delivered.
  • Beyond the paper, the same three pillars extend naturally to non-oncology indications and to multi-language EHR documents, where documentation variability will stress the human-reference assumption most.
  • Beyond the paper, the missing absolute thresholds could be derived empirically by calibrating variable-level LLM-human gaps against replication concordance across many datasets.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper introduces VALID, a framework for evaluating the accuracy and reliability of clinical data extracted from electronic health records (EHRs) by large language models (LLMs) and other machine learning models. The framework has three pillars: variable-level performance metrics benchmarked against expert human abstraction, automated verification checks for internal consistency and plausibility, and replication/benchmarking analyses against internal or external reference datasets. It also describes how these components can be applied for bias assessment by stratifying metrics across demographic and clinical subgroups. The authors position the paper as filling a gap in existing RWD and AI quality-assurance guidance, and they draw on their experience at Flatiron Health. The manuscript is a proposal: it contains no empirical data, no case study, and no numerical results beyond illustrative examples in Tables 2 and 3. The discussion acknowledges key limitations, including the need for a high-quality reference standard and the resource constraints of adjudication.

Significance. If the framework were validated, it would provide a practical, integrated approach to quality assurance for LLM-extracted RWD in oncology, which is a timely and important problem. The paper's strengths are its clear articulation of three complementary pillars, the emphasis on benchmarking relative to human abstraction, the acknowledgment of bias assessment via stratification, and a transparent discussion of limitations. The authors honestly state that the framework is a proposal and that reference-data quality is a prerequisite. However, the central claim that this framework 'ensures reliability' or enables detection of latent errors is not evidenced by any empirical demonstration. The manuscript is best understood as a structured position statement rather than a validated method. As such, its significance depends on follow-up work that applies the framework to real datasets and tests its operating characteristics.

major comments (3)
  1. [Abstract; §2; Tables 2–3] The central claim that the framework enables identification of variables most in need of improvement, systematic detection of latent errors, and confirmation of dataset fitness-for-purpose is not supported by any empirical data or case study; Tables 2 and 3 are explicitly illustrative. The reader cannot tell whether the three pillars jointly deliver the claimed benefits, how the outputs are interpreted in practice, or whether the workflow is feasible at scale. Please either add at least one worked example on a real or synthetic LLM-extracted dataset, with results from all three pillars, or revise the abstract and title to describe a proposed framework rather than one that 'ensures reliability.'
  2. [§2, Table 1] Duplicate abstraction defines Abstractor 2 as the reference standard for both the LLM and Abstractor 1, so the 'abstraction performance' column in Table 2 is effectively a measure of inter-abstractor agreement, not accuracy against an external truth. Double adjudication resolves disagreements but cannot detect systematic errors where the model and abstractor agree on the same wrong value, as the paper notes. The Discussion states that reference data must be 'of known and sufficient quality' but provides no method to verify that quality and no sensitivity analysis. Because variable prioritization is based on relative LLM-minus-human performance, reference-label noise could change the ranking of variables or the attribution of errors. Please add a calibration step, a small adjudicated gold-standard set, or a simulation under plausible label-noise rates to show that the framework's outputs are stable.
  3. [§3, Table 3] The verification checks in Table 3 are presented as qualitative flags ('may reflect', 'may be enriched') with no thresholds, no operating characteristics, and no criteria for deciding which conflicts require chart adjudication. Without a measure of sensitivity/specificity or at least a worked example of how flagged rates are compared across variables and cohorts, the claim that these checks 'systematically detect latent errors' is not substantiated. Please define how check results are aggregated and what action rules are triggered, or temper the claim to state that these are illustrative candidate checks.
minor comments (4)
  1. [Table 2 footnote and Table 3 abbreviations] The abbreviation 'HER2' is expanded incorrectly as 'human estrogen receptor 2' in both table footnotes; the correct expansion is 'human epidermal growth factor receptor 2', as used in the main text.
  2. [Table 3, patient-level checks] The check 'Patients should not have both a positive and negative gBRCA 1 result' should be written as 'gBRCA1/2 result' for clarity and technical accuracy.
  3. [Figure 1] Figure 1 is referenced in the text as 'Figure 1' but the caption is missing and no legend describes the three components and their interconnections; please add a descriptive caption.
  4. [References] Reference 10 (the HELM project) lacks full bibliographic details (no authors, no arXiv number or report number), which hinders readers from locating the source.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: VALID is a procedural QA framework, not a derivation whose outputs reduce to its inputs.

full rationale

The paper proposes a three-pillar framework (variable-level metrics, verification checks, replication) for assessing LLM-extracted EHR data. There is no fitted parameter presented as a prediction and no equation that equates an output with an input by construction. The use of expert human abstraction as a reference standard is an explicitly acknowledged domain assumption ('effective evaluation of model performance against human abstraction is only meaningful when the reference data is of known and sufficient quality'), not a hidden equivocation or a self-definitional move. The illustrative metrics in Table 2 are hypothetical examples, not fitted results. Self-citations to prior Flatiron work ([11-24], [25], [26]) provide context, replication examples, and verification-check sources; no central claim is forced by a self-citation chain or by an imported uniqueness theorem. The framework is self-contained as a set of recommendations, so the correct finding is no significant circularity.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The framework introduces no free parameters or invented entities. It rests on assumptions about the validity of human abstraction as a reference standard and about the interpretability of automated verification checks.

assumptions (3)
  • domain assumption Expert human abstraction provides a valid reference standard for LLM-extracted clinical data.
    Section 1 and Table 1 treat human-abstracted labels as the ground truth against which LLM and abstractor performance are measured.
  • domain assumption Reference data from human abstraction has sufficient and known quality for benchmarking.
    The Discussion states that effective evaluation is only meaningful when reference data is of known and sufficient quality; this is a load-bearing premise.
  • domain assumption Automated verification checks serve as a proxy for accuracy and can identify model errors.
    Section 2 states that these checks 'serve as a proxy for accuracy' and implies that certain data conflicts are more likely to be model errors than real clinical scenarios.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Ensuring Reliability of Curated EHR-Derived Data: The Validation of Accuracy for LLM/ML-Extracted Information and Data (VALID) Framework." pith.science (2026). https://pith.science/paper/JM2RU62O

@misc{pith2026250608231,
  author       = {Pith},
  title        = {Pith review of: Ensuring Reliability of Curated EHR-Derived Data: The Validation of Accuracy for LLM/ML-Extracted Information and Data (VALID) Framework},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/JM2RU62O}},
  note         = {Machine review of arXiv:2506.08231}
}
read the original abstract

Large language models (LLMs) are increasingly used to extract clinical data from electronic health records (EHRs), offering significant improvements in scalability and efficiency for real-world data (RWD) curation in oncology. However, the adoption of LLMs introduces new challenges in ensuring the reliability, accuracy, and fairness of extracted data, which are essential for research, regulatory, and clinical applications. Existing quality assurance frameworks for RWD and artificial intelligence do not fully address the unique error modes and complexities associated with LLM-extracted data. In this paper, we propose a comprehensive framework for evaluating the quality of clinical data extracted by LLMs. The framework integrates variable-level performance benchmarking against expert human abstraction, automated verification checks for internal consistency and plausibility, and replication analyses comparing LLM-extracted data to human-abstracted datasets or external standards. This multidimensional approach enables the identification of variables most in need of improvement, systematic detection of latent errors, and confirmation of dataset fitness-for-purpose in real-world research. Additionally, the framework supports bias assessment by stratifying metrics across demographic subgroups. By providing a rigorous and transparent method for assessing LLM-extracted RWD, this framework advances industry standards and supports the trustworthy use of AI-powered evidence generation in oncology research and practice.

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

29 extracted references · 25 canonical work pages

  1. [1]

    good enough

    Ensuring Reliability of Curated EHR-Derived Data: The Validation of Accuracy for LLM/ML-Extracted Information and Data (VALID) Framework 1Melissa Estevez, MS; 2Nisha Singh, MS; 2Lauren Dyson, MS; 2Blythe Adamson, PhD, MPH; 2Qianyu Yuan, PhD; 2Megan W. Hildner; 2Erin Fidyk, MSN, MBA; 2Olive Mbah, PhD, MHS; 2Farhad Khan, MSPH; 3Kathi Seidl-Rathkopf, PhD; 2A...

  2. [2]

    Use of Generative AI to Identify Helmet Status Among Patients With Micromobility-Related Injuries From Unstructured Clinical Notes

    Burford, K.G.; Itzkowitz, N.G.; Ortega, A.G.; Teitler, J.O.; Rundle, A.G. Use of Generative AI to Identify Helmet Status Among Patients With Micromobility-Related Injuries From Unstructured Clinical Notes. JAMA Netw. Open 2024, 7, e2425981, doi:10.1001/jamanetworkopen.2024.25981

  3. [3]

    Opportunities and Challenges for ChatGPT and Large Language Models in Biomedicine and Health

    Tian, S.; Jin, Q.; Yeganova, L.; Lai, P.-T.; Zhu, Q.; Chen, X.; Yang, Y.; Chen, Q.; Kim, W.; Comeau, D.C.; et al. Opportunities and Challenges for ChatGPT and Large Language Models in Biomedicine and Health. Brief. Bioinform. 2024, 25, bbad493, doi:10.1093/bib/bbad493

  4. [4]

    Large Language Models Encode Clinical Knowledge

    Singhal, K.; Azizi, S.; Tu, T.; Mahdavi, S.S.; Wei, J.; Chung, H.W.; Scales, N.; Tanwani, A.; Cole-Lewis, H.; Pfohl, S.; et al. Large Language Models Encode Clinical Knowledge. Nature 2023, 620, 172–180, doi:10.1038/s41586-023-06291-2

  5. [5]

    A Large Language Model for Electronic Health Records

    Yang, X.; Chen, A.; PourNejatian, N.; Shin, H.C.; Smith, K.E.; Parisien, C.; Compas, C.; Martin, C.; Costa, A.B.; Flores, M.G.; et al. A Large Language Model for Electronic Health Records. npj Digit. Med. 2022, 5, 194, doi:10.1038/s41746-022-00742-2

  6. [6]

    Food and Drug Administration. Considerations for the Use of Artificial Intelligence To Support Regulatory Decision-Making for Drug and Biological Products Available online: https://www.fda.gov/regulatory-information/search-fda-guidance-documents/considerations-use-artificial-intelligence-support-regulatory-decision-making-drug-and-biological (accessed on ...

  7. [7]

    Machine Learning Methods in Health Economics and Outcomes Research—The PALISADE Checklist: A Good Practices Report of an ISPOR Task Force

    Padula, W.V.; Kreif, N.; Vanness, D.J.; Adamson, B.; Rueda, J.-D.; Felizzi, F.; Jonsson, P.; IJzerman, M.J.; Butte, A.; Crown, W. Machine Learning Methods in Health Economics and Outcomes Research—The PALISADE Checklist: A Good Practices Report of an ISPOR Task Force. Value Heal. 2022, 25, 1063–1080, doi:10.1016/j.jval.2022.03.022

  8. [8]

    Assessing Real-World Data From Electronic Health Records for Health Technology Assessment: The SUITABILITY Checklist: A Good Practices Report of an ISPOR Task Force

    Fleurence, R.L.; Kent, S.; Adamson, B.; Tcheng, J.; Balicer, R.; Ross, J.S.; Haynes, K.; Muller, P.; Campbell, J.; Bouée-Benhamiche, E.; et al. Assessing Real-World Data From Electronic Health Records for Health Technology Assessment: The SUITABILITY Checklist: A Good Practices Report of an ISPOR Task Force. Value Heal. 2024, 27, 692–701, doi:10.1016/j.jv...

Show all 29 references
  1. [9]

    Implementing Accuracy, Completeness, and Traceability for Data Reliability

    Riskin, D.J.; Monda, K.L.; Gagne, J.J.; Reynolds, R.; Garan, A.R.; Dreyer, N.; Muntner, P.; Bradbury, B.D. Implementing Accuracy, Completeness, and Traceability for Data Reliability. JAMA Netw. Open 2025, 8, e250128, doi:10.1001/jamanetworkopen.2025.0128

  2. [10]

    A Framework for Evaluating Clinical Artificial Intelligence Systems without Ground-Truth Annotations

    Kiyasseh, D.; Cohen, A.; Jiang, C.; Altieri, N. A Framework for Evaluating Clinical Artificial Intelligence Systems without Ground-Truth Annotations. Nat. Commun. 2024, 15, 1808, doi:10.1038/s41467-024-46000-9

  3. [11]

    A Reproducible and Transparent Framework for Evaluating Foundation Models (accessed on 21 May 2025)

    Stanford University. A Reproducible and Transparent Framework for Evaluating Foundation Models (accessed on 21 May 2025)

  4. [12]

    Approach to Machine Learning for Extraction of Real-World Data Variables from Electronic Health Records

    Adamson, B.; Waskom, M.; Blarre, A.; Kelly, J.; Krismer, K.; Nemeth, S.; Gippetti, J.; Ritten, J.; Harrison, K.; Ho, G.; et al. Approach to Machine Learning for Extraction of Real-World Data Variables from Electronic Health Records. Front. Pharmacol. 2023, 14, 1180962, doi:10....

  5. [13]

    Large Language Model Extraction of PD-L1 Biomarker Testing Details From Electronic Health Records

    Cohen, A.B.; Adamson, B.; Larch, J.K.; Amster, G. Large Language Model Extraction of PD-L1 Biomarker Testing Details From Electronic Health Records. AI Precis. Oncol. 2025, doi:10.1089/aipo.2024.0043

  6. [14]

    Replication of Real-World Evidence in Oncology Using Electronic Health Record Data Extracted by Machine Learning

    Benedum, C.M.; Sondhi, A.; Fidyk, E.; Cohen, A.B.; Nemeth, S.; Adamson, B.; Estévez, M.; Bozkurt, S. Replication of Real-World Evidence in Oncology Using Electronic Health Record Data Extracted by Machine Learning. Cancers (Basel) 2023, 15, doi:10.3390/cancers15061853

  7. [15]

    A Natural Language Processing Algorithm to Improve Completeness of ECOG Performance Status in Real-World Data

    Cohen, A.B.; Rosic, A.; Harrison, K.; Richey, M.; Nemeth, S.; Ambwani, G.; Miksad, R.; Haaland, B.; Jiang, C. A Natural Language Processing Algorithm to Improve Completeness of ECOG Performance Status in Real-World Data. Appl. Sci. 2023, 13, 6209, doi:10.3390/app13106209

  8. [16]

    Impact of Carboplatin and Cisplatin Shortages on Treatment Patterns in Patients with Metastatic Solid Tumors

    Castellanos, E.; Yuan, Q.; Wadé, N.; Patel, K.; Rinaldi, C.; Reiss, S.; Hankinson, E.; Cohen, A.B.; Estevez, M. Impact of Carboplatin and Cisplatin Shortages on Treatment Patterns in Patients with Metastatic Solid Tumors. J. Clin. Oncol. 2024, 42, 11149–11149, doi:10.1200/jco....

  9. [17]

    Real-World CtDNA Testing Patterns, Associated Biomarkers and Sites of Metastasis in Early Stage Colorectal Cancer

    Fidyk, E.; Kalesinskas, L.; Krismer, K.; Blarre, A.; Bouzit, L.; Ritten, J.J.; Marinescu, A.; Kelly, J.; Harrison, K.; Cohen, A.B. Real-World CtDNA Testing Patterns, Associated Biomarkers and Sites of Metastasis in Early Stage Colorectal Cancer. J. Clin. Oncol. 2024, 42, 3610–...

  10. [18]

    PPM8 A MACHINE LEARNING MODEL FOR CANCER BIOMARKER IDENTIFICATION IN ELECTRONIC HEALTH RECORDS

    Ambwani, G.; Cohen, A.; Estévez, M.; Singh, N.; Adamson, B.; Nussbaum, N.C.; Birnbaum, B. PPM8 A MACHINE LEARNING MODEL FOR CANCER BIOMARKER IDENTIFICATION IN ELECTRONIC HEALTH RECORDS. Value in Health 2019, 22, S334, doi:10.1016/j.jval.2019.04.1631

  11. [19]

    TIFTI: A Framework for Extracting Drug Intervals from Longitudinal Clinic Notes

    Agrawal, M.; Adams, G.; Nussbaum, N.; Birnbaum, B. TIFTI: A Framework for Extracting Drug Intervals from Longitudinal Clinic Notes. arXiv 2018, doi:10.48550/arxiv.1811.12793

  12. [20]

    A Hybrid Approach to Scalable Real-World Data Curation by Machine Learning and Human Experts

    Waskom, M.L.; Tan, K.; Wiberg, H.; Cohen, A.B.; Wittmershaus, B.; Shapiro, W. A Hybrid Approach to Scalable Real-World Data Curation by Machine Learning and Human Experts. medRxiv 2023, 2023.03.06.23286770, doi:10.1101/2023.03.06.23286770

  13. [21]

    Cohen, A.B.; Estevez, M.; Kelly, J.G.; Gippetti, J.N.; Fidyk, E.L.; Brandstadter, J.D.; Fajgenbaum, D.C. Clinical Characteristics, Treatment Trends, and Outcomes of Patients with HHV-8-Negative/Idiopathic Multicentric Castleman Disease Treated with Siltuximab in a Machine Lear...

  14. [22]

    Yuan, Q.; Dolor, A.; Qian, Y.; Donnelly, D.; Estevez, M.; Kuznetsova, Y.; Singh, N.; Yerram, P. Leveraging Machine Learning to Assess the Association of Rash and Survival in Patients With Advanced NSCLC Available online: https://www.ispor.org/heor-resources/presentations-datab...

  15. [23]

    Magee, K.; Yuan, Q.; Blarre, A.; Cohen, A.B.; Dolor, A.; Krismer, K.; Williams, T.; Zhang, Q. Performance Assessment and Validation of Real-World Response Data Generated Using a Deep Learning-Based Natural Language Processing Model Across Multiple Solid Tumors Available online...

  16. [24]

    Concordance of Response-Based Clinical Trial and Machine Learning–Generated Real-World End Points

    Zhang, Q.; Krismer, K.; Lu, Y.; Yuan, Q.; Dolor, A.; Blarre, A.; Cohen, A.B.; Williams, T.; Maund, S.; Srivastava, M.K.; et al. Concordance of Response-Based Clinical Trial and Machine Learning–Generated Real-World End Points. J. Clin. Oncol. 2025, 43, doi:10.1200/jco.2025.43....

  17. [25]

    Real-World (Rw) CtDNA Testing Trends and Associated Outcomes in Patients (Pts) with Early Stage Breast Cancer (EBC)

    Fidyk, E.; Ward, P.J.; Estevez, M.; Krismer, K.; Ritten, J.J.; Marinescu, A.; Cohen, A.B. Real-World (Rw) CtDNA Testing Trends and Associated Outcomes in Patients (Pts) with Early Stage Breast Cancer (EBC). J. Clin. Oncol. 2025, 43, 555–555, doi:10.1200/jco.2025.43.16_suppl.555

  18. [26]

    Considerations for the Use of Machine Learning Extracted Real-World Data to Support Evidence Generation: A Research-Centric Evaluation Framework

    Estevez, M.; Benedum, C.M.; Jiang, C.; Cohen, A.B.; Phadke, S.; Sarkar, S.; Bozkurt, S. Considerations for the Use of Machine Learning Extracted Real-World Data to Support Evidence Generation: A Research-Centric Evaluation Framework. Cancers (Basel) 2022, 14, doi:10.3390/cance...

  19. [27]

    Raising the Bar for Real-World Data in Oncology: Approaches to Quality Across Multiple Dimensions

    Castellanos, E.H.; Wittmershaus, B.K.; Chandwani, S. Raising the Bar for Real-World Data in Oncology: Approaches to Quality Across Multiple Dimensions. JCO Clin. Cancer Inform. 2024, 8, e2300046, doi:10.1200/cci.23.00046

  20. [28]

    Misclassification in Administrative Claims Data: Quantifying the Impact on Treatment Effect Estimates

    Funk, M.J.; Landi, S.N. Misclassification in Administrative Claims Data: Quantifying the Impact on Treatment Effect Estimates. Curr. Epidemiology Rep. 2014, 1, 175–185, doi:10.1007/s40471-014-0027-z

  21. [29]

    Good Practices for Quantitative Bias Analysis

    Lash, T.L.; Fox, M.P.; MacLehose, R.F.; Maldonado, G.; McCandless, L.C.; Greenland, S. Good Practices for Quantitative Bias Analysis. Int. J. Epidemiology 2014, 43, 1969–1985, doi:10.1093/ije/dyu149. Table

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.