REVIEW 2 major objections 5 minor 45 references
Do It Right! A Methodology for Successful NLP System Development
T0 review · 2 major / 5 minor · reviewed 2026-07-11 · grok-4.5
Pith's one-line read Algorithm skill alone does not make clinical NLP succeed; managing the full Systems Development Life Cycle does, and large language models make that discipline more necessary, not less.
desk verdict Solid process synthesis for clinical IE that correctly updates SDLC for LLM failure modes; useful guidance, not a new result or causal proof. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The Systems Development Life Cycle (SDLC) as a managed sequence of phases—Planning, Analysis, Design, Implementation, Testing, Deployment, Maintenance—applied as an iterative cycle to clinical information extraction, with concept sheets, feasibility analysis of semantic and contextual ambiguity, and matched annotation as the load-bearing early steps.
What would settle it
A controlled comparison of matched clinical extraction projects—one cohort run under explicit SDLC discipline (documented concept sheets, feasibility review, matched annotation, held-out test, revalidation logs) and one under algorithm- or prompt-first practice—would show whether measured failure, rework, or error rates fall when the cycle is followed.
Extended reading notes
Core claim
Regardless of project size, deliberately carrying a clinical information-extraction system through the Systems Development Life Cycle—especially purpose and scope definition, feasibility analysis that information exists and is machine-accessible, annotation that matches the system’s true data access, clean train/test separation, effectiveness evaluation, governed deployment, and maintenance against linguistic and model-version drift—raises the chance of a usable, trustworthy system. Large language models expand what is possible but introduce hallucination, prompt sensitivity, non-determinism, and model-version drift that make the same process more, not less, necessary.
Load-bearing premise
The paper assumes that the main reason clinical NLP projects fail is the same mismanagement already documented in general software and chart-abstraction work, so transplanting the classic SDLC sequence is the right primary remedy, without controlled evidence that SDLC adherence itself cuts clinical NLP failure rates.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This methodology paper argues that clinical NLP information-extraction projects fail for the same process reasons documented in MIS software and chart-abstraction literature, and that deliberately applying the classic SDLC phases—planning, analysis, design, implementation, testing, deployment, and maintenance—improves the chance of a usable system. It maps each phase to concrete clinical-NLP practices (concept sheets, semantic/contextual ambiguity checks, matched annotation, pipeline design, train/test separation, efficacy vs. effectiveness, error analysis, and revalidation against linguistic and model-version drift) and argues that LLMs expand capability while introducing failure modes (hallucination, prompt sensitivity, non-determinism, model-version drift, data governance) that make process discipline more, not less, necessary. The paper is explicitly literature- and practice-based rather than a new empirical trial.
Significance. If the process framing holds, the paper is a useful, timely synthesis for clinical informatics: it counters the common impression that algorithms or LLM APIs alone suffice, and it gives multidisciplinary teams a shared checklist that spans feasibility, annotation design, evaluation, deployment governance, and maintenance. Strengths include the explicit parallel to chart abstraction, the insistence that annotators see only what the system will see, the distinction between efficacy and effectiveness, the systematic/random/hallucination error taxonomy, and the treatment of model-version drift as a maintenance concern. The contribution is pedagogical and organizational rather than a new algorithm or causal trial; that is appropriate for a methodology paper but limits claims of demonstrated impact.
major comments (2)
- Introduction, §2, and Conclusions: The central claim that SDLC adherence improves clinical NLP success rests on transplanting MIS failure literature [14, 22] and chart-review methodology [15–17] without controlled or even systematic observational evidence that process mismanagement is the dominant failure mode in clinical NLP, or that SDLC adherence reduces failure relative to algorithm- or LLM-first practice. For a methodology paper this is acceptable if framed as structured guidance, but the manuscript should state the evidential limit more explicitly and avoid language that reads as a demonstrated causal remedy.
- §4.1–4.3 and §8: Feasibility and evaluation recommendations are sound in principle, but the paper does not specify operational decision rules (e.g., how much manual review, what agreement threshold, what residual error rate is acceptable for a given purpose in Table 1). Without such criteria, teams cannot know when to stop iterating or when to abandon automation. Adding illustrative thresholds or worked decision points would make the guidance actionable rather than only conceptual.
minor comments (5)
- References [28]–[31] list incomplete bibliographic detail (“Venue pending verification,” “Authors TBD”). These should be completed or replaced with citable sources before publication.
- Affiliation is a placeholder; corresponding-author contact is present but institutional affiliation should be filled for the published version.
- Figure 1 and Figure 3 are described but not fully self-contained in the text; ensure captions and in-text callouts make the productivity pyramid and the cyclic SDLC readable without the figure alone.
- Table 1 usefully contrasts three purpose scenarios; a short note on how success metrics differ across rows (publication novelty vs. dataset accuracy vs. operational robustness) would tighten the link to later testing/effectiveness sections.
- Minor wording: occasional missing spaces after punctuation in the source (e.g., “Italsorevealsthekinds”) should be cleaned in production.
Circularity Check
No circularity: methodology paper with no fitted parameters, no self-referential derivations, and recommendations grounded in external literature.
full rationale
This is a process-methodology paper that maps the classic Systems Development Life Cycle (planning, analysis, design, implementation, testing, deployment, maintenance) onto clinical information-extraction NLP, with additional discussion of LLM-specific failure modes (hallucination, prompt sensitivity, non-determinism, model-version drift, data governance). It advances no quantitative prediction, no fitted parameter renamed as a result, and no uniqueness theorem. Load-bearing support is drawn from external MIS project-failure literature (e.g., Kappelman et al., Nelson), chart-abstraction methodology, and standard clinical NLP practice; author-affiliated prior work appears only as illustrative experience (e.g., advanced basal-cell-carcinoma extraction, sublanguage clustering) and is not required to force the SDLC claim. There is therefore no self-definitional step, no fitted-input-called-prediction, no load-bearing self-citation chain, and no renaming of a known result as a new derivation. Circularity score is 0.
Assumptions & free parameters
assumptions (5)
- domain assumption Mismanagement is a primary cause of software project failure and clinical NLP projects share that risk profile.
- domain assumption Manual chart-abstraction quality practices (definitions, training, adjudication, agreement) transfer to LLM-based extraction because LLMs approximate human reading of narrative notes.
- domain assumption Target clinical information may be absent, ambiguous, or only present as surrogates in care-oriented documentation, so feasibility must be assessed before system building.
- domain assumption Reference standards must be produced under the same information access constraints the NLP system will have, or measured performance is misleading.
- domain assumption Clinical language and hosted LLM model versions both drift over time, so deployed systems require scheduled revalidation.
Cite this review
Pith. "Pith review of Do It Right! A Methodology for Successful NLP System Development." pith.science (2026). https://pith.science/paper/AGX5M2MX
@misc{pith2026260705644,
author = {Pith},
title = {Pith review of: Do It Right! A Methodology for Successful NLP System Development},
year = {2026},
howpublished = {\url{https://pith.science/paper/AGX5M2MX}},
note = {Machine review of arXiv:2607.05644}
}
read the original abstract
Natural language processing (NLP) is a common method for supplying data to clinical research and decision making by extracting information from electronic medical records. Numerous textbooks and tutorials describe specific algorithms and applications for text processing, yet algorithmic knowledge is only one ingredient of a successful NLP project. Drawing on the available literature, this paper presents a stepwise approach that applies the Systems Development Life Cycle (SDLC) to projects that rely on data extraction through language processing.
Figures
Reference graph
Works this paper leans on
-
[1]
A. Névéol, P. Zweigenbaum, Clinical Natural Language Processing in 2014: Foundational Methods Supporting Efficient Healthcare., Year- book of medical informatics 10 (1) (2015) 194–8.doi:10.15265/ IY-2015-035. URLhttp://www.ncbi.nlm.nih.gov/pubmed/26293868http://www. pubmedcentral.nih.gov/articlerender.fcgi?artid=PMC4587052
-
[2]
A. Névéol, P. Zweigenbaum, Clinical Natural Language Processing in 2015: Leveraging the Variety of Texts of Clinical Interest., Yearbook of medical informatics (1) (2016) 234–239.doi:10.15265/IY-2016-049. URLhttp://www.ncbi.nlm.nih.gov/pubmed/27830256
-
[3]
C. Friedman, S. B. Johnson, Natural Language and Text Processing in Biomedicine, Springer New York, 2006, pp. 312–343.doi:10.1007/ 0-387-36278-9_8. URLhttp://link.springer.com/10.1007/0-387-36278-9_8
-
[4]
K. S. Jones, Natural language processing: a historical review, in: Cur- rent Issues in Computational Linguistics: in Honour of Don Walker, Vol. 7, 1994, pp. 3–16.doi:10.1007/978-0-585-35958-8_1. 24
-
[5]
N. Sager, C. Friedman, M. S. Lyman, Medical Language Processing: Computer Management of Narrative Data, New York: Addison-Wesley, Reading, Mass, 1987. URLhttp://www.worldcat.org/oclc/13947442http://www. aclweb.org/anthology/J/J89/J89-3009.pdfhttp://portal.acm. org/citation.cfm?id=49103.1046397
-
[6]
S. Doan, M. Conway, T. M. Phuong, L. Ohno-Machado, Natural language processing in biomedicine: a unified system architecture overview., Methods in molecular biology (Clifton, N.J.) 1168 (2014) 275–94.arXiv:1401.0569,doi:10.1007/978-1-4939-0847-9_16. URLhttp://www.ncbi.nlm.nih.gov/pubmed/24870142%5Cnhttp: //arxiv.org/abs/1401.0569http://www.ncbi.nlm.nih.go...
work page Pith review arXiv doi:10.1007/978-1-4939-0847-9_16 2014
-
[7]
P. M. Nadkarni, L. Ohno-Machado, W. W. Chapman, Natural language processing: an introduction, Journal of the Amer- ican Medical Informatics Association 18 (5) (2011) 544–551. doi:10.1136/amiajnl-2011-000464. URLhttp://jamia.bmj.com/cgi/doi/10.1136/ amiajnl-2011-000464http://jamia.oxfordjournals.org/lookup/ doi/10.1136/amiajnl-2011-000464
-
[8]
R. Grishman, Information extraction: Techniques and challenges, Infor- mation Extraction A Multidisciplinary Approach to an Emerging Infor- mation Technology (1997) 10–27doi:10.1007/3-540-63438-X_2
Show all 45 references
-
[9]
M. J. Schuemie, J. A. Kors, B. Mons, Word sense disambiguation in the biomedical domain: an overview., Journal of computational biology : a journal of computational molecular cell biology 12 (5) (2005) 554–65. doi:10.1089/cmb.2005.12.554. URLhttp://www.ncbi.nlm.nih.gov/pubmed/15952878
2005 doi
-
[10]
Sarawagi, Information Extraction, Foundations and Trends in Databases 1 (3) (2008) 261–377.doi:10.1561/1900000003
S. Sarawagi, Information Extraction, Foundations and Trends in Databases 1 (3) (2008) 261–377.doi:10.1561/1900000003. URLhttp://www.nowpublishers.com/product.aspx?product=DBS& doi=1900000003 25
2008 doi
-
[11]
S. M. Meystre, G. K. Savova, K. C. Kipper-Schuler, J. F. Hurdle, Ex- tracting information from textual documents in the electronic health record: a review of recent research., Yearbook of medical informatics (2008) 128–44. URLhttp://www.ncbi.nlm.nih.gov/pubmed/18660887
2008
-
[12]
Demner-Fushman, W
D. Demner-Fushman, W. W. Chapman, C. J. McDonald, What can natural language processing do for clinical decision support?, Journal of biomedical informatics 42 (5) (2009) 760–72.doi:10.1016/j.jbi. 2009.08.007. URLhttp://www.ncbi.nlm.nih.gov/pubmed/19683066
2009 doi
-
[13]
Spyns, Natural language processing in medicine: an overview., Methods of information in medicine 35 (4-5) (1996) 285–301
P. Spyns, Natural language processing in medicine: an overview., Methods of information in medicine 35 (4-5) (1996) 285–301. URLhttp://www.ncbi.nlm.nih.gov/pubmed/9019092?dopt= Citationhttp://www.ncbi.nlm.nih.gov/pubmed/9019092
1996
-
[14]
L. A. Kappelman, R. McKeeman, L. Zhang, Early Warn- ing Signs of it Project Failure: The Dominant Dozen, In- formation Systems Management 23 (4) (2006) 31–36.doi: 10.1201/1078.10580530/46352.23.4.20060901/95110.4. URLhttp://web.a.ebscohost.com/ehost/detail/detail?sid= 27084ed3...
2006 doi
-
[15]
Vassar, M
M. Vassar, M. Holzmann, The retrospective chart review: important methodological considerations, J Educ Eval Health Prof 10 (2013).doi: 10.3352/jeehp.2013.10.12
2013 doi
-
[16]
R. E. Gearing, I. A. Mian, J. Barber, A. Ickowicz, A methodology for conducting retrospective chart review research in child and adolescent psychiatry., Journal of the Canadian Academy of Child and Adolescent Psychiatry = Journal de l’Académie canadienne de psychiatrie de l’en...
2006
-
[17]
L. A. Knake, M. Ahuja, E. L. McDonald, K. K. Ryckman, N. Weathers, T. Burstain, J. M. Dagle, J. C. Murray, P. Nadkarni, Quality of EHR data extractions for studies of preterm birth in a tertiary care center: guidelines for obtaining reliable data., BMC pediatrics 16 (2016) 59....
2016 doi
-
[18]
W. Yeoh, A. Koronios, Critical success factors for business intelligence systems, in: Journal of computer information systems, Vol. 50, 2010, pp. 23–32.doi:abs/10.1080/08874417.2010.11645404. URLhttps://www.tandfonline.com/doi/abs/10.1080/08874417. 2010.11645404
2010 doi
-
[19]
Asosheh, S
A. Asosheh, S. Nalchigar, M. Jamporazmey, Information technology project evaluation: An integrated data envelopment analysis and balanced scorecard approach, Expert Systems with Applications 37 (8) (2010) 5931–5938.doi:10.1016/j.eswa.2010.02.012. URLhttp://linkinghub.elsevier....
2010 doi
-
[20]
Baccarini, Logical Framework Method for Defining Project Success, Project Management Journal (1999)
D. Baccarini, Logical Framework Method for Defining Project Success, Project Management Journal (1999). URLhttps://www.pmi.org/learning/library/ logical-framework-method-defining-project-success-5309
1999
-
[21]
Davis, Logical framework analysis: a methodology to turn vision into reality, in: AIPM National Conference, 2005
K. Davis, Logical framework analysis: a methodology to turn vision into reality, in: AIPM National Conference, 2005
2005
-
[22]
R. R. Nelson, IT Project Management: Infamous Failures, Classic Mistakes, and Best Practices, MIS Quarterly Executive 6 (2) (2007) 67–78. URLhttp://misqe.org/ojs2/index.php/misqe/article/view/ 128http://www2.comm.virginia.edu/cmit/Research/MISQE6-07. pdf
2007
-
[23]
D. K. Iwamoto, W. M. Liu, The impact of racial identity, ethnic iden- tity, asian values and race-related stress on Asian Americans and Asian internationalcollegestudents’psychologicalwell-being., Journalofcoun- seling psychology 57 (1) (2010) 79–91.doi:10.1037/a0017393. 27 UR...
2010 doi
-
[24]
S. M. Meystre, Y. Kim, G. T. Gobbel, M. E. Matheny, A. Redd, B. E. Bray, J. H. Garvin, Congestive heart failure information extraction framework for automated treatment performance measures assessment., Journal of the American Medical Informatics Association : JAMIA 24 (e1) (2...
2017 doi
-
[25]
Divita, T
G. Divita, T. E. Workman, M. E. Carter, A. Redd, M. H. Samore, A. V. Gundlapalli, PlateRunner: A Search Engine to Identify EMR Boilerplates., Studies in health technology and informatics 226 (2016) 33–6. URLhttp://www.pubmedcentral.nih.gov/articlerender.fcgi? artid=3900197&too...
2016
-
[26]
V. Liu, M. P. Clark, M. Mendoza, R. Saket, M. N. Gardner, B. J. Turk, G. J. Escobar, Automated identification of pneumonia in chest radio- graph reports in critically ill patients., BMC medical informatics and decision making 13 (2013) 90.doi:10.1186/1472-6947-13-90. URLhttp:/...
2013 doi
-
[27]
Topaz, K
M. Topaz, K. Lai, D. Dowding, V. J. Lei, A. Zisberg, K. H. Bowles, L. Zhou, Automated identification of wound information in clinical notes of patients with heart diseases: Developing and validating a nat- ural language processing application., International journal of nursing...
2016 doi
-
[28]
Guo, et al., Improving large language models for clinical named en- tity recognition via prompt engineeringPMC11339492
Z. Guo, et al., Improving large language models for clinical named en- tity recognition via prompt engineeringPMC11339492. Venue pending verification. (2024)
2024
-
[29]
Syrstad, et al., Harnessing large language models for efficient data extraction in systematic reviews: the role of prompt engineer- ingPMC12559671
O. Syrstad, et al., Harnessing large language models for efficient data extraction in systematic reviews: the role of prompt engineer- ingPMC12559671. Venue pending verification. (2024)
2024
-
[30]
Venue pending ver- ification
Authors TBD, Retrieval augmented generation for large language mod- els in healthcare: a systematic reviewPMC12157099. Venue pending ver- ification. (2024)
2024
-
[31]
Authors TBD, Retrieval-augmented generation (RAG) in healthcare: a comprehensive review, AI (MDPI) 6 (9) (2025) 226
2025
-
[32]
T. Wolf, L. Debut, V. Sanh, et al., HuggingFace’s Transformers: State- of-the-art natural language processing, in: Proceedings of the 2020 Con- ference on Empirical Methods in Natural Language Processing: System Demonstrations, Association for Computational Linguistics, 2020, ...
2020
- [33]
-
[34]
Honnibal, I
M. Honnibal, I. Montani, S. Van Landeghem, A. Boyd, spaCy: Industrial-strength natural language processing in Python, Zenodo, 2020.doi:10.5281/zenodo.1212303
2020 doi
-
[35]
M. J. Jurafsky D., Speech and Language Processing, 2nd (2008)
2008
- [36]
-
[37]
Lupyan, R
G. Lupyan, R. Dale, The role of adaptation in understanding linguistic diversity, Language Structure and Environment Social, cultural, and natural factors Edited by Rik De Busser and Randy J. LaPolla (2015) 184doi:10.1075/clscc.6.11lup. URLhttps://benjamins.com/catalog/clscc.6...
2015 doi
-
[38]
W. L. Hamilton, J. Leskovec, D. Jurafsky, Cultural Shift or Linguistic Drift? Comparing Two Computational Measures of Semantic Change., Proceedings of the Conference on Empirical Methods in Natural Lan- guage Processing. Conference on Empirical Methods in Natural Lan- guage Pr...
2016
-
[39]
Gentner, I
D. Gentner, I. M. France, The Verb Mutability Effect: Stud- ies of the Combinatorial Semantics of Nouns and Verbs, in: Lexical Ambiguity Resolution, Elsevier, 1988, pp. 343–382. doi:10.1016/B978-0-08-051013-2.50018-5. URLhttp://linkinghub.elsevier.com/retrieve/pii/ B9780080510...
1988 doi
-
[40]
Z. S. Harris, A Theory of Language and Information: A Mathematical Approach, Clarendon Press, 1991. URLhttp://www.dmi.columbia.edu/zellig/langinf.html
1991
-
[41]
Friedman, P
C. Friedman, P. Kra, A. Rzhetsky, Two biomedical sublanguages: a description based on the theories of Zellig Harris., Journal of biomed- ical informatics 35 (4) (2002) 222–35.doi:10.1016/S1532-0464(03) 00012-1. URLhttp://www.ncbi.nlm.nih.gov/pubmed/12755517
2002 doi
-
[42]
Patterson, J
O. Patterson, J. F. Hurdle, Document clustering of clinical narratives: a systematic study of clinical sublanguages., AMIA ... Annual Sympo- sium proceedings / AMIA Symposium. AMIA Symposium 2011 (2011) 1099–1107. URLhttp://www.pubmedcentral.nih.gov/articlerender.fcgi? artid=3...
2011
-
[43]
Doing-Harris, O
K. Doing-Harris, O. Patterson, S. Igo, J. Hurdle, Document Sublan- guage Clustering to Detect Medical Specialty in Cross-institutional Clinical Texts., in: Proceedings of the ACM ... International Workshop on Data and Text Mining in Biomedical Informatics . ACM Interna- tional...
2013 doi
-
[44]
Bernhardt, S
P. Bernhardt, S. M. Humphrey, T. C. Rindflesch, Determining promi- nent subdomains in medicine, AMIA Annu Symp Proc (2005) 46–50. URLhttp://www.ncbi.nlm.nih.gov/pubmed/16778999?dopt= Citation
2005
-
[45]
D. S. Carrell, R. E. Schoen, D. A. Leffler, M. Morris, S. Rose, A. Baer, S. D. Crockett, R. A. Gourevitch, K. M. Dean, A. Mehrotra, Challenges in adapting existing clinical natural language processing systems to multiple, diverse health care settings., Journal of the American ...
2017 doi
Reviewed July 11, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.