REVIEW 3 major objections 4 minor 34 references
A human-in-the-loop AI agent can serve as the first-pass author of a scholar's translational impact record, shifting staff from collecting and writing to reviewing.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 19:19 UTC pith:6PU3HPY4
load-bearing objection A genuine deployment study of an LLM agent for impact reporting with honest reporting, but the recall gap means the strongest conclusion runs ahead of the evidence. the 3 major comments →
Real-World Evaluation of an AI Agent Drafting Translational Impact Summaries
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
On the paper's own terms, the central discovery is that an AI agent can draft a scholar's impact record well enough that expert staff spend their time curating rather than compiling. Two independent reviewers in the hub's actual workflow accepted or edited 414 of 507 agent-proposed findings, with a median review time of 14 minutes per scholar. The agent drew evidence from indexed databases and open-web sources, covering clinical, community, economic, and policy benefits, and proposed non-scholarly impacts that closed-source pipelines miss. The authors conclude that the human-in-the-loop agent can be the first-pass author of a scholar's impact record.
What carries the argument
The central object is the human-in-the-loop AI agent itself. Starting from a scholar seed—name, institution, and any internal grant records—the agent runs up to three research runs, each alternating a gather stage (up to 25 turns of planning and searching public research databases and the open web), an assemble stage that writes dossier sections and one-sentence TSBM impact summaries, and a critique stage that flags gaps, duplicates, contradictions, and weak sources. A register-revise stage then tunes summaries toward a target style. A two-model design, a large model for planning and drafting and a small model for high-volume page reading, keeps the context small and limits untrusted content
Load-bearing premise
The load-bearing premise is that the agent's recall of impact evidence is good enough for first-pass drafting, but the study does not measure completeness: the review judged only what the agent proposed, and the profile-discovery recall is relative to a pooled reference that includes the agent's own findings, so a finding every search missed is invisible to the metric.
What would settle it
Construct a ground-truth list of impacts for a few scholars by interviewing the scholars themselves and exhaustively searching sources the agent does not use, then run the agent and measure recall against that list. If recall is much lower than the 81.7% usable rate implies, the first-pass-author claim fails even if the findings the agent does propose are mostly usable.
If this is right
- Staff effort shifts from an estimated 15 hours of assembly per scholar to minutes of review, making cohort-scale impact reporting feasible within a reporting cycle.
- The agent surfaces non-scholarly impact evidence—clinical-program leadership, community engagement, cost or adoption data, policy uptake—that indexed-only pipelines miss.
- Human review remains necessary: at least one reviewer rejected 93 of 507 findings, mostly for weak sources or failing the reviewer's impact threshold, so the tool concentrates expert time on judgment rather than search.
- Profile discovery recall was close to human search (0.82 versus 0.82 against human-only search; 0.76 versus 0.85 against AI-guided search), with the gap largely from one profile type the agent cannot query.
- The study offers an in-workflow evaluation method that other AI-assisted reporting systems could use, measuring usable rate, review time, and inter-rater agreement.
Where Pith is reading between the lines
- The same draft-then-curate division of labor could generalize to other expert reporting tasks—tenure dossiers, grant progress reports, or program evaluations—where the bottleneck is assembling a defensible first draft from scattered sources.
- Because the study does not measure the completeness of impact evidence, the agent's true recall ceiling is unknown; a natural next test would build a complete ground-truth impact record for a few scholars by interviewing them and exhaustively searching, then compare what the agent retrieves.
- Wiring structured policy-citation and patent databases into the agent, as the paper's discussion suggests, is a testable extension that would likely strengthen economic and policy coverage and reduce dependence on open-web fallback.
- The moderate inter-rater agreement implies reviewers draw the 'completed impact' line differently; a short calibration step with example thresholds before reviewing could raise the unanimous usable rate.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper reports a real-world evaluation of a human-in-the-loop AI agent that assembles evidence dossiers and drafts one-sentence TSBM impact summaries for CTSA KL2/K12 scholars. Two staff reviewers independently reviewed all 507 findings generated for 10 scholars, with a unanimous usable rate of 81.7% and a median review time of 14 minutes per scholar. A separate profile discovery study compared the agent's recall against human-only and AI-guided searches. The authors conclude that the agent can serve as a first-pass author, shifting staff effort from collection and writing to reviewing, and thus making cohort-scale impact reporting feasible. The paper is transparent about its limitations, including the single-site sample, the estimated rather than measured manual baseline, and the fact that completeness of impact evidence is not established.
Significance. The study is a valuable contribution to the emerging literature on real-world evaluation of AI systems in live workflows. Its strengths include independent double coding of every proposed finding, a conservative unanimous usable rate, release of code/prompts/configuration for reproducibility, and an honest discussion of limitations. If replicated, the findings would provide concrete evidence for the feasibility of AI-assisted impact reporting and shift the debate from model capability to workflow integration. The principal risk is that the agent's recall of impact evidence is unmeasured; although the paper acknowledges this, the claimed relocation of staff effort depends on it.
major comments (3)
- [§2.2, §2.3, §4 (Discussion, Limitations)] The central claim that the agent shifts staff from collecting/writing to reviewing requires that the agent does not miss substantial impact evidence. The review study measures usability only among the agent's proposed findings; it does not measure what the agent missed. The profile discovery study checks only profiles, not the TSBM impact sections, and its pooled reference includes the agent's own findings, so recall is partly self-referential. The Discussion correctly states 'the completeness of the impact evidence is not established,' but this is a load-bearing gap. A concrete remedy would be a recall audit on a subset of scholars using an independent comprehensive manual search (or a second independent agent/search strategy), with recall of impact findings computed against that reference. Without such evidence, the 14-minute review time cannot be taken as the full staff effort.
- [§2.3, §4, Abstract] The claimed effort savings ('replacing an estimated 15 hours of manual assembly') rests on an unmeasured staff estimate, as the paper acknowledges. Because the headline benefit is the effort shift, the quantitative comparison is weak. A measured baseline, even for one or two scholars using the same verification standard, would substantially strengthen the claim that the workflow shifts staff from collecting to reviewing.
- [§4 (Limitations), Abstract Conclusion] The sample comprises 10 scholars at one hub with two reviewers. The paper frames the study as formative, but the abstract states that cohort-scale impact reporting is feasible. While the authors limit their claim with a single-site caveat, the evidence supports feasibility at one site only. A multi-site replication is future work, and the current data cannot establish generalizability across hubs, staff, or scholar cohorts.
minor comments (4)
- [Abstract, §1] The 'estimated 15 hours' is stated without a source or derivation. Please cite the hub's internal estimate or provide the basis for the figure.
- [§2.3, Table 1] The phrase 'Both reviewers accepted or edited 81.7%' is slightly ambiguous; consider 'The unanimous usable rate was 81.7%' to match the table header.
- [Figure 2] If space permits, increase the resolution of the diagram; the nested 'repeat' arrows are easy to misread in a compressed version.
- [§4] The term 'CRIS' (Current Research Information Systems) is used in the Discussion without definition; the abbreviation is unnecessary and could be replaced with the full term.
Circularity Check
No circularity: the central results are measured human-review outcomes; the acknowledged pooled-recall limitation is a gold-standard gap, not a circular reduction.
full rationale
The paper's central claim is an empirical workflow evaluation: staff review every agent-proposed finding, and the headline measures (81.7% unanimous usable rate, median 14 minutes per scholar, kappa 0.43) are directly observed outcomes of that review, not derived from the agent's own outputs. The review study is not circular because the reviewers independently accept, edit, or reject all 507 findings against their sources; the paper explicitly frames what the review can and cannot show: 'Because they judged the agent's own list, the review reveals what the agent proposed but not what it missed.' The profile discovery study is the only place with a self-referential element: recall is computed relative to a pooled reference that includes the agent's findings, with the paper stating 'Recall is therefore relative to this pool, so a profile every search missed is invisible to it.' That is transparently a limitation in the completeness of the relevance judgments, not a construction that forces the agent's recall to equal its input; the reported recall values differ across arms (0.82 vs 0.82; 0.76 vs 0.85), so the comparison retains contingent content. The paper also explicitly acknowledges the load-bearing gap in the Discussion: 'the completeness of the impact evidence is not established.' This is an honest statement of an unmeasured condition for the effort-shift claim, but it is not a circular derivation. No load-bearing self-citations by the present authors appear; the cited comparators are independent published studies and the reproducibility section releases code, prompts, and configuration. The 15-hour manual baseline is a staff estimate, not a fitted parameter renamed as a prediction. Accordingly, there is no exhibited circular step and the score is 0.
Axiom & Free-Parameter Ledger
axioms (4)
- domain assumption TSBM is a valid framework for classifying translational impact.
- domain assumption The unanimous usable rate is a valid measure of agent utility in the workflow.
- domain assumption The staff estimate of 15 hours manual effort per scholar is accurate.
- domain assumption The pooled-reference recall design provides an unbiased measure of discovery completeness.
read the original abstract
Introduction. Clinical and Translational Science Award (CTSA) programs must document their scholars' research impact, but assembling each scholar's record by hand takes staff an estimated 15 hours and does not scale to a full cohort. An artificial intelligence (AI) agent could serve as a tool to gather scholar data across platforms and disciplines. Methods. We built a human-in-the-loop AI agent that assembles a dossier of sourced evidence for each scholar and drafts one-sentence Translational Science Benefits Model (TSBM) impact summaries for staff review. We evaluated it in the impact-reporting workflow of one CTSA hub across 10 career-development (KL2/K12) scholars. Two evaluation staff independently coded all 507 findings as accept, edit, or reject; the primary measure was the unanimous usable rate, defined as the share both accepted or edited. Results. Both reviewers accepted or edited 81.7% of the agent's findings. Reviewers each spent a median of 14 minutes per scholar, replacing an estimated 15 hours of manual assembly. Inter-rater agreement was moderate (Cohen's kappa 0.43 on the usable-versus-reject decision). A profile discovery study found the agent's recall close to human search. The agent's impact evidence spanned all four TSBM domains, and about a third of the reviewed findings fell in non-scholarly categories that routine processes tend to miss. Reviewers rated synthesis accuracy 4.5 and usefulness 4.8 on a 5-point scale. Conclusions. A human-in-the-loop AI agent can serve as the first-pass author of a scholar's impact record, shifting staff from collecting and writing to reviewing, and making cohort-scale impact reporting feasible.
Figures
Reference graph
Works this paper leans on
-
[1]
Conte and M
Marisa L. Conte and M. Bishr Omary. NIH Career Development Awards: conversion to research grants and regional distribution.The Journal of Clinical Investigation, 128(12):5187–5190, December 2018. ISSN 0021-
2018
-
[2]
The Impact of Individual Mentored Career Development (K) Awards on the Research Trajectories of Early-Career Scientists.Academic Medicine, 94(5):708–714, May 2019
Silda Nikaj and P Kay Lund. The Impact of Individual Mentored Career Development (K) Awards on the Research Trajectories of Early-Career Scientists.Academic Medicine, 94(5):708–714, May 2019. ISSN 1040-
2019
-
[3]
URLhttps://grants.nih.gov/grants/ policy/nihgps/html5/section_12/12.2.5_institutional_scientist_development_programs.h tm
12.2.5 Institutional Scientist Development Programs, March 2026. URLhttps://grants.nih.gov/grants/ policy/nihgps/html5/section_12/12.2.5_institutional_scientist_development_programs.h tm. 12 Real-World Evaluation of an AI Agent Drafting Translational Impact SummariesA PREPRINT
2026
-
[4]
Sorkness, Linda Scholl, Alecia M
Christine A. Sorkness, Linda Scholl, Alecia M. Fair, and Jason G. Umans. KL2 mentored career development programs at clinical and translational science award hubs: Practices and outcomes.Journal of Clinical and Translational Science, 4(1):43–52, February 2020. ISSN 2059-8661. doi: 10.1017/cts.2019.424. URLhttps: //www.cambridge.org/core/journals/journal-o...
-
[5]
Kelli Qua, Fei Yu, Tanha Patel, Gaurav Dave, Katherine Cornelius, and Clara M. Pelfrey. Scholarly Pro- ductivity Evaluation of KL2 Scholars Using Bibliometrics and Federal Follow-on Funding: Cross-Institution Study.Journal of Medical Internet Research, 23(9):e29239, September 2021. doi: 10.2196/29239. URL https://www.jmir.org/2021/9/e29239
-
[6]
Eric J. Nehl, Clara M. Pelfrey, Deborah DiazGranados, Gaurav Dave, and Nicole M. Llewellyn. Academic influencers: Clinical and Translational Science scholars and trainees at the intersection of influential scholarship and public attention.Journal of Clinical and Translational Science, 9(1):e130, January 2025. ISSN 2059-8661. doi: 10.1017/cts.2025.10067. U...
arXiv 2025
-
[7]
Rebecca Helton, Scott Pearson, and Katherine Hartmann. 112 Flight Tracker: A REDCap Tool to Streamline Career Development Grant Preparation and Reporting.Journal of Clinical and Translational Science, 8(s1):32, April 2024. ISSN 2059-8661. doi: 10.1017/cts.2024.110. URLhttps://www.cambridge.org/core/journ als/journal-of-clinical-and-translational-science/a...
-
[8]
Douglas A. Luke, Cathy C. Sarli, Amy M. Suiter, Bobbi J. Carothers, Todd B. Combs, Jae L. Allen, Courtney E. Beers, and Bradley A. Evanoff. The Translational Science Benefits Model: A New Framework for Assessing the Health and Societal Benefits of Clinical and Translational Sciences.Clinical and Translational Science, 11(1): 77–84, January 2018. ISSN 1752...
-
[9]
We Should Evaluate Real-World Impact.Computational Linguistics, 51(4):1419–1431, December
Ehud Reiter. We Should Evaluate Real-World Impact.Computational Linguistics, 51(4):1419–1431, December
-
[10]
Glicksberg, Panagiotis Korfiatis, Girish N
Yaara Artsi, Vera Sorin, Benjamin S. Glicksberg, Panagiotis Korfiatis, Girish N. Nadkarni, and Eyal Klang. Large language models in real-world clinical workflows: a systematic review of applications and implementation. Frontiers in Digital Health, 7:1659134, 2025. ISSN 2673-253X. doi: 10.3389/fdgth.2025.1659134. URLhttp s://www.frontiersin.org/journals/di...
arXiv 2025
-
[11]
Building Effective AI Agents, 2024
Anthropic. Building Effective AI Agents, 2024. URLhttps://www.anthropic.com/engineering/buildi ng-effective-agents
2024
-
[12]
Sahar Abdelnabi, Kai Greshake, Shailesh Mishra, Christoph Endres, Thorsten Holz, and Mario Fritz. Not What You’ve Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection. In Maura Pintor, Xinyun Chen, and Florian Tramèr, editors,Proceedings of the 16th ACM Workshop on Artificial Intelligence and Security, AISec 20...
arXiv 2023
-
[13]
Defending Against Indirect Prompt Injection Attacks With Spotlighting
Keegan Hines, Gary Lopez, Matthew Hall, Federico Zarfati, Yonatan Zunger, and Emre Kiciman. Defending Against Indirect Prompt Injection Attacks With Spotlighting. In Rachel Allen, Sagar Samtani, Edward Raff, and Ethan M. Rudd, editors,Proceedings of the Conference on Applied Machine Learning in Information Security (CAMLIS 2024), Arlington, Virginia, USA,...
2024
-
[14]
Language Models are Few-Shot Learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-V oss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott G...
1901
-
[15]
V oorhees and Donna K
Ellen M. V oorhees and Donna K. Harman.TREC: Experiment and Evaluation in Information Retrieval (Digital Libraries and Electronic Publishing). The MIT Press, August 2005. ISBN 978-0-262-22073-6. 13 Real-World Evaluation of an AI Agent Drafting Translational Impact SummariesA PREPRINT
2005
-
[16]
Jacob Cohen. A Coefficient of Agreement for Nominal Scales.Educational and Psychological Measurement, 20(1):37–46, April 1960. ISSN 0013-1644. doi: 10.1177/001316446002000104. URLhttps://doi.org/10 .1177/001316446002000104
-
[17]
J. R. Landis and G. G. Koch. The measurement of observer agreement for categorical data.Biometrics, 33(1): 159–174, March 1977. ISSN 0006-341X
1977
-
[18]
Stadnick, Isaac Bouchard, Zeying Du, Lauren Brookman-Frazee, Gregory A
Kera Swanson, Nicole A. Stadnick, Isaac Bouchard, Zeying Du, Lauren Brookman-Frazee, Gregory A. A. Aarons, Emily Treichler, Maryam Gholami, and Borsika A. A. Rabin. An iterative approach to evaluating impact of CTSA projects using the translational science benefits model.Frontiers in Health Services, 5:1535693, 2025. ISSN 2813-0146. doi: 10.3389/frhs.2025...
arXiv 2025
-
[19]
Tejaswini Manjunath, Eline Appelmans, Sinem Balta, Dominick DiMercurio, Claudia Avalos, and Karen Stark. Topic analysis on publications and patents toward fully automated translational science benefits model impact extraction.Frontiers in Research Metrics and Analytics, 10:1596687, September 2025. ISSN 2504-0537. doi: 10.3389/frma.2025.1596687. URLhttps:/...
arXiv 2025
-
[20]
Kristopher Bough, Francisco Leyva, Monica Donerson, and Soju Chang. 243 An AI-driven, Translational Sci- ence Benefits Model (TSBM) approach to assess the real-world impacts of NCATS’ CTSA Collaborative and Innovative Acceleration Award Initiative.Journal of Clinical and Translational Science, 10(s1):81–82, April
-
[21]
Dillon, and Deborah DiazGranados
Andrea Molzhon, Pamela M. Dillon, and Deborah DiazGranados. Leveraging the translational science benefits model to enhance planning and evaluation of impact in CTSA hub-supported research.Frontiers in Public Health, 13:1593920, June 2025. ISSN 2296-2565. doi: 10.3389/fpubh.2025.1593920. URLhttps://www.fr ontiersin.org/journals/public-health/articles/10.33...
arXiv 2025
-
[22]
Treadwell, Jung Min Han, Jesse Wagner, Eric A
Gerald Gartlehner, Shannon Kugley, Karen Crotty, Meera Viswanathan, Andreea Dobrescu, Barbara Nussbaumer-Streit, Graham Booth, Jonathan R. Treadwell, Jung Min Han, Jesse Wagner, Eric A. Apaydin, Erin L. Coppola, Margaret Maglione, Rainer Hilscher, Robert Chew, Meagan Pilar, Bryan Swanton, and Leila C. Kahwati. Artificial Intelligence–Assisted Data Extract...
2025
-
[23]
Yawen Guo, Di Hu, Ziqi Yang, Seungjun Kim, Brian Tran, Jamie Lee, Sitha Vallabhaneni, Rachael Zehrung, Sairam Sutari, Steven Tam, Emilie Chow, Danielle Perret, Deepti Pandita, and Kai Zheng. What do clinicians edit in ambient AI-drafted clinical documentation? A qualitative content analysis.Journal of the American Medical Informatics Association, page oca...
-
[24]
Charlotte M. H. H. T. Bootsma-Robroeks, Jessica D. Workum, Stephanie C. E. Schuit, Anne Hoekman, Tarannom Mehri, Job N. Doornberg, Tom P. van der Laan, and Rosanne C. Schoonbeek. AI-generated draft replies to patient messages: exploring effects of implementation.Frontiers in Digital Health, 7:1588143, June 2025. ISSN 2673- 253X. doi: 10.3389/fdgth.2025.15...
arXiv 2025
-
[25]
Soumik Mandal, Batia M. Wiesenfeld, Adam C. Szerencsy, William R. Small, Vincent Major, Safiya Richard- son, Antoinette Schoenthaler, Devin Mann, and Oded Nov. Utilization of Generative AI-drafted Responses for Managing Patient-Provider Communication.npj Digital Medicine, 8(1):591, October 2025. ISSN 2398-6352. doi: 10.1038/s41746-025-01972-w. URLhttps://...
-
[26]
The Effects of Generative AI on High-Skilled Work: Evidence from Three Field Experiments with Software Developers, August
Zheyuan (Kevin) Cui, Mert Demirer, Sonia Jaffe, Leon Musolff, Sida Peng, and Tobias Salz. The Effects of Generative AI on High-Skilled Work: Evidence from Three Field Experiments with Software Developers, August
-
[27]
Martin Szomszor and Euan Adie. Overton: A bibliometric database of policy document citations.Quantitative Science Studies, 3(3):624–650, November 2022. ISSN 2641-3337. doi: 10.1162/qss_a_00204. URLhttps: //doi.org/10.1162/qss_a_00204
-
[28]
Nicole M Llewellyn, Amber A Weber, Clara M Pelfrey, Deborah DiazGranados, and Eric J Nehl. Translat- ing Scientific Discovery Into Health Policy Impact: Innovative Bibliometrics Bridge Translational Research 14 Real-World Evaluation of an AI Agent Drafting Translational Impact SummariesA PREPRINT Publications to Policy Literature.Academic Medicine, 98(8):...
-
[32]
URLhttps://papers.ssrn.com/abstract=4945566
-
[2020]
URLhttps://papers.nips.cc/paper/2020/hash/1457c0d6bfcb4967418bfb8ac142f64a-Abs tract.html
2020
-
[2025]
ISSN 0891-2017. doi: 10.1162/COLI.a.18. URLhttps://doi.org/10.1162/COLI.a.18
-
[2026]
ISSN 2059-8661. doi: 10.1017/cts.2026.10445. URLhttps://www.cambridge.org/core/journal s/journal-of-clinical-and-translational-science/article/243-an-aidriven-translation al-science-benefits-model-tsbm-approach-to-assess-the-realworld-impacts-of-ncats-cts a-collaborative-and-innovative-acceleration-award-initiative/71AFFBF1BE517C4C3150C4 2E4E9FB5C9
arXiv 2059
-
[2446]
URLhttps://doi.org/10.1097/ACM.0000000000002543
doi: 10.1097/ACM.0000000000002543. URLhttps://doi.org/10.1097/ACM.0000000000002543
-
[9738]
URLhttps://www.jci.org/articles/view/123875
doi: 10.1172/JCI123875. URLhttps://www.jci.org/articles/view/123875
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.