Pith. sign in

REVIEW 2 major objections 4 minor 44 references

Roadmap for using large language models (LLMs) to accelerate cross-disciplinary research with an example from computational biology

T0 review · 2 major / 4 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read This paper argues that large language models accelerate cross-disciplinary research when used as augmentative assistants under expert supervision, and it demonstrates the argument with a computational-biology case study on HIV rebound…

desk verdict A sensible, well-written roadmap; the HIV case study is honest but retrospective, and the 'substantially accelerate' claim outruns the evidence. read the letter →

arxiv 2507.03722 v1 pith:YPMKWOWF submitted 2025-07-04 cs.AI q-bio.OT

classification cs.AIq-bio.OT
keywords largelanguagemodelscross-disciplinaryresearchhuman-in-the-loopChatGPTcomputationalbiologyHIVrebounddynamicsscientificworkflowsacceleration
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that large language models work best in cross-disciplinary research as augmentative assistants inside a human-in-the-loop workflow, rather than as autonomous generators of scientific conclusions. It supports the argument with a worked computational-biology example in which iterative interactions with ChatGPT carry a project from literature review, through data cleaning and statistical testing, to mathematical model construction and manuscript drafting. Across each stage the authors document where LLM output was correct, where it was incomplete or wrong, and where an expert had to correct or verify it. The takeaway is that LLMs can lower the cost of crossing disciplinary boundaries, but only when researchers bring enough domain expertise to steer, check, and debug the outputs.

What carries the argument

The carrying mechanism is an iterative prompt-response-evaluation loop, organized as a roadmap with four stages: literature review and idea generation, data analysis and visualization, method selection and model development, and drafting and polishing. At each stage the LLM performs a routine, language-heavy task, such as synthesizing papers, writing code, suggesting statistical tests, or drafting text, while a domain expert assesses the output and feeds corrections back into the next prompt. The HIV example illustrates the loop's load-bearing role: the LLM's ODE model was "reasonable, but not fully correct," and only became nearly identical to the published model after the authors supplied detailed biological corrections.

What would settle it

A blinded prospective trial would test the claim: give one group of researchers a cross-disciplinary question they have never studied and access to the paper's roadmap with an LLM, give a matched control group traditional tools, and compare the time to a validated analysis and the number of undetected errors. If LLM-assisted teams are not faster or are more error-prone when the correct answer is genuinely unknown, the acceleration claim would not generalize.

Watch

Extended reading notes

Core claim

The central claim is that the most effective and responsible use of LLMs in cross-disciplinary research is as augmentative tools within a human-in-the-loop framework. In the HIV rebound modeling case study, ChatGPT-generated literature syntheses, statistical tests, plotting code, parameter tables, ODE models, and draft manuscript sections; each output was useful but imperfect. Expert corrections, such as supplying the missing target-cell dynamics in the ODE model or choosing which parameters to fit, were required to turn the drafts into a usable analysis. The authors conclude that LLMs can accelerate cross-disciplinary work by translating jargon, generating code, and proposing methods, while the critical tasks of judging, correcting, and taking responsibility for the science remain with human experts.

Load-bearing premise

The case study is retrospective, so the authors already knew the correct model, statistical tests, and answers; the demonstration assumes that users who do not have that prior knowledge can still replicate the same acceleration by following the roadmap.

Editorial extensions

If this is right

  • Researchers new to a field can use LLM-generated literature overviews and plain-language explanations to build shared vocabulary with collaborators from other disciplines, reducing the communication burden of cross-disciplinary teams.
  • LLM-generated code for data cleaning, statistical testing, visualization, and model fitting can cut setup time substantially, but every script must be reviewed and tested by someone who understands the underlying method.
  • Domain experts can quickly steer LLM-suggested models toward correct formulations, compressing the model-development phase of a project.
  • Manuscript drafts and language polishing can be handled faster with LLMs, but the content must be written or revised by the researchers so that results are real and the paper reflects their thinking.
  • As agentic systems improve, more routine tasks may be automated, shifting the human role toward oversight and creative decisions rather than reducing the need for experts.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A direct test of the paper's claim would be a prospective, blinded study in which teams without prior knowledge of a result use the same roadmap on a new problem; this would separate the effect of the tool from the authors' own expertise and hidden target.
  • The human-in-the-loop pattern likely transfers beyond computational biology to any field where domain jargon and code generation are bottlenecks, such as materials science, ecology, or public health, but the specific failure modes will differ by discipline.
  • The roadmap implies an educational use: LLMs could teach students the vocabulary and methods of multiple fields quickly, with the caveat that students must be trained to verify outputs rather than trust fluent text.
  • If future LLMs gain deeper contextual understanding and reliable citation, the case for autonomous multi-agent discovery strengthens, but the paper's own evidence suggests hallucination and code errors will persist, so expert validation is likely to remain the rate-limiting step.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 4 minor

Summary. The manuscript presents a roadmap for integrating large language models (LLMs) into cross-disciplinary research, centered on a human-in-the-loop framework. The authors argue that LLMs should be used as augmentative tools rather than autonomous agents, and they illustrate this through a detailed case study in computational biology: the development of a mathematical model of HIV rebound dynamics using ChatGPT. The paper walks through four stages of research—literature review, data analysis, model development, and manuscript drafting—providing example prompts, describing ChatGPT's outputs, and evaluating their quality. The authors are candid about failures, such as erroneous Monolix code and fabricated details in a draft Results section, and they repeatedly emphasize that domain expertise is essential for verifying and correcting LLM outputs. The central claim is that iterative interactions with LLMs can facilitate interdisciplinary collaboration and accelerate research, provided experts supervise each stage.

Significance. If accepted, this paper serves a useful practical purpose: it offers a structured, stage-by-stage guide for researchers new to using LLMs in interdisciplinary projects, with concrete prompt examples and honest documentation of both strengths and limitations. The case study is not a controlled evaluation—it is a retrospective illustration based on work previously published by the same authors—but it is consistent with the broader literature on LLM capabilities and limitations. The authors explicitly credit human oversight and document multiple instances where expert knowledge was decisive, which strengthens the credibility of their recommendations. The paper does not provide quantitative evidence of acceleration, and its generalizable claims should be tempered accordingly. Overall, the roadmap is sensible and likely to be helpful to its target audience.

major comments (2)
  1. [Abstract; Conclusion and Outlook] The abstract states that responsible LLM use 'will ... substantially accelerate scientific discoveries,' and the Conclusion repeats this expectation. The case study provides only qualitative, retrospective evidence: no wall-clock times, token costs, success rates, or comparison with a non-LLM workflow are reported. Please qualify these strong claims to 'may accelerate' or 'has the potential to accelerate,' and explicitly note that the demonstration is illustrative rather than a quantitative evaluation.
  2. [Example (page 5)] The case study's successful outputs depended heavily on the authors' prior knowledge: they knew the published model [25], supplied detailed expert corrections (e.g., the missing ODE for target cells and the latent reservoir proliferation/death terms), and evaluated every response with full knowledge of the correct answer. The manuscript mentions this in passing ('good knowledge of the different aspects was essential for debugging'), but does not discuss its implications for the roadmap's generalizability to users without such expertise. Please add a paragraph in the Limitations or Discussion section that explicitly addresses the retrospective design and expert-dependence of the demonstration.
minor comments (4)
  1. [Example (page 9)] The statement 'The response from LLMs is accurate' should be relativized to 'In this instance, the response was accurate,' since only a single statistical test is discussed.
  2. [Supplementary Information / Data Availability] The paper cites Supplementary Information containing the full prompts and ChatGPT responses, but no supplementary material statement or data availability section is included in the manuscript. Please add one that specifies how the supplementary material can be accessed.
  3. [Table 1, Category C1] There is a capitalization typo in the example prompt: 'the dataset has a larger number...' should begin with 'The' after a period. A final proofread would be helpful.
  4. [Table 2] Table 2 is not explicitly referenced in the main text; please add a citation where code generation is discussed (e.g., page 8, 'generating codes for data analysis').

Circularity Check

0 steps flagged · score 2.0 of 10

No derivation-level circularity: the central claim is a recommendation supported by a disclosed retrospective case study, with only a minor, non-load-bearing self-reference to the authors' prior paper.

full rationale

The paper makes no formal derivation or prediction that could reduce to its own inputs. Its central claim, that LLMs are best used as augmentative tools within a human-in-the-loop framework, is an argument based on a documented case study, not a quantity fitted from data and later re-labelled as a finding. The case study is retrospective: the authors already knew the modeling outcome from their own prior publication [25] and supplied expert corrections, such as the missing target-cell ODE and the reservoir proliferation and death terms, before the LLM produced a system nearly identical to [25]. This is a real limitation on the example's evidentiary force because the evaluators knew the target and authored the ground truth. However, the paper does not conceal this; it states that the work was originally performed without LLM assistance and published previously [25], and it explicitly notes that the authors already had a clear idea of what they wanted to do and that good knowledge of the different aspects was essential for debugging. Those admissions show the example is offered as an illustration of expert-guided iteration, not as an independent test of the conclusion. No equation is defined in terms of another to produce the claimed result, no fitted parameter is renamed as a prediction, and no uniqueness theorem or ansatz is imported from the authors' prior work to force the choice. The self-citation to [25] is used as a target reference for evaluation, not as the sole justification for the roadmap, so it is not load-bearing in the circularity sense. The lack of quantitative acceleration measurements and the known-answer design are better categorized as evidence-strength concerns than as circular derivation; thus the paper receives a low score for circularity despite legitimate concerns about the strength of the case study.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The paper introduces no free parameters or invented entities. Its central claims rest on unproven assumptions about LLM capabilities, the sufficiency of expert oversight, and generalizability from a single case study.

assumptions (3)
  • domain assumption LLMs such as ChatGPT with Deep Research are capable of performing the literature review, data analysis, code generation, and drafting tasks described in the paper.
    The entire roadmap rests on this capability claim, which is supported only by anecdotal evidence from the authors' case study.
  • domain assumption Expert human oversight is necessary and sufficient to detect and correct LLM errors (hallucinations, code bugs, fabricated results).
    The paper repeatedly asserts human-in-the-loop is needed, but does not provide evidence that this oversight reliably catches all critical errors.
  • domain assumption The case study outcomes generalize to other cross-disciplinary research projects and to other LLMs.
    The paper generalizes from a single HIV rebound modeling project using one proprietary LLM (ChatGPT).

how reviews work

0 comments
Cite this review

Pith. "Pith review of Roadmap for using large language models (LLMs) to accelerate cross-disciplinary research with an example from computational biology." pith.science (2026). https://pith.science/paper/YPMKWOWF

@misc{pith2026250703722,
  author       = {Pith},
  title        = {Pith review of: Roadmap for using large language models (LLMs) to accelerate cross-disciplinary research with an example from computational biology},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/YPMKWOWF}},
  note         = {Machine review of arXiv:2507.03722}
}
read the original abstract

Large language models (LLMs) are powerful artificial intelligence (AI) tools transforming how research is conducted. However, their use in research has been met with skepticism, due to concerns about hallucinations, biases and potential harms to research. These emphasize the importance of clearly understanding the strengths and weaknesses of LLMs to ensure their effective and responsible use. Here, we present a roadmap for integrating LLMs into cross-disciplinary research, where effective communication, knowledge transfer and collaboration across diverse fields are essential but often challenging. We examine the capabilities and limitations of LLMs and provide a detailed computational biology case study (on modeling HIV rebound dynamics) demonstrating how iterative interactions with an LLM (ChatGPT) can facilitate interdisciplinary collaboration and research. We argue that LLMs are best used as augmentative tools within a human-in-the-loop framework. Looking forward, we envisage that the responsible use of LLMs will enhance innovative cross-disciplinary research and substantially accelerate scientific discoveries.

Figures

Figures reproduced from arXiv: 2507.03722 by the authors.

Figure 1
Figure 1. Integration of LLMs into cross-disciplinary research workflows. Effective and responsible use of LLMs with domain expert oversight facilitates communication among researchers from diverse scientific disciplines, guides the selection of suitable computational and statistical methods to analyze heterogeneous datasets, and ultimately accelerates scientific discovery. The figure is adapted from an image generated by Cha… view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

44 extracted references · 38 canonical work pages

  1. [25]

    PLoS Pathog, 2024

    Phan, T., et al., Understanding early HIV-1 rebound dynamics following antiretroviral therapy interruption: The importance of effector cell expansion. PLoS Pathog, 2024. 20(7): p. e1012236

  2. [1]

    Nature, 2023

    Wang, H., et al., Scientific discovery in the age of artificial intelligence. Nature, 2023. 620(7972): p. 47-60

  3. [2]

    Jumper, J., et al., Highly accurate protein structure prediction with AlphaFold. Nature,

  4. [3]

    Language Models are Few-Shot Learners

    Brown, T.B., et al. Language Models are Few-Shot Learners. 2020. arXiv:2005.14165 DOI: 10.48550/arXiv.2005.14165

  5. [4]

    Akhavan, A. and M.S. Jalali, Generative AI and simulation modeling: how should you (not) use large language models like ChatGPT. System Dynamics Review, 2024. 40(3)

  6. [5]

    J Stomatol Oral Maxillofac Surg, 2024

    Alyasiri, O.M., et al., ChatGPT revisited: Using ChatGPT-4 for finding references and editing language in medical scientific articles. J Stomatol Oral Maxillofac Surg, 2024. 125(5S2): p. 101842

  7. [6]

    42(4): p

    Gill, Y., Will AI Write Scientific Papers in the Future? AI Magazine, 2022. 42(4): p. 3-15

  8. [7]

    Digit Discov, 2023

    Jablonka, K.M., et al., 14 examples of how LLMs can transform materials science and chemistry: a reflection on a large language model hackathon. Digit Discov, 2023. 2(5): p. 1233-1250

Show all 44 references
  1. [8]

    Khalifa, M. and M. Albadawy, Using artificial intelligence in academic writing and research: An essential productivity tool. Computer Methods and Programs in Biomedicine Update, 2024. 5

  2. [9]

    BioData Min, 2023

    Meyer, J.G., et al., ChatGPT and large language models in academia: opportunities and challenges. BioData Min, 2023. 16(1): p. 20

  3. [10]

    Clin Transl Sci, 2025

    Lu, J., et al., Large Language Models and Their Applications in Drug Discovery and Development: A Primer. Clin Transl Sci, 2025. 18(4): p. e70205

  4. [11]

    Trotter, and J.Y

    Song, K., A. Trotter, and J.Y. Chen LLM Agent Swarm for Hypothesis-Driven Drug Discovery. 2025. arXiv:2504.17967 DOI: 10.48550/arXiv.2504.17967

  5. [12]

    Bender, E.M., et al., On the Dangers of Stochastic Parrots: Can Language Models Be Too Big? Proceedings of the 2021 Acm Conference on Fairness, Accountability, and Transparency, Facct 2021, 2021: p. 610-623

  6. [13]

    Nature Reviews Physics, 2023

    Birhane, A., et al., Science in the age of large language models. Nature Reviews Physics, 2023. 5(5): p. 277-280

  7. [14]

    Messeri, L. and M.J. Crockett, Artificial intelligence and illusions of understanding in scientific research. Nature, 2024. 627(8002): p. 49-58

  8. [15]

    Bergstrom, C.T. and J. Bak-Coleman, AI, peer review and the human activity of science. Nature, 2025

  9. [16]

    Padmanabhan, and K

    Maleki, N., B. Padmanabhan, and K. Dutta, AI Hallucinations: A Misnomer Worth Clarifying. 2024 Ieee Conference on Artificial Intelligence, Cai 2024, 2024: p. 133-138

  10. [17]

    Bryson, and A

    Caliskan, A., J.J. Bryson, and A. Narayanan, Semantics derived automatically from language corpora contain human-like biases. Science, 2017. 356(6334): p. 183-186

  11. [18]

    Narayanan, A. and S. Kapoor, Why an overreliance on AI-driven modelling is bad for science. Nature, 2025. 640(8058): p. 312-314

  12. [19]

    Res Social Adm Pharm, 2023

    Alqahtani, T., et al., The emergent role of artificial intelligence, natural learning processing, and large language models in higher education and research. Res Social Adm Pharm, 2023. 19(8): p. 1236-1242

  13. [20]

    Stephens, and F

    Seckel, E., B.Y. Stephens, and F. Rodriguez, Ten simple rules to leverage large language models for getting grants. PLoS Comput Biol, 2024. 20(3): p. e1011863

  14. [21]

    PLoS Comput Biol, 2024

    Smith, G.R., et al., Ten simple rules for using large language models in science, version 1.0. PLoS Comput Biol, 2024. 20(1): p. e1011767

  15. [22]

    Am Psychol, 2018

    Hall, K.L., et al., The science of team science: A review of the empirical evidence and research gaps on collaboration in science. Am Psychol, 2018. 73(4): p. 532-548. 18

  16. [23]

    Wang, and J.A

    Wu, L., D. Wang, and J.A. Evans, Large teams develop and small teams disrupt science and technology. Nature, 2019. 566(7744): p. 378-382

  17. [24]

    Jones, and B

    Wuchty, S., B.F. Jones, and B. Uzzi, The increasing dominance of teams in production of knowledge. Science, 2007. 316(5827): p. 1036-9

  18. [26]

    Virus Evol, 2024

    Feng, Y., et al., CovTransformer: A transformer model for SARS-CoV-2 lineage frequency forecasting. Virus Evol, 2024. 10(1): p. veae086

  19. [27]

    AgentCoder: Multi-Agent-based Code Generation with Iterative Testing and Optimisation

    Huang, D., et al. AgentCoder: Multi-Agent-based Code Generation with Iterative Testing and Optimisation. 2023. arXiv:2312.13010 DOI: 10.48550/arXiv.2312.13010

  20. [28]

    1 edition ed

    Lavielle, M., Mixed Effects Models for the Population Approach: Models, Tasks, Methods and Tools. 1 edition ed. 2014, Boca Raton: Chapman and Hall/CRC. 383

  21. [29]

    Nature Reviews Physics, 2024

    AI is no substitute for having something to say. Nature Reviews Physics, 2024. 6(3): p. 151-151

  22. [30]

    Nat Hum Behav, 2024

    Alvarez, A., et al., Science communication with generative AI. Nat Hum Behav, 2024. 8(4): p. 625-627

  23. [31]

    Nature Reviews Physics, 2024

    Biyela, S., et al., Generative AI and science communication in the physical sciences. Nature Reviews Physics, 2024. 6(3): p. 162-165

  24. [32]

    bioRxiv, 2025: p

    Guan, Y., et al., AI-assisted Drug Re-purposing for Human Liver Fibrosis. bioRxiv, 2025: p. 2025.04.29.651320

  25. [33]

    Exploring Large Language Model based Intelligent Agents: Definitions, Methods, and Prospects

    Cheng, Y., et al. Exploring Large Language Model based Intelligent Agents: Definitions, Methods, and Prospects. 2024. arXiv:2401.03428 DOI: 10.48550/arXiv.2401.03428

  26. [34]

    Nature Machine Intelligence, 2024

    Qiu, J.N., et al., LLM-based agentic systems in medicine and healthcare. Nature Machine Intelligence, 2024. 6(12): p. 1418-1420

  27. [35]

    Towards an AI co-scientist

    Gottweis, J., et al. Towards an AI co-scientist. 2025. arXiv:2502.18864 DOI: 10.48550/arXiv.2502.18864. Acknowledgements: We thank Emma Goldberg and many scientists at Los Alamos National Laboratory for helpful discussions. The work was supported by NIH grants R01 -AI152703 (R...

  28. [39]

    Generate Python code to load a CSV file, remove rows with null values, and display the first 5 rows

    Automating Code Generation - Generate code snippets from natural language descriptions "Generate Python code to load a CSV file, remove rows with null values, and display the first 5 rows."

  29. [40]

    Create Python code using pandas and matplotlib to compute basic descriptive statistics and plot a histogram of the 'age' column

    Facilitating Data Analysis & Visualization - Propose scripts for exploratory data analysis - Recommend appropriate visualizations (e.g., histograms, scatter plots) - Summarize data with descriptive stats "Create Python code using pandas and matplotlib to compute basic descript...

  30. [41]

    Provide TensorFlow code to build and train a neural network on the MNIST dataset, including guidance on hyperparameter tuning

    Assisting in Advanced Computational Tasks - Provide scaffolding for simulations and machine learning pipelines - Suggest hyperparameter tuning strategies "Provide TensorFlow code to build and train a neural network on the MNIST dataset, including guidance on hyperparameter tuning."

  31. [42]

    Review the following Python script and suggest improvements to optimize it for faster performance on large datasets

    Debugging and Code Optimization - Identify potential errors or inefficiencies in code - Offer algorithmic improvements "Review the following Python script and suggest improvements to optimize it for faster performance on large datasets."

  32. [43]

    Generate a Python script for data cleaning with detailed inline comments and guidelines for Git-based version control

    Enhancing Reproducibility & Accessibility - Generate well-documented, standardized code - Integrate best practices for version control "Generate a Python script for data cleaning with detailed inline comments and guidelines for Git-based version control."

  33. [44]

    Explain the logic behind this R script in simple terms and provide an equivalent Python version for the same statistical analysis

    Bridging Interdisciplinary Gaps - Translate complex code into more accessible language - Convert scripts between different programming languages - Explain sophisticated algorithms to non-experts "Explain the logic behind this R script in simple terms and provide an equivalent ...

  34. [200]

    Can you write an R script to read in the dataset? B2

    There is missing data in the dataset. Can you write an R script to read in the dataset? B2. Run statistical test and visualize results The second method, i.e. [], you suggested is great. Can you provide R code to read the data from the csv file and perform the test you suggest...

  35. [2021]

    596(7873): p. 583-589

  36. [6000]

    words?’ D2. Revising a manuscript ‘Can you proofread my manuscript (to be published as a research article in the field of []), correct grammar mistakes and revise it to be suitable for a [] audience? Please highlight the revisions you made.’ * [] denotes places where users sha...

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.