Pith. sign in

REVIEW 2 major objections 6 minor 27 references

Effective, critical, and responsible use of LLMs in research rests on eight competencies, above all domain expertise and oversight of AI outputs.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 21:22 UTC pith:BT2EEOUX

load-bearing objection Useful synthesis of AI-research competencies, but the paper's headline frequency counts rest on pooling two incompatible datasets; the qualitative framework survives, the 'most prevalent' ordering doesn't as currently computed. the 2 major comments →

arxiv 2607.16083 v1 pith:BT2EEOUX submitted 2026-07-17 cs.SE

What Does It Take to Research with AI? A Rapid Review of Competencies to Train LLM-Literate Researchers

classification cs.SE
keywords LLM literacyresearch competenciesrapid reviewAI-assisted researchdomain expertiseacademic integrityprompt engineeringreproducibility
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This rapid review sets out to name what researchers and graduate students must be able to do—not just what AI can do—to use large language models well across the research lifecycle. Analyzing 40 studies drawn from 194 retrieved documents, the authors identify eight competencies and, through pooled manual and AI-assisted coding, find that the most frequently invoked is domain expertise plus oversight of AI outputs: knowing your field well enough to catch hallucinations, verify sources, and keep responsibility for conclusions. The other competencies are metacognitive decisions about when to use AI, ethics and academic integrity, prompt engineering, reproducibility of AI use, methodological design, AI literacy, and data analysis with AI. The point of the synthesis is to show that AI-assisted research is less a technical skill problem than a judgment and accountability problem, which would redirect training away from tool mastery toward epistemic and ethical preparation.

Core claim

On the authors' own account, the study's central discovery is that the literature on AI in research, when coded thematically, converges on eight competencies that span the whole research process: domain expertise and oversight of AI outputs, metacognition and decision making about AI use, ethics, privacy and academic integrity, prompt engineering for research, reproducibility and reporting of AI use, methodological and experimental design with AI, AI literacy and technical knowledge, and data analysis and interpretation with AI. The most prevalent by mention count is domain expertise and oversight, with 123 pooled instances, more than double the next competency. The authors further claim tha

What carries the argument

The load-bearing object is the eight-competency framework itself, built by an inductive thematic analysis of 147 manually extracted competency instances (26 codes, 7 themes) merged with an independent LLM-assisted analysis of 199 occurrences (13 codes, 9 themes) across a smaller article subset; conceptually equivalent themes were consolidated into 8 final competencies, with pooled frequencies (Σn) used as indicators of relative salience. The selection machinery is a dual human screening of 194 articles (two independent groups, inclusion/exclusion criteria, with an agreement statistic AC1 of 0.76–0.83) plus an AI-supported validation pass. The framework does the work of converting scattered o

Load-bearing premise

The whole competency ordering rests on the assumption that a quick, convenience-based sample of 40 articles retrieved from AI-assisted and general web search fairly represents the literature on AI-assisted research competencies, and that simply pooling how often each competency is mentioned across two overlapping analyses tells us which ones truly matter.

What would settle it

Conduct a systematic search across curated bibliographic databases with the same 2022–2025 inclusion criteria, code the retrieved studies with the same thematic protocol, and check whether the same eight competencies emerge and whether domain expertise and oversight still outranks the others; a different ordering or an additional competency would falsify the review's central ranking.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

Share X Bluesky LinkedIn Reddit HN

If this is right

  • Graduate research training should treat AI literacy as one component of a wider competency profile centered on domain expertise, oversight, ethics, and reproducibility, not as a standalone technical subject.
  • Instructor-designed research methods courses and AI literacy initiatives can use the eight competencies as learning objectives and assessment targets.
  • In software engineering research, reporting norms for LLM use—prompts, model versions, configurations, validation—are part of the reproducibility competency and should be incorporated into study guidelines.
  • Researchers remain the accountable agents for knowledge produced with AI; the framework implies that delegating epistemic responsibility to the model is a competency failure, not a workflow choice.
  • The relative salience ordering offers a starting point for prioritising where to invest in researcher training, such as more in oversight and decision-making than in advanced prompt techniques.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Because the paper pools non-de-duplicated mentions from two differently sized analyses, the numerical ordering should be read cautiously; a re-analysis that deduplicates instances might shift the ordering of the lower-ranked competencies, though the dominance of domain expertise is likely robust.
  • If domain expertise is genuinely the prerequisite for critical evaluation, then novice researchers are the most exposed to AI errors; a testable extension would be to compare error-detection rates between domain experts and novices under identical prompt conditions.
  • The same eight-competency structure may transfer to other expert knowledge work, such as medicine, law, or policy analysis, where automated outputs must be verified by domain judgment, suggesting a general professional oversight literacy beyond academic research.
  • An intervention study could test whether training that pairs prompt engineering with source-verification protocols yields larger gains in output quality than prompt training alone, directly testing the paper's central emphasis.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 6 minor

Summary. This rapid review aims to identify and consolidate the competencies required for the effective, critical, and responsible use of large language models (LLMs) in scientific research. The authors retrieved 194 articles from Elicit and Google Scholar (2022–2025), screened them through independent dual human review and an AI-assisted procedure, and selected 40 articles for thematic analysis. Combining 147 manually extracted instances and 199 AI-assisted occurrences, they propose eight competencies, with 'domain expertise and oversight of AI outputs' reported as the most prevalent (Σn=123), followed by metacognition/decision making, ethics/integrity, prompt engineering, reproducibility, methodological design, AI literacy, and data analysis. The paper contends that AI-assisted research depends less on technical AI mastery and more on domain judgment, accountability, and methodological rigor, with implications for graduate education and AI literacy initiatives.

Significance. If the synthesis is accepted, the paper makes a timely and useful contribution by consolidating dispersed discussions on AI literacy, research integrity, reproducibility, and AI-assisted research into a single competency framework. The authors deserve credit for dual independent screening with agreement coefficients, a transparent limitations section, and a promised replication package. The substantive conclusion — that human accountability and domain expertise are central to responsible AI-assisted research — is plausible and important for curriculum design. However, the quantitative 'most prevalent' claim currently rests on non-de-duplicated pooled counts from two analyses of different scale and extraction density, and the manuscript contains an unresolved inconsistency in the AI-assisted screening counts. The qualitative framework is defensible, but the headline ordering needs to be recomputed or explicitly qualified.

major comments (2)
  1. [Abstract; §3.3; Table 1; §4.9] The headline claim that 'domain expertise and oversight of AI outputs' is the most prevalent competency (Σn=123) is based on pooling 147 manual instances across 40 articles with 199 AI-assisted occurrences across 17 articles. As §3.3 itself concedes, the two analyses have different extraction densities (≈3.7 vs ≈11.7 per article) and the counts were not de-duplicated, so the pooled sums are not commensurable across the two pipelines. This is not merely a technical nuance: the abstract and §4.9 use the ordering to support the conclusion that AI-assisted research depends less on technical mastery and more on domain judgment. The AI subset's higher per-article density could by itself drive the 123 total. Please recompute the Σn values from the artifact's raw coding in a de-duplicated and/or per-article-normalized way, report manual and AI counts separately, or explicitly reframe the abstrac
  2. [§3.1; Figure 1(a); §3.2] There is an inconsistency in the AI-assisted screening counts that affects the corpus used for the AI-assisted analysis. The text in §3.1 states that the AI application 'identified 17 articles already included in the manual review and 1 additional article,' which implies an 18-article AI selection. Figure 1(a), however, shows 17 AI-selected articles with 16 overlapping the manual set and 1 AI-only article, and §3.2 says the AI-assisted analysis operated on the 17 articles it selected. These numbers cannot all be correct. Since the article sets feeding the two coding pipelines determine the pooled counts, this discrepancy must be resolved and the corrected numbers reflected in both the text and Figure 1.
minor comments (6)
  1. [Abstract] The abstract says 'This rapid review analyzed 194 articles,' but only 40 articles were selected for extraction and thematic analysis; the 194 were screened, not analyzed in depth. Consider writing 'screened 194 articles and selected 40 for analysis.'
  2. [§3.1 (inclusion criteria)] The sentence 'the paper that not attend the inclusion criteria are excludes' is ungrammatical and should be revised, e.g., 'papers that do not meet the inclusion criteria are excluded.'
  3. [References; §3.1] The reference to Landis and Koch appears as 'Landis JRKoch, G.' and in the text as 'Landis JRKoch 1977,' missing the space and initials. Please format as 'Landis, J. R., and Koch, G. G.' and verify the citation style throughout.
  4. [§3.1] The phrase 'already excluding duplicates also returned by Google Scholar' is ambiguous. State whether the 94 Elicit records are unique after removing duplicates against the Google Scholar results and, ideally, report the number of duplicates removed.
  5. [§7; Artifacts] The artifact identifier 'zenodo.21313656' is not a resolvable link. Please provide a full DOI/URL (e.g., https://doi.org/...) and ensure the coding files, mapping, and AI prompts are accessible to reviewers.
  6. [Figure 1(a)] The 'Selection overlap 24 16 1' row is difficult to parse. Use a Venn diagram or a clearly labeled table with 'manual-only,' 'both,' and 'AI-only' so the reader can verify the 40-article final set.

Circularity Check

0 steps flagged

No significant circularity: the competency claims are an external-literature synthesis; self-citations are present but not load-bearing.

full rationale

The paper's central claim is a synthesis of an external corpus: 40 articles screened by humans and 17 articles selected by an LLM-assisted pass. The eight competencies were generated by thematic analysis and then consolidated by human researchers; the AI-assisted analysis is explicitly characterized as a complementary coding lens rather than an autonomous source of conclusions. The key counts (Σn) are empirical frequencies extracted from the reviewed studies, not parameters fitted to the conclusion, and the paper itself cautions in §3.3 and Table 1 that the pooled counts were not de-duplicated and should be read only as an indication of relative salience, not a precise measure of prevalence. That caveat weakens the abstract's 'most prevalent' wording, but it is a validity limitation rather than a circular reduction: the count 123 is not defined in terms of the claim it supports. The self-citations are minor and non-load-bearing: Cartaxo et al. 2020 is cited for the rapid-review method and includes a co-author of this paper, and Santos et al. 2026 is cited as background on academic integrity concerns; neither supplies the competency framework or the reported ordering. No fitted input is renamed as a prediction, no self-citation chain forces the central claim, and no uniqueness theorem or ansatz is imported from the authors' prior work. Therefore no qualifying circular step exists under the required evidence standard.

Axiom & Free-Parameter Ledger

0 free parameters · 4 axioms · 0 invented entities

The framework introduces no new entities, forces, or mechanisms. Its epistemic load is in the search-representativeness, coding-validity, and prevalence-indexing assumptions above.

axioms (4)
  • domain assumption Primary studies in the review accurately report competencies and practices; their claims are treated as evidence.
    The synthesis codes statements from 40 papers as instances of competencies without independently verifying the underlying empirical claims (§3.2, Table 1).
  • domain assumption The retrieval set from Elicit and first 10 Google Scholar pages is representative of relevant literature.
    This is a stated premise of the rapid review, explicitly flagged as a threat in §3.3.
  • standard math Inter-rater reliability statistics (Gwet AC1, Cohen's kappa, Krippendorff's alpha) measure screening reliability as intended.
    Used in §3.1 to support the reliability of screening; standard statistical assumptions.
  • domain assumption Pooled, un-de-duplicated counts can index relative salience of competencies.
    The paper asserts this in §3.3/§4 but does not justify it. Different extraction densities (3.7 vs 11.7 per article) and double counting make this assumption fragile.

pith-pipeline@v1.3.0-alltime-deepseek · 10782 in / 10964 out tokens · 94036 ms · 2026-08-01T21:22:41.505164+00:00 · methodology

0 comments
Cite this review

Pith. "Pith review of What Does It Take to Research with AI? A Rapid Review of Competencies to Train LLM-Literate Researchers." pith.science (2026). https://pith.science/paper/BT2EEOUX

@misc{pith2026260716083,
  author       = {Pith},
  title        = {Pith review of: What Does It Take to Research with AI? A Rapid Review of Competencies to Train LLM-Literate Researchers},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/BT2EEOUX}},
  note         = {Machine review of arXiv:2607.16083}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

The growing adoption of Large Language Models in scientific research has created a need to understand what competencies researchers and graduate students require to use these tools critically and responsibly. This rapid review analyzed 194 articles retrieved from Elicit and Google Scholar (2022 to 2025), from which 40 were selected for competency extraction and thematic analysis following independent dual screening (Gwet AC1: 0.76 to 0.83). Eight competencies were identified. The most prevalent was domain expertise and oversight of AI outputs (n = 123), encompassing subject matter mastery, systematic skepticism, source verification, and researcher accountability. Other key competencies include metacognition and decision making about AI use (n = 55), ethics and academic integrity (n = 53), prompt engineering for research (n = 38), and reproducibility of AI use (n = 29). AI literacy and technical knowledge (n = 16) was explicitly identified as a risk factor when absent, with domain expertise treated as a prerequisite for meaningful critical evaluation. The findings suggest that preparing researchers to use LLMs goes beyond technical instruction, requiring an integrated set of epistemic, ethical, and methodological competencies centered on human accountability for the knowledge produced. These results have direct implications for the design of graduate programs and AI literacy initiatives.

Figures

Figures reproduced from arXiv: 2607.16083 by Breno Andrade, Danilo Monteiro Ribeiro, Gilberto Hida, Gustavo Pinto, Julia Alencar, Rafael Batista Duarte, Rodrigo Siqueira, Ronnie de Souza Santos.

Figure 1
Figure 1. Figure 1: Overview of the review process Note. (a) Article selection: 194 retrieved articles were screened independently by humans (40 selected) and by an AI-assisted procedure (17 selected). (b) Thematic analysis: 7 manual and 9 AI-assisted themes were merged into 8 consolidated competencies. searchers decided not to include it in the final corpus, as it did not fully meet the accep￾tance criteria. Therefore, the f… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

27 extracted references · 2 linked inside Pith

  1. [1]

    Challenges, and Ethical Considerations (September 02, 2024) , year=

    Navigating the future of large language models in scientific research: Opportunities, challenges, and ethical considerations , author=. Challenges, and Ethical Considerations (September 02, 2024) , year=

  2. [2]

    Proceedings of the AAAI conference on artificial intelligence , volume=

    Towards scientific discovery with generative ai: Progress, opportunities, and challenges , author=. Proceedings of the AAAI conference on artificial intelligence , volume=

  3. [3]

    Proceedings of the National Academy of Sciences , volume=

    Protecting scientific integrity in an age of generative AI , author=. Proceedings of the National Academy of Sciences , volume=. 2024 , publisher=

  4. [4]

    arXiv preprint arXiv:2508.15503 , year=

    Guidelines for empirical studies in software engineering involving large language models , author=. arXiv preprint arXiv:2508.15503 , year=

  5. [5]

    Journal of Education, Management and Development Studies , volume=

    ChatGPT and academic research: A review and recommendations based on practical examples , author=. Journal of Education, Management and Development Studies , volume=

  6. [6]

    BioData mining , volume=

    ChatGPT and large language models in academia: opportunities and challenges , author=. BioData mining , volume=. 2023 , publisher=

  7. [7]

    IEEE access , volume=

    Artificial intelligence in education: A review , author=. IEEE access , volume=. 2020 , publisher=

  8. [8]

    Journal of Computers in Education , pages=

    Understanding influential factors for college instructors’ adoption of LLM-based applications using analytic hierarchy process , author=. Journal of Computers in Education , pages=. 2025 , publisher=

  9. [9]

    Proceedings of the 55th ACM technical symposium on computer science education v

    Software engineering education must adapt and evolve for an llm environment , author=. Proceedings of the 55th ACM technical symposium on computer science education v. 1 , pages=

  10. [10]

    International Conference on the Foundations of Software Engineering (FSE) , year=

    LLM Use, Cheating, and Academic Integrity in Software Engineering Education , author=. International Conference on the Foundations of Software Engineering (FSE) , year=

  11. [11]

    Computers and Education: Artificial Intelligence , volume=

    Conceptualizing AI literacy: An exploratory review , author=. Computers and Education: Artificial Intelligence , volume=. 2021 , publisher=

  12. [12]

    Education in the Knowledge Society (EKS) , volume=

    Research competencies mediated by technologies: A systematic mapping of the literature , author=. Education in the Knowledge Society (EKS) , volume=

  13. [13]

    Procedia-Social and Behavioral Sciences , volume=

    Measuring graduate students research skills , author=. Procedia-Social and Behavioral Sciences , volume=. 2012 , publisher=

  14. [14]

    Empirical Software Engineering Issues

    Empirical software engineering: Teaching methods and conducting studies , author=. Empirical Software Engineering Issues. Critical Assessment and Future Directions: International Workshop, Dagstuhl Castle, Germany, June 26-30, 2006. Revised Papers , pages=. 2007 , organization=

  15. [15]

    Empirical Software Engineering , volume=

    Training students in evidence-based software engineering and systematic reviews: a systematic review and empirical study , author=. Empirical Software Engineering , volume=. 2021 , publisher=

  16. [16]

    Information and software technology , volume=

    Systematic literature reviews in software engineering--a systematic literature review , author=. Information and software technology , volume=. 2009 , publisher=

  17. [17]

    Proceedings

    Evidence-based software engineering , author=. Proceedings. 26th International Conference on Software Engineering , pages=. 2004 , organization=

  18. [18]

    arXiv preprint arXiv:2604.11184 , year=

    Taking a Pulse on How Generative AI is Reshaping the Software Engineering Research Landscape , author=. arXiv preprint arXiv:2604.11184 , year=

  19. [19]

    , author=

    Practical Application of AI and Large Language Models in Software Engineering Education. , author=. International Journal of Advanced Computer Science & Applications , volume=

  20. [20]

    2023 IEEE/ACM International Conference on Software Engineering: Future of Software Engineering (ICSE-FoSE) , pages=

    Large language models for software engineering: Survey and open problems , author=. 2023 IEEE/ACM International Conference on Software Engineering: Future of Software Engineering (ICSE-FoSE) , pages=. 2023 , organization=

  21. [21]

    Proceedings of the 33rd ACM International Conference on the Foundations of Software Engineering , pages=

    Get on the train or be left on the station: Using llms for software engineering research , author=. Proceedings of the 33rd ACM International Conference on the Foundations of Software Engineering , pages=

  22. [22]

    Contemporary empirical methods in software engineering , pages=

    Rapid reviews in software engineering , author=. Contemporary empirical methods in software engineering , pages=. 2020 , publisher=

  23. [23]

    2011 international symposium on empirical software engineering and measurement , pages=

    Recommended steps for thematic synthesis in software engineering , author=. 2011 international symposium on empirical software engineering and measurement , pages=. 2011 , organization=

  24. [24]

    British Journal of Mathematical and Statistical Psychology , volume=

    Computing inter-rater reliability and its variance in the presence of high agreement , author=. British Journal of Mathematical and Statistical Psychology , volume=. 2008 , publisher=

  25. [25]

    , author=

    Weighted kappa: Nominal scale agreement provision for scaled disagreement or partial credit. , author=. Psychological bulletin , volume=. 1968 , publisher=

  26. [26]

    2018 , publisher=

    Content analysis: An introduction to its methodology , author=. 2018 , publisher=

  27. [27]

    Biometrics , volume=

    The measurement of observer agreement for categorical data , author=. Biometrics , volume=