Pith. sign in

REVIEW 4 major objections 5 minor 24 references

By analyzing 1,383 Reddit posts, this paper claims that graduate students are systematically outsourcing the cognitive effort that builds research skills to LLMs, ending up with neither the expected results nor the necessary competence.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-03 05:22 UTC pith:TJINGFPE

load-bearing objection A credible thematic synthesis of Reddit posts about LLM use in graduate research, undermined by an abstract that overstates what 54 self-selected posts can support. the 4 major comments →

arxiv 2607.29519 v1 pith:TJINGFPE submitted 2026-07-31 cs.SE

Students' Practices and Skills in the LLM-Era: "You Can't Outsource the Struggle and Still Get the Skill"

classification cs.SE
keywords AI competenciescomputer science educationgrey literature reviewgenerative AIgraduate researchLLM skillscognitive outsourcing
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Drawing on a thematic synthesis of 54 Reddit posts drawn from 1,383 in five research-focused subreddits, this paper argues that graduate students are systematically using LLMs to outsource the cognitive effort that builds research skills, ending up with neither the expected results nor the underlying competence. The paper names five skill families that graduate education should teach: autonomous learning, programming foundations, research skills, LLM context management, and problem solving. The contribution matters because it reframes the AI-in-education debate from a binary of ban versus adopt to a question of which parts of research work can be delegated, which can be supported, and which must remain under human responsibility. If correct, the finding implies that LLMs do not replace the need for foundational training; they amplify its value.

Core claim

The paper's central discovery is that, in the reports of graduate students on research-focused subreddits, the use of LLMs is pervasive but frequently unaccompanied by the validation and judgment needed to turn outputs into knowledge. Students describe using LLMs to fabricate references they do not verify, to generate code they cannot reproduce, and to complete tasks without understanding the content. The paper interprets these reports as a systematic outsourcing of cognitive effort: the struggle that produces skill is being delegated to the tool, so the student gets neither the final quality nor the competence. On the skills side, the paper identifies five families — autonomous learning, pr

What carries the argument

The argument is carried by a thematic synthesis of 54 Reddit posts (from a corpus of 1,383) that codes reports of LLM use into two dimensions: LLM-based practices (RQ1) and LLM-based skills (RQ2). The load-bearing distinction is between delegating a task to an LLM and using an LLM as support while retaining the ability to validate, explain, and reproduce the output. That distinction yields the paper's core normative claim: the cognitive effort of producing and verifying results is the mechanism through which research skills are built, and outsourcing that effort to an LLM leaves the student without the skill.

Load-bearing premise

The 54 Reddit posts that met the inclusion criteria are treated as representative of graduate students' actual LLM practices and skill gaps; if Reddit users are atypical or over-report failures, the 'systematic outsourcing' finding may not generalize to the broader student population.

What would settle it

A representative survey or observational study of graduate students in computing-related research that measures both their LLM usage patterns and their ability to reproduce, debug, or validate outputs without AI assistance would either confirm or refute the claim. If the majority of heavy LLM users still demonstrate the named skills in independent tests, the claim that students systematically outsource and fail to gain competence collapses.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • Graduate programs should teach a named set of skills — autonomous learning, programming foundations, research skills, LLM context management, and problem solving — instead of settling for a ban-or-adopt policy on LLMs.
  • LLM use in research should be governed by a division of labor: tasks can be delegated, supported, or must remain under human responsibility, depending on whether validation and judgment are preserved.
  • The researcher's role shifts from passive code consumer to reviewer, tester, and maintainer of AI-generated artifacts, so curricula should train students to evaluate and verify rather than merely produce.
  • LLMs amplify the value of strong foundational skills rather than replacing them, meaning that foundational training in programming, critical thinking, and research design becomes more important, not less.
  • Validation and hallucination detection are core competencies for graduate research, and should be taught explicitly alongside prompt construction and context management.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If cognitive effort is indeed the mechanism for skill acquisition, a testable corollary is that students who use LLMs for explanation or for 'boring parts' of coding (the paper's own suggestion) would retain skills better than those who delegate entire problem-solving tasks; this could be tested in a controlled learning study.
  • The finding that online discussions emphasize problematic experiences suggests the 'systematic outsourcing' may overstate the prevalence among the general graduate population; a representative survey of actual LLM usage and independent skill demonstration would clarify the baseline.
  • The paper's framework implies that assessment design should shift from evaluating final outputs to evaluating the process of validation — for instance, asking students to find and fix hallucinated references or reproduce a generated result without assistance.
  • The two-dimensional practices-skills taxonomy could be extended to other research fields (e.g., social sciences) where the balance of delegation and skill may differ, offering a comparative test of the outsourcing claim.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. This paper reports a grey-literature review of Reddit posts about graduate students' use of large language models (LLMs) in research. The authors collected 1,383 posts from five research-oriented subreddits, screened them with explicit inclusion/exclusion criteria, and applied thematic synthesis to the 54 posts that survived. They answer two research questions: RQ1 identifies LLM-based practices (e.g., false content and lack of validation, learning to code with AI, automated testing and quality assurance), and RQ2 identifies needed skills (e.g., autonomous learning, programming, research skills, managing LLM context, problem solving). The paper's headline claim is that 'students systematically outsource the cognitive effort required to develop research skills and end up with neither the expected results nor the necessary competence.' The authors discuss implications for graduate curricula, arguing for a middle path between banning and unrestricted adoption of LLMs.

Significance. If the claims are appropriately scoped, the paper provides a useful qualitative map of the kinds of LLM-related practices and skill gaps that surface in graduate-student discussions, and it names a set of competencies that could inform curriculum design. The study is transparent about its search and screening process, and the replication package (Zenodo) is a concrete strength. The chief value is exploratory: it generates hypotheses about how graduate students use LLMs and where they struggle. However, the paper's current significance is undermined by a systematic mismatch between the strength of the headline claims and the inferential power of the design. The data can support claims about what some Reddit users report; they cannot support population-level statements about 'students' or 'systematically' without additional evidence such as base rates, frequencies, or complementary methods. If revised to align the claims with the evidence, the paper could be a useful addition to the AI-education literature.

major comments (4)
  1. [Abstract and §5.1] The central claim in the abstract — 'students systematically outsource the cognitive effort required to develop research skills and end up with neither the expected results nor the necessary competence' — is not supported by the study design. The 54 analyzed posts were selected under inclusion criterion 1 (§2.2), which requires a post to 'discuss AI-related skills, competencies, or practices in the context of research activities.' This selection rule, combined with Reddit's self-selection and the platform's bias toward problematic experiences (as the paper itself concedes in §5.1), guarantees an enrichment for posts about struggles and skill gaps. The abstract also says 'analyzing 1,383 posts,' but only 54 posts were actually analyzed thematically. The authors should reframe the finding as 'some Reddit users report outsourcing' and explicitly state that no prevalence or systematicity cla
  2. [§2.3 and §3] The paper repeatedly refers to the 'most common practices' (§3) and implies a ranking of skills, but it never reports the number of posts supporting each code or category. Frequencies are not required for qualitative thematic synthesis, but if the authors use comparative terms such as 'most common' or claim that a practice is 'systematic,' they must report the underlying counts. In addition, the two authors coded independently (§2.3) but no inter-rater reliability statistic or disagreement-resolution rate is reported. At minimum, the authors should report how many posts contributed to each theme and provide a codebook or coding table, so readers can judge the evidentiary weight for each practice and skill.
  3. [§2.2, footnote 3, and §3] The inclusion criteria in §2.2 state that a post must have been 'posted between January 2020 and April 2026' (criterion 4). However, the year distribution in §3 includes two posts from 2019 and individual posts from other pre-2020 years. Footnote 3 says that after initial screening the authors decided that posts not meeting the date criterion could be included if they met all other criteria. This post-hoc protocol deviation is not justified and undermines the reproducibility of the screening process. The authors should either enforce the original date criterion, explicitly revise it in the methods section, or explain why the pre-2020 posts were necessary. As written, the methods and results are internally inconsistent on a procedural point that is part of the study's stated rigor.
  4. [§2.2 and §3] There is a partial circularity between the inclusion criterion and the research questions. RQ1 asks about 'LLM-based practices' and RQ2 asks about 'skills,' while inclusion criterion 1 requires posts to 'discuss AI-related skills, competencies, or practices in the context of research activities.' The answers to RQ1 and RQ2 are therefore built into the selection rule. This is not necessarily fatal for a scoped qualitative review, but the authors should acknowledge that their findings describe only posts that already self-identified as being about AI-related skills or practices. The current framing presents the practices and skills as emerging independently from the data, which overstates the inductive power of the analysis.
minor comments (5)
  1. [§2.2] The categories 'AI tools' and 'Non-CC' are used in the screening counts but never defined. The reader has to infer that 'AI tools' means posts about AI tools without skill-gap discussion and 'Non-CC' means non-cognitive-content or unrelated posts. Please define these labels in the text.
  2. [Figure 1] The caption contains a typo: 'T otal vs Included Posts.' Also, the figure does not show the classification breakdown (included, excluded, AI tools, Non-CC, uncertain) by subreddit; reporting those counts would help readers assess the screening process.
  3. [§3] The sentence 'By year, there were thirty-three posts in 2026, six in 2025, five in 2024, three in 2022, two in 2019, and one post in each of the remaining years' does not sum to 54 unless 'remaining years' includes 2020, 2021, and 2023 with one post each (33+6+5+3+2+1+1+1 = 52). Please clarify the exact counts per year.
  4. [General] Several phrases, such as 'Due to space limitations' in §3 and §4, are more appropriate for a conference version than a journal article. The authors should either expand the presentation of all identified practices and skills or explain the editorial constraint.
  5. [References] Some references are incomplete or inconsistently formatted. For example, reference [16] lists the DOI as '10.5753/jserd.2019' without the full article DOI, and reference [20] is only tangentially related to the text. A careful copyedit of the reference list is needed.

Circularity Check

1 steps flagged

Partial self-definitional circularity: the inclusion rule requires posts to discuss AI skills/practices and skill gaps, so the class-level RQ1/RQ2 findings are precommitted; only the specific thematic labels are emergent.

specific steps
  1. self definitional [Section 2.2 (selection criteria); Section 1 (research questions); Abstract]
    "a post is included if it satisfies all of the following items: 1) it discusses AI-related skills, competencies, or practices in the context of research activities ... a post is excluded if it matches at least one of the following items: ... 2) it discusses AI tools without addressing skill gaps or research practices. ... RQ1: What are the LLM-based practices that graduate students employ in their research? RQ2: What are the skills that graduate students must have ..."

    The final corpus is constructed from posts that already discuss AI-related skills/competencies/practices and that address skill gaps or research practices. Therefore a thematic synthesis of that corpus cannot fail to find that graduate students exhibit AI-related practices and need skills—the class-level answer to RQ1/RQ2 is written into the sampling rule. The abstract's conclusion ('students systematically outsource the cognitive effort required to develop research skills and end up with neither the expected results nor the necessary competence') presents this selection-enforced topical content as an empirical discovery. The individual categories (false content, autonomous learning, programming, etc.) are not verbatim in the inclusion rule, so the circularity is partial rather than total.

full rationale

The paper is a qualitative grey-literature synthesis: 1,383 posts retrieved, 54 coded, themes derived. The strongest circularity issue is the selection rule in Section 2.2, which requires posts to discuss the exact constructs the research questions ask about (AI-related practices/skills, skill gaps). The abstract then reports these presence-of-practice/skill findings as a population-level discovery, while Section 5.1 itself concedes that Reddit users may not be representative and that online discussions overrepresent problematic experiences. That concession confirms an external-validity limitation, which is closely related to the selection issue but not identical to a derivation circle. The specific thematic codes are not forced by the inclusion criterion, so the paper retains independent content. Self-citations by co-author Gustavo Pinto (references [15], [16]) support grey-literature credibility but are not the load-bearing evidence; external guidelines [8] and Reddit-methodology literature [7], [18] are also cited. No fitted parameters, equations, or uniqueness theorems are involved. Overall: partial self-definitional circularity at the class level (score 5), not full circularity.

Axiom & Free-Parameter Ledger

0 free parameters · 4 axioms · 0 invented entities

No numerical free parameters are fitted; the hand-chosen design decisions (five subreddits, eighteen queries, date range, inclusion/exclusion criteria) are recorded as domain assumptions because they shape the corpus. No new entities are introduced; this is a qualitative study.

axioms (4)
  • domain assumption Reddit posts are candid, self-reported accounts of real graduate-student experiences and skill gaps.
    The entire analysis treats pseudonymous posts as evidence about actual student behavior; Section 5.1 concedes this cannot be verified.
  • domain assumption The inclusion criterion that a post 'discusses AI-related skills, competencies, or practices in the context of research activities' yields a corpus appropriate for answering RQ1 and RQ2.
    Section 2.2, item 1; this choice pre-selects posts about practices/skills, shaping the findings.
  • domain assumption Thematic synthesis of 54 posts by two coders, with disagreements resolved by discussion, produces reliable categories without a reported inter-rater reliability measure.
    Section 2.3; no agreement statistic (e.g., Cohen's kappa) is reported.
  • domain assumption Posts from r/MachineLearning, r/academia, r/GradSchool, r/PhD, and r/compsci are sufficiently representative of graduate students doing empirical CS/SE research.
    Section 2.1; authors acknowledge Reddit users are not representative of the broader graduate-student population (Section 5.1).

pith-pipeline@v1.3.0-daily-deepseek · 8747 in / 11387 out tokens · 102626 ms · 2026-08-03T05:22:11.765272+00:00 · methodology

0 comments
read the original abstract

Generative AI tools have been rapidly learned in the daily workflow of graduate students in Software Engineering, but little is known about what AI-related skills they actually need for effective use in empirical research. Without this understanding, graduate programs cannot prepare students to conduct rig-orous research in the LLM era, risking creating a generation of researchers who delegate tasks without the necessary expertise. By analyzing 1,383 posts from five research-focused subreddits, we found that students systematically outsource the cognitive effort required to develop research skills and end up with neither the expected results nor the necessary competence. Naming these missing skills is the first step toward curricula that teach graduate students to work \emph{with} LLMs without being replaced by them.

Figures

Figures reproduced from arXiv: 2607.29519 by Danilo Monteiro, Enne Rebeca Silva de Freitas, Gustavo Pinto.

Figure 1
Figure 1. Figure 1: Total vs included posts by subreddit. 3 RQ1: The LLM-based practices Due to space limitations, we will discuss the three most common practices below. False content and lack of validation This category highlights the use of LLMs as assistants in fabricating false references and citations in a “student’s graduation project. [...] Out of ten sources, four led nowhere – broken links, nonexistent books, actual … view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

24 extracted references · 2 canonical work pages

  1. [1]

    Generative AI literacy: Twelve defining competencies.Digital Government: Research and Practice, 6(1): 8 1–21, 2025

    Ravinithesh Annapureddy, Alessandro Fornaroli, and Daniel Gatica-Perez. Generative AI literacy: Twelve defining competencies.Digital Government: Research and Practice, 6(1): 8 1–21, 2025. doi: 10.1145/3685680

  2. [2]

    Generative AI can harm learning.SSRN Electronic Journal, 2024

    Hamsa Bastani, Osbert Bastani, Alp Sungu, Haosen Ge, Özge Kabakcı, and Rei Mariman. Generative AI can harm learning.SSRN Electronic Journal, 2024. doi: 10.2139/ssrn.4895486. The Wharton School Research Paper

  3. [3]

    Connell Pensky, Jordan H

    Allison E. Connell Pensky, Jordan H. Usdan, and Harley Chang. Generative AI’s impact on graduate student professional writing productivity and quality.International Journal of Artificial Intelligence in Education, 2025. doi: 10.1007/s40593-025-00528-z

  4. [4]

    Cruzes and Tore Dybå

    Daniela S. Cruzes and Tore Dybå. Recommended steps for thematic synthesis in software engineering. In2011 International Symposium on Empirical Software Engineering and Mea- surement (ESEM ’11), pages 275–284. IEEE, 2011. doi: 10.1109/ESEM.2011.36

  5. [5]

    so what if ChatGPT wrote it?

    Yogesh K. Dwivedi, Nir Kshetri, Laurie Hughes, Emma Louise Slade, Anand Jeyaraj, Arpan Ku- mar Kar, Abdullah M. Baabdullah, Alex Koohang, Vishnupriya Raghavan, Manju Ahuja, et al. Opinion paper: “so what if ChatGPT wrote it?” multidisciplinary perspectives on opportunities, challenges and implications of generative conversational AI for research, practice...

  6. [6]

    Yizhou Fan, Luzhen Tang, Huixiao Le, Kejie Shen, Shufang Tan, Yueying Zhao, Yuan Shen, Xinyu Li, and Dragan Gaševi´c. Beware of metacognitive laziness: Effects of generative artificial intelligence on learning motivation, processes, and performance.British Journal of Educational Technology, 56(2):489–530, 2025. doi: 10.1111/bjet.13544

  7. [7]

    Gilbert, and Naiyan Jones

    Casey Fiesler, Michael Zimmer, Nicholas Proferes, Sarah A. Gilbert, and Naiyan Jones. Remem- ber the human: A systematic review of ethical considerations in reddit research.Proceedings of the ACM on Human-Computer Interaction, 8(GROUP):1–33, 2024. doi: 10.1145/3633070

  8. [8]

    Vahid Garousi, Michael Felderer, and Mika V . Mäntylä. Guidelines for including grey literature and conducting multivocal literature reviews in software engineering.Information and Software Technology, 106:101–121, 2019. doi: 10.1016/j.infsof.2018.09.006

  9. [9]

    AI tools in society: Impacts on cognitive offloading and the future of critical thinking.Societies, 15(1):6, 2025

    Michael Gerlich. AI tools in society: Impacts on cognitive offloading and the future of critical thinking.Societies, 15(1):6, 2025. doi: 10.3390/soc15010006

  10. [10]

    Nitter, jeito alternativo de acessar o X, é descontinuado [in portuguese]

    Rodrigo Ghedin. Nitter, jeito alternativo de acessar o X, é descontinuado [in portuguese]. Núcleo Jornalismo. https://nucleo.jor.br/curtas/2024/01/29/nitter-twitter-x-fim/ ,

  11. [11]

    LLM-assisted empirical software engineering: Systematic literature review and research agenda, 2026

    Victoria Gomes, Delaney Selb, Fabio Palomba, Rodrigo Spinola, and David Lo. LLM-assisted empirical software engineering: Systematic literature review and research agenda, 2026. URL https://arxiv.org/abs/2604.26192

  12. [12]

    Susmita Haldar, Mary Pierce, and Luiz Fernando Capretz. Exploring the integration of gen- erative AI tools in software testing education: A case study on ChatGPT and Copilot for preparatory testing artifacts in postgraduate learning.IEEE Access, 13:46070–46090, 2025. doi: 10.1109/ACCESS.2025.3545882

  13. [13]

    LLM-as-a-judge for software engineering: Literature review, vision, and the road ahead.ACM Transactions on Software Engineering and Methodology, 2026

    Junda He, Jieke Shi, Terry Yue Zhuo, Christoph Treude, Jiamou Sun, Zhenchang Xing, Xiaoning Du, and David Lo. LLM-as-a-judge for software engineering: Literature review, vision, and the road ahead.ACM Transactions on Software Engineering and Methodology, 2026. doi: 10.1145/3797276

  14. [14]

    The next frontier of LLM applications: Open ecosystems and hardware synergy.CoRR, abs/2503.04596, 2025

    Xinyi Hou, Yanjie Zhao, and Haoyu Wang. The next frontier of LLM applications: Open ecosystems and hardware synergy.CoRR, abs/2503.04596, 2025. doi: 10.48550/arXiv.2503.04 596

  15. [15]

    On the use of grey literature: A survey with the brazilian software engineering research community.Brazilian Symposium on Software Engineering (SBES), 2020

    Fernando Kamei, Igor Wiese, Gustavo Pinto, Márcio Ribeiro, and Sérgio Soares. On the use of grey literature: A survey with the brazilian software engineering research community.Brazilian Symposium on Software Engineering (SBES), 2020. doi: 10.1145/3422392.3422442. 9

  16. [16]

    Assessing the credibility of grey literature: A study with brazilian software engineering researchers.Journal of Software Engineering Research and Development, 2022

    Fernando Kamei, Igor Wiese, Gustavo Pinto, Waldemar Ferreira, Márcio Ribeiro, Renata Souza, and Sérgio Soares. Assessing the credibility of grey literature: A study with brazilian software engineering researchers.Journal of Software Engineering Research and Development, 2022. doi: 10.5753/jserd.2019

  17. [17]

    Your brain on ChatGPT: Accumula- tion of cognitive debt when using an AI assistant for essay writing task.arXiv preprint arXiv:2506.08872, 2025

    Nataliya Kosmyna, Eugene Hauptmann, Ye Tong Yuan, Jessica Situ, Xian-Hao Liao, Ashly Vi- vian Beresnitzky, Iris Braunstein, and Pattie Maes. Your brain on ChatGPT: Accumula- tion of cognitive debt when using an AI assistant for essay writing task.arXiv preprint arXiv:2506.08872, 2025. doi: 10.48550/arXiv.2506.08872

  18. [18]

    Studying Reddit: A systematic overview of disciplines, approaches, methods, and ethics.Social Media + Society, 7(2), 2021

    Nicholas Proferes, Naiyan Jones, Sarah Gilbert, Casey Fiesler, and Michael Zimmer. Studying Reddit: A systematic overview of disciplines, approaches, methods, and ethics.Social Media + Society, 7(2), 2021. doi: 10.1177/20563051211019004

  19. [19]

    Griswold, and Adal- bert Gerald Sosai Raj

    Anshul Shah, Anya Chernova, Elena Tomson, Leo Porter, William G. Griswold, and Adal- bert Gerald Sosai Raj. Students’ use of GitHub Copilot for working with large code bases. In Proceedings of the 56th ACM Technical Symposium on Computer Science Education (SIGCSE TS 2025), page 7, 2025. doi: 10.1145/3641554.3701800

  20. [20]

    Creating large language model applications utilizing LangChain: A primer on developing LLM apps fast

    Oguzhan Topsakal and Tahir Cetin Akinci. Creating large language model applications utilizing LangChain: A primer on developing LLM apps fast. InInternational Conference on Applied Engineering and Natural Sciences (ICAENS), volume 1, pages 1050–1056, 2023. doi: 10.592 87/icaens.1127

  21. [21]

    Taking a pulse on how gen- erative AI is reshaping the software engineering research landscape, 2026

    Bianca Trinkenreich, Fabio Calefato, Kelly Blincoe, Viggo Tellefsen Wivestad, Antonio Pedro Santos Alves, Júlia Condé Araújo, Marina Condé Araújo, Paolo Tell, Marcos Kali- nowski, Thomas Zimmermann, and Margaret-Anne Storey. Taking a pulse on how gen- erative AI is reshaping the software engineering research landscape, 2026. URL https: //arxiv.org/abs/2604.11184

  22. [22]

    Yoshija Walter. Embracing the future of artificial intelligence in the classroom: The relevance of AI literacy, prompt engineering, and critical thinking in modern education.International Journal of Educational Technology in Higher Education, 21(1):15, 2024. doi: 10.1186/s41239 -024-00448-3

  23. [23]

    Promises and challenges of generative artificial intelligence for human learning.Nature Human Behaviour, 8(10): 1839–1850, 2024

    Lixiang Yan, Samuel Greiff, Ziwen Teuber, and Dragan Gaševi ´c. Promises and challenges of generative artificial intelligence for human learning.Nature Human Behaviour, 8(10): 1839–1850, 2024. doi: 10.1038/s41562-024-02004-5. 10

  24. [2024]

    Accessed: 2026-05-22