Pith. sign in

REVIEW 4 major objections 6 minor 22 references

Structured assignment design steers students to use AI for requirements analysis and critique, not automation, with clearer gains on concrete quality criteria than on interpretive ones.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-31 15:34 UTC pith:UY3WB72U

load-bearing objection Solid SE-education empirical study: staged TPACK assignment produces clear selective-AI behavior; learning-gain claims are weaker without a control arm. the 4 major comments →

arxiv 2607.28176 v1 pith:UY3WB72U submitted 2026-07-30 cs.SE cs.AI

Integrating AI into Requirements Quality Learning in Software Engineering Education: A TPACK-Guided Empirical Study

classification cs.SE cs.AI
keywords requirements engineering educationAI tooluser story qualityTPACKINVEST frameworksoftware engineering educationgenerative AImixed-methods
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper asks how generative AI should be built into requirements-engineering education so that students still learn to judge quality rather than outsource the work. In a master’s course assignment, a multi-agent AI tool was embedded in a fixed sequence: students first improved user stories by hand against the INVEST quality criteria, then generated and refined AI alternatives, compared the two, peer-reviewed, and reflected. Across 72 analysed submissions, students approved only about half of AI stories, refined most of them, and treated outputs as drafts. Measured alignment with instructor reference judgments rose most for structurally concrete criteria such as value and testability; negotiability did not improve and sometimes fell. Students reported conditional trust and greater awareness of quality rules, alongside usability friction. The claim is that aligning technology, teaching sequence, and content knowledge can make AI a scaffold for analytical skill rather than a substitute for it.

Core claim

When a multi-agent AI requirements tool is introduced only after manual revision and is wrapped in comparison, approval, peer review, and reflection anchored in INVEST, master’s students use it selectively for analysis and evaluation rather than as an automation engine; learning alignment improves most on concrete quality dimensions (valuable, testable) while interpretive ones such as negotiable show mixed or negative change, and students report conditional trust with active refinement.

What carries the argument

TPACK-guided assignment workflow: a staged sequence (manual INVEST revision first, then multi-agent generation, iterative refine-and-approve, contrastive comparison, justified peer review, and structured reflection) that forces technological affordances to serve pedagogical goals and requirements content knowledge.

Load-bearing premise

Changes in how students score four fixed user stories before and after the assignment are mainly caused by the AI-integrated workflow, not by the rest of the course, practice effects, or students skipping the manual-first steps.

What would settle it

Run the same pre/post INVEST scoring on four user stories with a parallel cohort that gets identical lectures and INVEST practice but no AI tool and no contrastive workflow; if their alignment gains match or exceed the AI cohort, the central attribution fails. Separately, log whether students actually completed manual revision before first AI use; if heavy skippers show the same selective-use and learning pattern, the scaffolding claim weakens.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • RE and SE courses should sequence human reasoning before AI generation if the goal is critical evaluation rather than speed.
  • AI-generated alternatives work best as contrastive stimuli scored against an explicit quality framework such as INVEST, not as final deliverables.
  • Instruction and scaffolding must be stronger for interpretive criteria (e.g., negotiability) than for structurally concrete ones (value, testability).
  • Peer review plus mandatory justification and reflection can keep student ownership of quality judgments when AI is present.
  • Design guidance for “responsible” AI in RE education can be stated as workflow principles rather than tool features alone.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The same manual-first, contrastive pattern could transfer to other SE artifacts (architecture decisions, test oracles, design rationales) where LLMs produce fluent but shallow drafts.
  • Negotiability’s flat or negative alignment may mean AI fluency raises awareness of ambiguity without supplying shared exemplars; courses may need annotated borderline cases, not more generation.
  • If commercial tools lack explicit multi-stakeholder agents and approval gates, instructors may need to simulate those gates with rubrics and forced comparison steps.
  • Longitudinal follow-up on later projects would test whether short-assignment conditional trust becomes durable analytical habit or fades once scaffolding is removed.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. This paper reports a mixed-methods study (N=100 enrolled; 72 consented submissions) of a TPACK-guided integration of a multi-agent AI requirements tool into a master’s RE assignment on user-story quality (INVEST). The assignment sequences manual revision before AI generation, then iterative refinement, comparison, peer review, and reflection. The authors claim that this scaffolding shaped selective, evaluative AI use (median 56% approval; frequent refinements) rather than automation; that pre/post agreement with instructor reference analyses improved most for structurally concrete dimensions (Valuable, Testable) while Negotiable showed non-positive change (Table II); and that students reported conditional trust, active editing, and moderate usability friction. Design principles for responsible AI integration in RE education are offered.

Significance. The work addresses a genuine gap: AI adoption in SE/RE education is widespread but often weakly grounded in instructional theory. Applying TPACK as both design principle and analytic lens, with a replicable staged workflow, interaction metrics (Table I), dimension-specific pre/post alignment (Table II), and triangulated reflections, is a concrete contribution beyond perception-only studies. The differential finding—stronger support for structurally explicit quality attributes than for interpretive ones such as negotiability—is pedagogically useful if it holds. Strengths include voluntary participation scale, transparent threats discussion (§6.5), and actionable design principles (manual-first sequencing, contrastive analysis, quality-framework anchoring). The main limit on significance is causal attribution of learning gains, which the manuscript itself flags.

major comments (4)
  1. [§5.2, Table II, §6.5] §5.2, Table II, and the RQ2 claim: learning impact is operationalized as Δ = p_after − p_pre agreement with instructor/TA reference labels on four fixed user stories. §6.5 correctly notes there is no control group, concurrent INVEST exposure in lectures/other assignments, and unverified adherence to the manual-first sequence (§4.1.1, Fig. 2). As written, the abstract and §6.1–6.2 still attribute dimension-specific “alignment improvements” primarily to the AI-integrated workflow. That attribution is load-bearing for half the central claim and is not supported by the design. Either (a) reframe RQ2 results strictly as pre/post association under a confounded course trajectory, with causal language removed from abstract/conclusion, or (b) add analyses that strengthen isolation (e.g., dose–response with refinement/time metrics, comparison to non-consenting or non-completing students if data ex
  2. [§4.3, Table II] §4.3 and Table II: only agreement proportions and raw Δ are reported; no uncertainty (CIs), no paired tests, and no correction for multiple dimension×story cells. With N=72 and story-specific mixed signs (e.g., Negotiable ΔUS2 = −0.292; Testable ΔUS1 = +0.108), readers cannot judge whether shifts exceed sampling noise or practice effects on the same four items. At minimum, report paired proportion tests or bootstrap CIs per cell and an overall sensitivity check; otherwise state explicitly that results are exploratory pattern description, not inferential evidence of learning gain.
  3. [§5.2, §6.2, Fig. 5] §5.2–§6.2 construct validity: treating instructor-defined INVEST reference labels as the alignment target is especially fragile for interpretive attributes (Negotiable), where the paper itself argues contextual judgment is required and where perceived support diverges from measured alignment (Fig. 5 vs Table II). The manuscript notes this in threats but still frames non-positive Negotiable Δ as a substantive finding about AI affordances. Clarify that divergence may reflect legitimate interpretive variance or reference-standard narrowness rather than failed learning, and avoid equating “agreement with reference” with “understanding/application of quality criteria” without qualification in RQ2 wording and abstract.
  4. [§3.2, §5.1, §6.5] §3.2, §6.5: the multi-agent tool is the authors’ research prototype ([3][5]), used in the same institutional setting. Interaction metrics (approval rate, refinement rounds) and usability themes may partly reflect prototype-specific behavior (e.g., product-owner bias quoted in §5.3) rather than general multi-agent AI affordances under TPACK scaffolding. The external-validity paragraph acknowledges the prototype but underplays how this couples TK evaluation with the pedagogical claim. Separate more clearly: (i) evidence that the staged workflow produced selective use, which is stronger, from (ii) claims about multi-agent AI tools in general. If possible, report any configuration logs (agent roles chosen) to show the multi-agent feature was actually enacted beyond generation.
minor comments (6)
  1. [Abstract, §1] Abstract and §1: “Alignment improvements were most evident for… value articulation and testability” overstates Table II, where Valuable is small/mixed (including −0.041 on US1) and gains are story-dependent. Soften to match the table.
  2. [§5.1] §5.1: “wheras” → “whereas”; “indicatesthat” → “indicates that”.
  3. [§3.2, Fig. 1] Figure 1 caption and body refer to “REQ” inconsistently with “RE course”; standardize terminology.
  4. [§4.3, Table III] Table III theme counts are useful; briefly state coding reliability procedure (single coder vs dual coding, agreement) given thematic analysis citation [20].
  5. [§2] §2: a few citation/typo glitches (e.g., “Guardadoet al.”, “Sahet al.”) need spacing fixes.
  6. [§4.1] Keywords and title emphasize TPACK; a short explicit mapping table (CK/PK/TK → assignment elements) in §4.1 would help readers who are not TPACK-fluent without adding length.

Circularity Check

1 steps flagged

No significant circularity: empirical mixed-methods study; student interaction and pre/post INVEST alignment are not forced by construction or by a load-bearing self-citation chain.

specific steps
  1. self citation load bearing [§3.2 A Multi-Agent AI Tool; refs [3][5]; also §6.5]
    "To support students' learning in the REQ, the course integrated a multi-agent AI-based requirements assistance tool... The tool [3][5] was developed as a research prototype in the Software Engineering Research Center (TASE) at Tampere University. ... Moreover, the AI tool was developed within the same research environment as the course, potentially introducing contextual bias"

    The intervention is the authors' own multi-agent prototype, justified by overlapping prior publications rather than an independent external system. This is institutional self-reference, not a derivation that makes student approval rates or INVEST Δ tables true by construction; outcomes remain free to vary. Treated as minor (not load-bearing) circularity risk only.

full rationale

This paper is an educational empirical study, not a first-principles derivation. Its central claims (selective AI use under staged TPACK scaffolding; dimension-specific INVEST alignment changes; conditional trust) rest on new observational data: interaction metrics (median 56% approval, refinement rounds), pre/post agreement deltas on four fixed user stories against an instructor/TA reference, Likert distributions, and thematic coding of reflections. Those outcomes are not definitionally equivalent to the inputs: students could have approved nearly all AI stories, made zero refinements, or shown uniform/no alignment change. Self-citation of the multi-agent tool prototype ([3][5], §3.2) and same-lab development are normal intervention description and are already flagged by the authors as a conclusion/external-validity threat (§6.5); they do not force the behavioral or learning results. The instructor-defined INVEST reference is a construct-validity choice (benchmark), not a circular reduction of prediction to fit. Absence of a control group and unverified workflow adherence weaken causal attribution but are not circularity. Score 1 only for minor same-lab tool self-reference that is not load-bearing on the educational claims.

Axiom & Free-Parameter Ledger

2 free parameters · 5 axioms · 1 invented entities

Load-bearing commitments are educational and measurement assumptions, not fitted physical constants. The central interpretation depends on treating TPACK alignment, INVEST reference labels, and the staged manual-first workflow as the right mediators of 'effective integration,' and on reading descriptive student behavior as evidence of pedagogical shaping.

free parameters (2)
  • Instructor/TA reference INVEST labels for US1–US4 = Reference analysis established by instructor and TAs (figshare link; not numerically itemized in text)
    Pre/post alignment Δ is defined against this benchmark; alternative defensible labels would change reported learning impact.
  • Approval and refinement operational thresholds/metrics = Median 16.5 generated, 9.5 approved, 56% approval, 1.5 refinement rounds
    Median generated/approved counts and 'edited' self-reports summarize enactment; coding boundaries affect the selective-use narrative.
axioms (5)
  • domain assumption Effective educational technology use requires deliberate alignment of technological, pedagogical, and content knowledge (TPACK).
    Used as both design principle and analytical lens from introduction through discussion (§1, §2.2, §4.1).
  • domain assumption INVEST (and related RE quality guidance) is an appropriate content standard for judging user-story quality in this course.
    Assignment scoring, peer review, and pre/post instruments are anchored in INVEST (§4.1, §5.2).
  • ad hoc to paper Sequencing manual revision before AI generation reduces automation-oriented use and preserves analytical responsibility.
    Core design hypothesis of the staged workflow (§4.1.1 Steps 1–3; §6.1 design principles).
  • ad hoc to paper Agreement change with instructor reference on four stories is a valid proxy for understanding/application of quality criteria.
    Primary RQ2 outcome definition (§4.3, §5.2); authors note construct limits for interpretive attributes in §6.5.
  • standard math Thematic coding of reflections plus descriptive interaction metrics can triangulate pedagogical enactment.
    Standard mixed-methods education research practice cited via Braun & Clarke and Creswell (§4.3).
invented entities (1)
  • TPACK-guided multi-agent AI assignment workflow for RE quality learning (manual-first → generate → refine/approve → compare → peer review → reflect) no independent evidence
    purpose: Operationalize responsible AI integration so AI acts as analytical scaffold rather than automation substitute.
    The specific six-step orchestration and role assignment is the paper's design object; not a new physical entity but a new instructional configuration evaluated empirically.

pith-pipeline@v1.2.0-daily-grok45 · 17753 in / 3230 out tokens · 53696 ms · 2026-07-31T15:34:03.304976+00:00 · methodology

0 comments
read the original abstract

The rapid adoption of generative Artificial Intelligence (AI) in software engineering (SE) practice creates a need for pedagogically grounded approaches to AI integration in SE education, especially in conceptually intensive subjects such as requirements engineering (RE). This study examines a TPACK-guided integration of a multi-agent AI tool into a master-level RE assignment on requirements quality analysis. Using a mixed-methods design (N=100; 72 submissions analysed), we examine how structured assignment design shaped students' AI use, affected their understanding of user story quality criteria, and influenced their perceptions of AI's benefits and limitations. Results show that students used the AI tool selectively, mainly as support for analysis and evaluation rather than automation. Alignment improvements were most evident for structurally concrete requirements quality dimensions, such as value articulation and testability, while negotiability showed mixed effects. Students reported conditional trust, active refinement, and increased awareness of quality criteria, alongside moderate usability challenges. The findings show that TPACK-guided scaffolding can align AI affordances with pedagogical goals and RE content, offering design guidance for responsible AI integration in RE education.

Figures

Figures reproduced from arXiv: 2607.28176 by Hansika Ekanayake Mudiyanselage, Malik Abdul Sami, Rohan Jai Dharmaraj, Zheying Zhang.

Figure 1
Figure 1. Figure 1: Screenshots of features implemented in the multi-agent AI tool [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Assignment workflow for AI tool integration in Assignment 4 [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Distribution of time spent using the AI tool during [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Distribution of the number of refinements made by [PITH_FULL_IMAGE:figures/full_fig_p007_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Distribution of student Likert-scale responses for [PITH_FULL_IMAGE:figures/full_fig_p007_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Distribution of student Likert-scale responses for [PITH_FULL_IMAGE:figures/full_fig_p008_6.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

22 extracted references

  1. [1]

    Generative ai for re- quirements engineering: A systematic literature review,

    H. Cheng, J. H. Husen, Y . Lu, T. Racharak, N. Yoshioka, N. Ubayashi, and H. Washizaki, “Generative ai for re- quirements engineering: A systematic literature review,” Software: Practice and Experience, 2025

  2. [2]

    Llm-based agents for automating the enhancement of user story quality: An early report,

    Z. Zhang, M. Rayhan, T. Herda, M. Goisauf, and P. Abrahamsson, “Llm-based agents for automating the enhancement of user story quality: An early report,” in International conference on agile software development. Springer, 2024, pp. 117–126

  3. [3]

    A multi-agent llm system for au- tomated requirements analysis: a study on user story generation and prioritization,

    M. A. Sami, Z. Zhang, M. Waseem, K.-K. Kemell, Z. Rasheed, T. Herda, M. T. Hasan, J. Rasku, and P. Abrahamsson, “A multi-agent llm system for au- tomated requirements analysis: a study on user story generation and prioritization,” inEuromicro Conference on Software Engineering and Advanced Applications. Springer, 2025, pp. 178–187

  4. [4]

    Towards implementing and evaluating ai-assisted pull requests in software en- gineering education,

    E. Parra and S. Willingham, “Towards implementing and evaluating ai-assisted pull requests in software en- gineering education,” in2025 IEEE/ACM 37th Interna- tional Conference on Software Engineering Education and Training (CSEE&T). IEEE, 2025, pp. 13–18

  5. [5]

    Bridging humans and llms: Investigating human-ai collaboration in multi-agent requirements analysis for organizational ai adoption,

    M. A. Sami, Z. Zhang, M. Waseem, K.-K. Kemell, Z. Rasheed, T. Herda, and P. Abrahamsson, “Bridging humans and llms: Investigating human-ai collaboration in multi-agent requirements analysis for organizational ai adoption,”e-Informatica Software Engineering Journal, vol. 20, no. 1, p. 260103, 2026

  6. [6]

    Students’ perceptions of the use of LLMs in requirements engineering education: A cross-university empirical study,

    S. Guardado, R. Parveen, Z. Zhang, M. Rayhan, and N. Tripathi, “Students’ perceptions of the use of LLMs in requirements engineering education: A cross-university empirical study,” inProceedings of the 33rd IEEE In- ternational Requirements Engineering Conference (RE). IEEE, 2025, pp. 130–141

  7. [7]

    Navigating the AI frontier: A critical literature review on integrating artificial intelligence into software engi- neering education,

    C. K. Sah, L. Xiaoli, M. M. Islam, and M. K. Islam, “Navigating the AI frontier: A critical literature review on integrating artificial intelligence into software engi- neering education,” inProceedings of the 36th Interna- tional Conference on Software Engineering Education and Training (CSEE&T). IEEE, 2024, pp. 1–5

  8. [8]

    Large language models (llms) in engineering education: A systematic review and sugges- tions for practical adoption,

    S. Filippi and B. Motyl, “Large language models (llms) in engineering education: A systematic review and sugges- tions for practical adoption,”Information, vol. 15, no. 6, p. 345, 2024

  9. [9]

    What is technological ped- agogical content knowledge (tpack)?

    M. Koehler and P. Mishra, “What is technological ped- agogical content knowledge (tpack)?”Contemporary is- sues in technology and teacher education, vol. 9, no. 1, pp. 60–70, 2009

  10. [10]

    Wake, b: Invest in good stories, and smart tasks,

    “Wake, b: Invest in good stories, and smart tasks,” https://xp123.com/articles/ invest-in-good-stories-and-smart-tasks/

  11. [11]

    Iso/iec/ieee international standard - systems and soft- ware engineering – life cycle processes –requirements engineering,

    “Iso/iec/ieee international standard - systems and soft- ware engineering – life cycle processes –requirements engineering,”ISO/IEC/IEEE 29148:2011(E), pp. 1–94, 2011

  12. [12]

    Leveraging llms for re- quirements engineering education: How to approach?

    S. Tiwari and S. S. Rathore, “Leveraging llms for re- quirements engineering education: How to approach?” in 2025 IEEE 33rd International Requirements Engineering Conference (RE). IEEE, 2025, pp. 458–466

  13. [13]

    Requirements are all you need: From require- ments to code with llms,

    B. Wei, “Requirements are all you need: From require- ments to code with llms,” in2024 IEEE 32nd In- ternational Requirements Engineering Conference (RE). IEEE, 2024, pp. 416–422

  14. [14]

    Towards integrating emerging ai applications in se education,

    M. Vierhauser, I. Groher, T. Antensteiner, and C. Sauer- wein, “Towards integrating emerging ai applications in se education,” in2024 36th International Confer- ence on Software Engineering Education and Training (CSEE&T). IEEE, 2024, pp. 1–5

  15. [15]

    Revo- lutionizing engineering education: The impact of ai tools on student learning,

    S. M. Vidalis, R. Subramanian, and F. T. Najafi, “Revo- lutionizing engineering education: The impact of ai tools on student learning,” in2024 ASEE Annual Conference & Exposition, 2024

  16. [16]

    Software engineering education must adapt and evolve for an llm environment,

    V . D. Kirova, C. S. Ku, J. R. Laracy, and T. J. Marlowe, “Software engineering education must adapt and evolve for an llm environment,” inProceedings of the 55th ACM Technical Symposium on Computer Science Education V . 1, 2024, pp. 666–672

  17. [17]

    Work in progress: Ai-powered engineering-bridging theory and practice,

    O. Levy, I. Dikman, N. Levy, and M. Winokur, “Work in progress: Ai-powered engineering-bridging theory and practice,” in2025 IEEE Engineering Education World Conference (EDUNINE). IEEE, 2025, pp. 1–4

  18. [18]

    Modeling teachers’ acceptance of generative artificial intelligence use in higher education: The role of ai literacy, intelligent tpack, and perceived trust,

    A. M. Al-Abdullatif, “Modeling teachers’ acceptance of generative artificial intelligence use in higher education: The role of ai literacy, intelligent tpack, and perceived trust,”Education Sciences, vol. 14, no. 11, p. 1209, 2024

  19. [19]

    Developing and validating an ai-tpack assess- ment framework: Enhancing teacher educators’ profes- sional practice through authentic artifacts,

    L. Eyal, “Developing and validating an ai-tpack assess- ment framework: Enhancing teacher educators’ profes- sional practice through authentic artifacts,”Education Sciences, vol. 15, no. 11, p. 1452, 2025

  20. [20]

    Using thematic analysis in psychology,

    V . Braun and V . Clarke, “Using thematic analysis in psychology,”Qualitative Research in Psychology, vol. 3, no. 2, pp. 77–101, 2006

  21. [21]

    J. W. Creswell and V . L. Plano Clark,Designing and Conducting Mixed Methods Research, 3rd ed. SAGE, 2018

  22. [22]

    Wohlin, P

    C. Wohlin, P. Runeson, M. H ¨ost, M. C. Ohlsson, B. Reg- nell, A. Wessl ´enet al.,Experimentation in software engineering. Springer, 2012, vol. 236