{"id":"28e90b03-fe52-4b8b-b604-454f06156132","arxiv_id":"2608.04148","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"AgentForge lets novices practice software repair by playing one of four roles alongside AI agents; in a 37-person study, completion and self-reported learning were high, while code review proved the hardest role.","lead":"AgentForge is a learning platform where novices take on one of four software-engineering roles while AI agents perform the other three in a code-repair workflow. A study of 37 novice developers found high task-completion rates and self-reported learning gains, but revealed that code review was the most demanding role and that many students accepted AI-generated drafts with little or no revision.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central learning claim is undercut by absent objective pre-test/control and by the paper's own log evidence that 85.4% of Coach-assisted human turns were exact copies, leaving the 'critically and effectively' component unsupported.","rationale":"The reader's formal weakest assumption is the fixed role–task mapping, which threatens cross-role comparisons but not the abstract's central claim that AgentForge improves novices' software-repair skills and critical AI collaboration. The more load-bearing concern is that the central learning claim itself rests on an unbaselined objective test and self-report, and the paper's own interaction logs show high verbatim acceptance of AI Coach drafts. The reader's rationale does mention the lack of a pre-test and control condition, so there is partial overlap, but the chosen weakest assumption is not the one that most directly undermines the headline contribution. The paper remains a useful systems contribution and the empirical findings are honestly reported, including the low test-evidence accuracy and widespread Coach copying. These issues do not warrant rejection, but they do require major revisions to the claim: the evidence supports a scaffolded environment in which novices complete AI-assisted repairs and report understanding gains, not demonstrated critical or effective agentic collaboration. Thus the conditional verdict is unchanged.","tokens_in":12436,"tokens_out":4642,"duration_ms":52193,"concrete_test":"Run a pre-registered controlled follow-up in which the 19-item knowledge test is administered both before and after AgentForge use, and in which a control group performs the same four BugsInPy repair tasks with the same underlying GPT-5 agents but without AgentForge's role briefings, handoffs, or AI Coach scaffolding. If the pre-test mean already approximates the reported 82.6% post-test mean, or if the control group shows equivalent pre–post gains and equivalent post-test scores, the central learning claim is unsupported; if AgentForge yields significantly larger objective gains and a measurably higher rate of substantive (non-verbatim) edits to AI drafts, the claim survives.","verdict_should_be":"UNCHANGED","load_bearing_attack":"To support the abstract's claim that AgentForge helps novices 'develop practical software-engineering skills while learning to collaborate with agentic AI more critically and effectively,' the study must show both learning gains beyond self-report and critical engagement. Neither is established. The 19-item objective knowledge test was administered only in the post-survey (Data Analysis, Performance and Learning Analysis), so the reported 82.6% mean has no pre-test baseline; the only pre–post outcome is a 9-item self-report scale, which is vulnerable to demand characteristics. The 'critically' component is directly weakened by the paper's own log analysis: of 130 human turns submitted after a Coach request, 85.4% exactly reproduced the Coach draft and 88.5% were at least 95% text-identical (Results, Perceptions). Additionally, Coach-use strategy was not significantly associated with response quality, workflow-understanding gains, post-test knowledge, task completion, or patch similarity. The weakest objective area was interpreting test evidence, with only 20.0% correct (Results, Performance and Learning Analysis), which is precisely the critical-validation skill the central claim implies is learned. Thus the evidence supports high task-completion under scaffolding, but not measured growth in critical, effective agent collaboration.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents AgentForge, a role-play platform in which novice developers take on one of four software-engineering roles (Task Planner, Patch Author, Code Reviewer, Test Runner) inside a multi-agent code-repair workflow, while AI agents perform the other roles. In a within-subject study with 37 novice developers, participants completed role-based practice sessions using BugsInPy repair tasks. The paper reports high task-completion rates with AI-agent support, significantly different interaction demands across practices (with Code Reviewer showing the heaviest load), self-reported pre-post gains in understanding of software repair and agent collaboration, and a high mean score on an objective post-test knowledge instrument. The authors claim that AgentForge can help novices develop practical software-engineering skills while learning to collaborate with agentic AI more critically and effectively.","tokens_in":12607,"tokens_out":5545,"duration_ms":59326,"significance":"If fully supported, AgentForge would be a valuable contribution to computing education and human-AI collaboration: it provides a concrete, implemented platform for making multi-agent workflow structure visible and for scaffolding role-specific responsibilities. The paper has notable strengths: a real system implementation, the use of BugsInPy repair tasks, a within-subject design with four roles, rich interaction-log analysis, and explicit reporting of Coach-use behavior, effect sizes, and adjusted p-values. However, the central learning claim rests on self-report gains without a control condition or an objective pretest, and the paper's own log data show high rates of verbatim acceptance of AI-generated Coach drafts. As an exploratory system description with descriptive findings about interaction patterns, the work has merit, but the abstract's causal and 'critical collaboration' claims are not supported by the currently reported evidence.","major_comments":[{"comment":"Retaining only the last attempt per participant-role pair after including retries can bias completion rates and interaction-load estimates. If an incomplete first attempt is followed by a successful retry, only the successful attempt is counted, while an incomplete final attempt is counted as incomplete, but the mixture is not reported separately. The paper should report the number of retries per role, present results with and without retries, or provide a sensitivity analysis, because this decision directly affects the completion-rate and interaction-burden claims in Tables 1 and 2.","section":"Experiment, final dataset paragraph; Tables 1 and 2"},{"comment":"The four roles are assigned to different BugsInPy tasks of different difficulty, with Task Planner and Patch Author paired with 'easy' tasks and Code Reviewer and Test Runner paired with 'medium' tasks. All cross-role comparisons, such as the higher interaction turns, reroutes, completion time, and perceived challenge for Code Reviewer in Table 2, are therefore confounded: these differences could reflect task difficulty or task-specific properties rather than the role itself. The paper should either use multiple tasks per role, include task difficulty as a factor in the analysis, or explicitly limit the claims to the specific role-task combinations studied.","section":"Task Design, fixed role-task mapping; Table 2"},{"comment":"The 19-item objective knowledge test was administered only in the post-survey, so the reported 82.6% mean has no pretest baseline and cannot establish learning gains. The only pre-post outcome is the nine-item self-report scale, which is vulnerable to demand characteristics. Without an objective pretest or a control condition, the evidence does not support the abstract's claim that AgentForge helps novices 'develop practical software-engineering skills.' The manuscript should either add such a baseline/control or revise the central claim to refer to reported or perceived gains.","section":"Data Analysis, Performance and Learning Analysis; Results, Performance and Learning Analysis"},{"comment":"The paper's own log analysis directly undercuts the 'critically and effectively' component of the central claim: of the 130 human turns submitted after a Coach request, 85.4% exactly reproduced the Coach draft and 88.5% were at least 95% text-identical, and Coach-use strategy was not significantly associated with response quality, workflow-understanding gains, post-test knowledge, task completion, or patch similarity. Additionally, the weakest objective-knowledge area was interpreting test evidence, with only 20.0% correct, which is exactly the critical-validation skill the claim implies is learned. The authors should soften the 'critically and effectively' claim or provide additional evidence of critical engagement, such as revision depth, justification quality, or performance on validation items.","section":"Results, Perceptions; Discussion and Conclusion"}],"minor_comments":[{"comment":"The response-quality rubric is described as an 'exploratory deterministic rubric' with 89.1% agreement between the automated and human assessments, but the paper does not report the inter-rater statistic or how disagreements were resolved. If rubric scores are reported descriptively and used in inferential tests, this information should be included.","section":"Data Analysis, Response-quality rubric"},{"comment":"The sentence beginning 'This may be because of the 130 human turns submitted following a Coach request...' is grammatically incomplete and should be rephrased, for example as 'This may be because, of the 130 human turns submitted following a Coach request, 85.4% exactly reproduced the Coach draft...'.","section":"Results, Perceptions"},{"comment":"For reproducibility, the paper should report the specific BugsInPy task identifiers (issue IDs or bug IDs) used for the four practices, rather than only describing them by topic and difficulty.","section":"Methodology, System Design and Task Design"},{"comment":"The abstract's final sentence overstates the evidence by moving from 'reported significant gains' to 'These findings suggest that AgentForge can help novices develop practical software-engineering skills while learning to collaborate... critically and effectively.' The claims should be aligned with the evidence, which supports descriptive findings about interaction and perceived value rather than measured critical collaboration.","section":"Abstract and Introduction"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a reasonable exploratory system study, but the abstract and conclusion overclaim the learning evidence. A revised version that reframes the claims as descriptive/exploratory, adds the requested sensitivity analyses, and acknowledges the role-task and measurement confounds would be suitable for publication. I do not see evidence of undisclosed duplication or citation problems, but the paper should be checked for consistency with prior work on role-based learning and AI scaffolds."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear Colleague,\n\nTwo things to know about AgentForge. The role-immersion design—learner plays one role while AI agents play the others—is a genuine new application of role-based learning to agentic AI education. The empirical finding that Code Reviewer produced the heaviest interaction and the weakest test interpretation (20% correct) is credible and useful. But the central learning claim in the abstract is not actually backed by the evidence: there is no control condition, no objective pretest, and the logs show that 85.4% of Coach-assisted human turns exactly reproduced the coach draft. So 'critically and effectively' is an aspiration, not a measured outcome.\n\nWhat the paper does well. The system is thoughtfully designed: structured handoffs, role briefings, visible intermediate artifacts, and an AI Coach that generates an editable prompt—while keeping the learner responsible for the final submission. The study uses real BugsInPy tasks, logs are rich, and the statistical analysis is careful with Holm corrections. The authors also report the copy-paste statistic and openly interpret it as cognitive offloading, which is honest. Their suggestion that scaffolding should fade is a reasonable design implication.\n\nWhere it is soft. First, the cross-role comparison is confounded by the fixed role–task mapping: Code Reviewer and Test Runner were paired with medium tasks, the other two with easy tasks. The paper never mentions this as a limitation, so claims like 'Code Reviewer practice required more turns' are really claims about the task-role package. Second, retaining only the last attempt per participant–role pair biases completion rates upward; a sensitivity analysis with all attempts would be easy to run. Third, learning is measured only by self-report (the nine pre–post items) and a post-only 19-item test. The absence of a pretest means the 82.6% average tells you nothing about gains. Combined with the Coach-copying result, the evidence supports 'high completion under scaffolding' but not 'developed critical collaboration skills.'\n\nIs it worth refereeing? Yes. The system and the exploratory results are valuable for computing-education and HCI audiences, and the paper is transparent enough that a referee can see exactly where the claims exceed the data. It needs major revision before acceptance: add a control or baseline, address the role–task confound explicitly, report all attempts, and temper the abstract's causal language. I'd send it to review and expect a constructive outcome.\n\nBest.","headline":"Novel role-immersion platform for teaching agentic AI, with honest data, but the abstract overclaims: no control, no pretest, and heavy coach copy-paste.","tokens_in":13204,"tokens_out":2226,"would_cite":true,"duration_ms":23205,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"AgentForge is an immersive role-playing platform in which novices practise software repair by taking on one of four engineering roles while AI agents handle the others; a 37-participant study reports high completion rates and significant…","keywords":["agentic AI","role-based learning","software engineering education","multi-agent workflow","human-AI collaboration","code review","metacognitive scaffolding","novice developers"],"falsifier":"Rotate the roles across the same set of tasks, or pair each role with tasks of matched difficulty, and re-measure interaction turns, reroutes, completion time, and perceived challenge; the role-based explanation survives only if the Code Reviewer practice remains the hardest regardless of which bug it is attached to.","tokens_in":12157,"feed_emoji":"🤖","tokens_out":5664,"duration_ms":50903,"temperature":0.7,"pith_summary":"AgentForge is a learning platform that places novice developers inside a multi-agent code-repair workflow: the learner takes on one of four software-engineering roles — Task Planner, Patch Author, Code Reviewer, or Test Runner — while AI agents carry out the other three. The paper reports a study with 37 novices in which task-completion rates were high (100% for Planner and Author, 93.5% for Test Runner, 80.6% for Reviewer), and participants reported significant gains in understanding software repair and agent collaboration ($p_{\\mathrm{adj}} < .001$). At the same time, the Code Reviewer practice required significantly more interaction turns, reroutes, and completion time, and was perceived as the most challenging. The paper argues that making agent handoffs, intermediate artifacts, and role-specific decisions visible helps novices learn to guide and evaluate AI agents rather than passively accept their outputs.","feed_headline":"AgentForge helps novices steer AI repair agents; review stays hard","feed_subtitle":"Learners who take one of four roles while AI does the rest report stronger grasp of multi-agent repair.","key_machinery":"The load-bearing mechanism is a four-role workflow — Task Planner, Patch Author, Code Reviewer, Test Runner — in which the human learner occupies one role and AI agents occupy the rest, with execution pausing at the learner's role for a structured response. Role handoffs make upstream artifacts, responsibilities, and expected response formats explicit; role briefings provide procedures and success criteria; an optional AI Coach generates an editable draft prompt. These scaffolds are what carry the argument: they externalize the work of planning, monitoring, and evaluating that the paper claims transfers to understanding of software repair and agent collaboration.","core_discovery":"On the paper's own terms, the central discovery is that role-based immersion can give novices a workable footing in agentic software engineering: with scaffolding and metacognitive support, novice developers complete realistic repair workflows at high rates and report improved understanding of how planning, patching, review, and testing fit together, as well as how to collaborate with multiple AI agents. The same data show that critical evaluation is the bottleneck: Code Reviewer was the most interaction-heavy and most challenging practice, and post-test responses were weakest on deciding whether to approve a patch and on interpreting test evidence. The paper further finds that the AI Coach was used often, but that many 'human' turns were near verbatim copies of the Coach draft, indicating that the system succeeds at workflow understanding more than at forcing critical engagement.","pith_inferences":["The fixed role-task mapping is an untested confound: the paper's cross-role comparisons (Reviewer hardest, Author most variable) could be explained by task difficulty or bug type rather than role, since Planner/Author tasks are classified easy and Reviewer/Test Runner tasks medium.","The 85.4% exact-reproduction rate of Coach drafts suggests that the platform's current design may train novices to copy AI suggestions into the workflow rather than to reason independently; a transfer test on an unseen patch would tell whether critical evaluation actually improved.","A natural extension is adaptive scaffold fading: as novices gain experience, reduce hint detail and Coach availability, and measure whether revision rates and review accuracy rise; the paper itself suggests this direction.","The same role-immersion pattern could be applied to other software activities, such as requirements analysis or design review, but the present evidence is restricted to code repair on a single benchmark repository."],"forward_implications":["With AgentForge-style scaffolding, novices can achieve high completion rates on real benchmark repair tasks while AI agents do most of the work, so the bottleneck shifts from code generation to judgment.","Code review and test interpretation are the practices that expose novice weakness, suggesting curricula should invest in reviewing and validating AI output rather than only in prompting.","Because most AI Coach uses produced near-verbatim drafts, systems that want critical engagement must either fade scaffolding or design outputs that are intentionally flawed.","Self-reported understanding of software repair and agent collaboration improved significantly, indicating that visible handoffs and role briefings can teach workflow-level concepts even without deep code changes."],"supporting_citations":[{"why":"ChatDev provides the role-decomposed multi-agent structure that AgentForge adapts for education.","marker":"(Qian et al. 2024)"},{"why":"MetaGPT supplies the standardized-operation-procedure approach that informs the role handoff and briefing design.","marker":"(Hong et al. 2024)"},{"why":"SWE-agent is the prior single-agent automated repair system the paper contrasts with its human-in-the-loop learning setup.","marker":"(Yang et al. 2024)"},{"why":"AutoCodeRover is another autonomous repair baseline that motivates the need for human oversight and validation.","marker":"(Zhang et al. 2024)"},{"why":"Provides the BugsInPy-based tasks and the student-contribution context used to select accessible repair tasks.","marker":"(Fang et al. 2023)"},{"why":"Supplies the LLM-assisted rubric method used to score response quality.","marker":"(Cohn et al. 2025)"},{"why":"Documents novice reliance on generative AI, the problem AgentForge's scaffolding targets.","marker":"(Prather et al. 2023)"}],"fun_headline_variants":["AgentForge role-play lifts novice repair skills; review stays tough","Novices master AI repair roles, but Code Reviewer proves hardest","AgentForge: high completion in AI repair, but review needs more turns","Role-based training with AI agents: review is the weak link"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The strongest assumption is that differences across the four practices are caused by the role itself, because each role is fixed to a particular BugsInPy task (Planner and Author on easy tasks, Reviewer and Test Runner on medium ones) and the paper does not control for task difficulty or content when comparing roles.","fun_headline_variants_meta":{"raw":{"variants":["AgentForge role-play lifts novice repair skills; review stays tough","Novices master AI repair roles, but Code Reviewer proves hardest","AgentForge: high completion in AI repair, but review needs more turns","Role-based training with AI agents: review is the weak link"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000208,"raw_usage":{"total_tokens":1428,"prompt_tokens":993,"completion_tokens":435,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":609,"completion_tokens_details":{"reasoning_tokens":360}},"tokens_in":609,"tokens_out":435,"duration_ms":5718,"temperature":1.0,"reasoning_tokens":360,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T00:20:55.619852+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Rotate the roles across the same set of tasks, or pair each role with tasks of matched difficulty, and re-measure interaction turns, reroutes, completion time, and perceived challenge; the role-based explanation survives only if the Code Reviewer practice remains the hardest regardless of which bug it is attached to.","supporting_citations":[],"review_version":1}