{"id":"d336f200-9275-46f1-b9fe-c084a50171ad","arxiv_id":"2505.04251","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A work-in-progress RACI-based role assignment framework for splitting tasks between humans and LLM agents in software development, aligned with Trustworthy AI principles.","lead":"This paper proposes a RACI-based framework to assign roles between humans and LLM agents in multi-agent software engineering. It aims to make human-agent collaboration more trustworthy and accountable, but it is not yet tested.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'artefact-based' precondition is load-bearing but unsatisfied in the worked example: sprint planning and product roadmap have no ground-truth oracle, so the accountability guarantee is unanchored.","rationale":"The reader identified Step-1's artefact-based precondition as the weakest assumption, and I agree that it is the load-bearing point. However, the concern is sharper than a mere limitation of scope: the paper's own worked example includes tasks that are not verifiable against ground truth, so the precondition is either vacuous or violated internally. This does not change the overall verdict of CONDITIONAL, because the framework is an explicitly work-in-progress proposal and the paper itself acknowledges the absence of empirical validation. The fix is feasible: clarify the definition of 'artefact-based', adjust the example to exclude non-verifiable tasks, or soften the central claim to state the scope explicitly. The planned groupware walkthrough should also test whether human validation of LLM-generated sprint plans is meaningful rather than nominal. None of this requires rejection, but it does require conditioning acceptance on the stated clarification and validation.","tokens_in":8085,"tokens_out":5205,"duration_ms":56795,"concrete_test":"Re-run §3.2 with an explicit definition: a task is artefact-based iff every candidate output can be automatically checked against ground truth (per §1). For each of the six rows in Table 1, specify the ground-truth oracle. If no oracle can be supplied for 'create product roadmap' or 'sprint planning and task allocation', then the worked example violates Step-1; the author must either revise Step-4 or explicitly restrict the framework to tasks with machine-checkable outputs before the accountability claim can stand.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that the framework 'ensures accountability' depends on Step-1 (§3.1), which requires that all tasks be artefact-based, and on §1's assertion that artefact-based outputs can be evaluated against ground truth. The paper never defines 'artefact-based' precisely, and the worked example appears to violate the precondition. In Step-4/Table 1, 'sprint planning and task allocation' and 'create product roadmap' are presented as automatable or partially automatable tasks, yet a sprint plan, task allocation, or product roadmap is a collaborative commitment with no objective ground truth to check against. If 'artefact-based' merely means 'produces a document', then the Step-1 filter is vacuous and the accountability guarantee—which relies on verifiable outputs—does not follow. If it means 'automatically verifiable against ground truth', then the example does not satisfy its own Step-1. Either way, the framework's core mechanism for accountability (a human Accountable actor validates the LLM output, Step-8) is not anchored for the tasks in its own example, leaving the central trust claim resting on an unexamined assumption.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This work-in-progress paper proposes a RACI-based framework (Responsible, Accountable, Consulted, Informed) for assigning roles between humans and LLM-based agents in software engineering (SE) workflows. The framework is presented as nine implementation steps plus a set of constraints, and a worked example builds a RACI matrix for the planning phase of a DevOps SDLC, involving three human actors and three LLM agents. The central claim is that this structured role assignment 'can facilitate efficient collaboration, ensure accountability, and mitigate potential risks associated with LLM-driven automation while aligning with the Trustworthy AI guidelines.' The paper explicitly notes that no empirical validation has been performed and outlines a planned groupware-walkthrough multi-case study as future work.","tokens_in":8321,"tokens_out":2932,"duration_ms":28552,"significance":"If the framework's central claim survives scrutiny, it would give practitioners a concrete, low-cost tool for injecting human oversight and accountability into LLM-driven SE automation, a pressing problem given the EU AI Act's oversight requirements. Strengths of the paper include the clearly enumerated nine-step process, the explicit set of framework constraints in §3.1.1, and a worked example whose RACI matrix is internally consistent with the accompanying narrative in most rows. The authors also deserve credit for openly stating that the framework has not been empirically validated and for identifying a concrete evaluation method. However, the paper's core accountability mechanism rests on the 'artefact-based' precondition, which is not precisely defined and appears violated by the very tasks in the example implementation; this undermines the strength of the central claim as currently stated.","major_comments":[{"comment":"The framework's Step-1 requires that all tasks be 'artefact-based,' and §1 argues that artefact-based evaluation against ground truth can be automated. Yet the example includes 'sprint planning and task allocation' and 'create product roadmap' as tasks where LLM-agents are Responsible or Consulted (Table 1, rows 5-6 and row 2). A sprint plan, task allocation, or product roadmap is a collaborative commitment with no objective ground-truth oracle, so the Step-8 human validation step has no verifiable target. This creates a dilemma: if 'artefact-based' means 'produces a document,' the Step-1 filter is vacuous and the accountability guarantee does not follow; if it means 'automatically verifiable against ground truth,' the example violates its own Step-1. The paper needs to define 'artefact-based' precisely and either restrict the framework to tasks with automatic verifiability or demonstrate how human validation alone anchors accountability without a ground-truth oracle.","section":"§3.1 Step-1 and §3.2 Example Implementation (Table 1)"},{"comment":"The framework's second constraint states that there must be at least one 'Accountable' assignment for each task, with an exception only when LLM-agents are assigned 'Informed' or no assignment. In Table 1, the 'Create product roadmap' row assigns LLM-agent A and LLM-agent C as 'Consulted,' yet assigns no 'Accountable' actor at all. Step-6 (point 2) adds a further inconsistency: it says 'There is no accountable assignment in this scenario as per the second framework constraint. The responsible actor is the accountable actor as well.' If the responsible actor is also the accountable actor, an A assignment should appear; if no A appears, the constraint is violated because the exception does not apply when LLM-agents are Consulted. This internal inconsistency directly affects the framework's internal validity and must be resolved in the example and in the statement of the exception.","section":"§3.1.1 Constraint (2) and §3.2 Step-6 / Table 1, 'Create product roadmap' row"},{"comment":"The paper claims that the AIA's obligations are satisfied whenever a human actor holds the Accountable role for a task in which an LLM-agent is Responsible (§3.2 Step-7), and this is used to support the overall claim of alignment with Trustworthy AI guidelines. The EU AI Act, however, imposes specific requirements (risk classification, transparency, human oversight, etc.) that are not mapped to framework elements anywhere in the paper. The assertion that a single human 'Accountable' assignment resolves all compliance obligations is presented without supporting analysis or reference to specific AIA provisions. This is a load-bearing step for the 'trustworthy' part of the central claim; the authors should either provide a concrete mapping between framework constructs and AIA obligations or temper the claim to indicate that the framework provides a structure that could support compliance rather than ensure it.","section":"§3.2 Step-7 and §4 Conclusion"},{"comment":"The paper states as a matter of fact that the framework 'ensures accountability' and 'mitigates potential risks associated with LLM-driven automation,' while also admitting in the same section that it 'has not yet undergone empirical validation.' For a work-in-progress paper, the abstract and conclusion should use modal or provisional language (e.g., 'is intended to facilitate,' 'may help ensure') so that the claims are not stronger than the evidentiary basis. This is not a fatal flaw, but it is important for readers to accurately calibrate the state of the research.","section":"§4 Conclusion"}],"minor_comments":[{"comment":"The abstract contains a grammatical error: 'challenges that spans' should be 'challenges that span.' Please proofread throughout.","section":"Abstract"},{"comment":"In Step-7, the phrase 'are in do not conflict' should be 'do not conflict.' In Step-9, 'If yes, return to Step-8' uses 'yes' where 'so' or 'conflicts exist' would be clearer. Similar typos appear in Step-7's reference to 'stepStep-5.'","section":"§3.1 Step-7 and Step-9"},{"comment":"The sentence 'Sprint planning and task allocation, which is are typically performed...' should be 'which is typically performed' or 'which are typically performed.'","section":"§3.2 Step-4 (5)"},{"comment":"The row 'Task allocation' is identical to 'Sprint planning' in the matrix but has no separate explanation in Step-6. Please clarify whether these are intended as distinct tasks or as a single combined task; if combined, the table should have one row.","section":"Table 1"},{"comment":"Reference [10] lists 'Tse-Husn' as an author; this appears to be a garbled name. The authors should verify the correct spelling, likely 'Tse-Hsun Chen.'","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper has a clear structure and a useful practical idea, but the central 'artefact-based' precondition is not well defined and the worked example appears to violate it, which weakens the core accountability claim. The inconsistency in the 'Create product roadmap' row (constraint 2 vs. Table 1) is a concrete internal problem that the authors must fix. The compliance claims (Step-7) also need more substance. I recommend major revision rather than rejection because these issues are addressable within the paper's scope: the authors can refine the definition, adjust the example, and soften the claims. The paper is explicitly a work in progress, so the lack of empirical validation alone is not grounds for rejection, but the claims should be presented as proposals rather than established results."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a solid, clearly-scoped work-in-progress paper. The RACI framework is a reasonable and internally consistent way to assign roles between humans and LLM agents in software engineering, and the worked DevOps example is concrete. But the paper overclaims in the abstract and conclusion: “ensuring accountability” and “mitigating potential risks” are not supported by empirical evidence, and the paper itself admits this.\n\nWhat’s actually new: applying the standard RACI matrix from project management to human-LLM agent collaboration, with nine implementation steps and a specific planning-phase matrix that includes a hypothetical scrum-assistant agent. The constraints in §3.1.1 are sensible, and the example respects them. The paper also grounds the agent roles in published agent PO/RE designs, which is honest.\n\nSoft spots: the artefact-based precondition in Step-1 is the most fragile piece. The stress-test note worries that sprint planning and product roadmap have no ground-truth oracle. I think that’s a partial concern: the framework’s accountability mechanism is human validation (the A role), not automated verification, so the absence of an oracle alone isn’t fatal. But “artefact-based” is never defined, and if it just means “produces a document”, it is nearly vacuous. The paper should either define it precisely or drop it from the guidelines. Also, the EU AI Act alignment is thin – it boils down to ensuring a human is Accountable, which is fine but not deeply analyzed. And the novelty is incremental; it’s an application of an existing management tool.\n\nWho it’s for: practitioners designing LLM-agent SDLC processes, and researchers in human-agent collaboration. A serious referee should engage, but the request should be for a revision that tempers the claims and sharpens the artefact-based definition.\n\nRecommendation: accept as a companion paper with minor revisions.","headline":"A modest but useful work-in-progress: the RACI framework is a sensible, internally consistent role-assignment tool for human-LLM collaboration, but the trust and accountability claims outrun the evidence.","tokens_in":8812,"tokens_out":2515,"would_cite":false,"duration_ms":25033,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A RACI-based framework can keep humans accountable when LLM agents do software work","keywords":["Large Language Models","LLM-Based Multi-Agent Systems","Software Engineering","Trustworthy AI","Human-Agent Collaboration","DevOps","RACI matrix","Accountability"],"falsifier":"A multi-case study with expert walkthroughs that applies the nine steps to SE tasks whose outputs cannot be automatically verified would show whether Step-1's artefact-based precondition holds. If teams cannot assign a human Accountable actor to such tasks without a verifiable output, the framework's central accountability guarantee would be limited to a subset of software work.","tokens_in":7863,"feed_emoji":"🤖","tokens_out":4729,"duration_ms":40265,"temperature":0.7,"pith_summary":"This work-in-progress paper proposes a RACI-based framework (Responsible, Accountable, Consulted, Informed) for assigning roles between humans and LLM-based multi-agent systems in software engineering. The framework provides a nine-step implementation guideline and a worked example for a DevOps planning phase. The paper argues that this structured role assignment can facilitate efficient collaboration, ensure accountability, and mitigate risks from LLM-driven automation, aligning with the EU's Trustworthy AI guidelines. A sympathetic reading takes the proposal as a concrete answer to the open problem of how to allocate tasks between humans and LMA systems.","feed_headline":"RACI matrix keeps humans accountable in LLM software teams","feed_subtitle":"Nine-step role framework lets LLM agents act while humans stay accountable for the outcome.","key_machinery":"The central object is the RACI responsibility-assignment matrix, a standard project-management tool here adapted to LMA-oriented software engineering. Responsibilities are defined as: Responsible does the work, Accountable delegates and validates the outcome, Consulted provides input, and Informed is kept in the loop. The machinery is the constraint set that forces a human Accountable for every LLM-Responsible task, plus the nine-step implementation guideline that connects tasks, actors, regulatory constraints, and workflow design. The matrix enforces human oversight at the point of output validation.","core_discovery":"The paper's central claim is that applying the RACI matrix to human-agent collaboration solves the task-allocation problem in LLM-based multi-agent software engineering. In this adapted RACI, humans hold the Accountable role whenever an LLM-agent is Responsible for producing an artefact, so every automated output passes through a human validation gate before the task is complete. The framework's three constraints enforce this: each task needs at least one Responsible and one Accountable actor, and no LLM-agent can be Responsible without at least one human being Accountable. The example DevOps planning matrix demonstrates how this works for tasks such as creating user stories, building the product backlog, and sprint planning.","pith_inferences":["If the framework proves out, it suggests a general principle for human-AI collaboration: accountability should be structurally assigned to humans at the validation point, not at the point of task execution, and this could extend to other artefact-producing AI domains such as data analysis or document drafting.","The artefact-based precondition means the framework may not cover early-SDLC activities like stakeholder negotiation or product vision alignment; extending it to such tasks would require specifying new validation mechanisms for non-verifiable outputs.","A testable extension would be to compare teams using RACI assignments against teams without them on metrics like the share of LLM-generated artefacts merged without human sign-off, or the number of incidents traced to unvalidated outputs.","The paper's planned groupware walkthrough could be sharpened to specifically probe Step-1, asking experts which SDLC tasks resist ground-truth verification and whether the framework still guarantees accountability for those."],"forward_implications":["Teams can use the nine-step method to produce a RACI matrix for any artefact-based task, phase, or the whole SDLC, giving a concrete division of labour.","Human validation becomes a mandatory step before any LLM-agent-produced artefact is accepted, which operationalises accountability in practice.","The framework offers a systematic way to demonstrate compliance with EU AI Act deployer obligations, since every LLM output has a named human accountable actor.","The example DevOps planning matrix can be reused as a template for other SDLC phases such as development, testing, and deployment.","The distinction between Responsible and Accountable clarifies when LLM agents may act autonomously versus when humans must stay in the loop."],"supporting_citations":[{"why":"Supplies the systematic review of LMA systems in SE and the identification of trustworthiness as a research gap that motivates the framework.","marker":"[5]"},{"why":"Provides the premise that SE collaboration is artefact-based and enables automated evaluation, which underwrites Step-1 of the framework.","marker":"[3]"},{"why":"Provides the original RACI definitions that the paper adapts to LMA-oriented SE.","marker":"[7]"},{"why":"Supplies the seven Trustworthy AI requirements that the framework claims to align with.","marker":"[14]"},{"why":"Supports the claim that SE collaboration is artefact-based, part of the rationale for Step-1.","marker":"[23]"},{"why":"Exemplifies an LMA system in the SDLC that motivates the role-allocation problem.","marker":"[6]"}],"fun_headline_variants":["RACI framework puts humans in control of LLM agents","Human accountability via RACI in LLM software teams","RACI model gives humans final say over LLM agents","LLM agents act, humans answer: RACI framework"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The framework assumes every task it governs is artefact-based, so its outputs can be compared against ground truth automatically; for tasks like stakeholder negotiation or product vision alignment, the paper offers no mechanism to keep humans accountable for LLM-agent output.","fun_headline_variants_meta":{"raw":{"variants":["RACI framework puts humans in control of LLM agents","Human accountability via RACI in LLM software teams","RACI model gives humans final say over LLM agents","LLM agents act, humans answer: RACI framework"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000184,"raw_usage":{"total_tokens":1263,"prompt_tokens":836,"completion_tokens":427,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":452,"completion_tokens_details":{"reasoning_tokens":361}},"tokens_in":452,"tokens_out":427,"duration_ms":4359,"temperature":1.0,"reasoning_tokens":361,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T23:32:43.336890+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A multi-case study with expert walkthroughs that applies the nine steps to SE tasks whose outputs cannot be automatically verified would show whether Step-1's artefact-based precondition holds. If teams cannot assign a human Accountable actor to such tasks without a verifiable output, the framework's central accountability guarantee would be limited to a subset of software work.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the original RACI definitions that the paper adapts to LMA-oriented SE."},{"cited_title":"2019.Ethics Guidelines for Trustworthy Artificial Intelligence (AI)","cited_arxiv_id":null,"evidence_quote":"Supplies the seven Trustworthy AI requirements that the framework claims to align with."}],"review_version":1}