{"id":"8486e42e-96c4-4fd0-aaa1-2caed2b09ce3","arxiv_id":"2411.14009","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Human experts and GPT agents emphasized different ethical concerns for multi-robot systems, with GPT outputs mirroring existing AI ethics guidelines rather than new context-specific issues.","lead":"This paper ran ethics brainstorming workshops with human experts and with teams of ChatGPT agents, then compared what each group worried about. It found human experts focused on corporate bad behavior and data privacy, while the AI agents mostly repeated standard AI ethics talking points.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The human-versus-GPT comparison is confounded: GPT Workshop 3 used a generic prompt unrelated to the domestic multi-robot scenario given to human experts, so observed theme differences may reflect task mismatch rather than reasoner type.","rationale":"The reader's CONDITIONAL verdict rests on the same mismatch I see as load-bearing, so I agree with the weakest_assumption. I want to add that the mismatch is even more specific than 'task specificity': the GPT prompt quoted in Section 3.3.3 does not mention multi-robot systems at all, making the referent of the GPT deliberation different from the human deliberation. This makes the abstract's comparison of 'concern profiles' vulnerable to the objection that the two groups answered different questions. A matched re-run is feasible and would settle the issue. Since the authors themselves propose future replications but do not report a matched comparison, the appropriate disposition remains conditional rather than full accept; I do not see grounds to reject outright because the qualitative observations about GPT behavior (politeness, solution-orientation) and the MORUL model development are independent contributions. Thus verdict_should_be is UNCHANGED with respect to the reader's CONDITIONAL.","tokens_in":26096,"tokens_out":5165,"duration_ms":49053,"concrete_test":"Run the GPT-agent workshop again exactly as in Section 3.3.3, but replace the generic project description with the Workshop 1 scenario text: two robot vacuum cleaners of the same model, brand and manufacturer, owned and used for some time, with the home owners then purchasing a new robot arm from a different company. Keep everything else fixed (gpt-3.5-turbo-16k, three agents per team, one judge, equal number of rounds for both teams, e.g., 15 each). Apply the same thematic analysis and compare the theme distributions. If GPT output now includes corporate dominance, maleficence, and deviance themes at rates similar to the human workshops, the original difference is a scenario artifact; if GPT output still clusters on standard AI ethics principles, the original claim survives.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract's comparative claim presupposes that the human and GPT deliberation tasks are equivalent except for who (or what) is reasoning. They are not. Human Workshops 1 and 2 (Sections 3.3.1 and 3.3.2) were built around a specific domestic scenario: two robot vacuum cleaners of the same brand and model, plus a newly purchased robot arm from a different company, with face-to-face or Jamboard mind-mapping and group discussion. Workshop 3 (Section 3.3.3) instead gave GPT-3.5 agents the generic project description 'build a team of large language model agents cooperating,' with no mention of robot vacuums, the home, or a brand mismatch; the two GPT teams also ran different numbers of rounds (5 vs 15). The human-emphasized themes—corporate dominance, maleficence, strategic withholding or falsifying information between brands—are plausibly direct consequences of the brand-mismatch scenario itself. A generic LLM-agent prompt would naturally evoke standard AI ethics principles (privacy, bias, transparency, accountability) that the paper attributes to GPT. Because the prompt, scenario, and procedure differ systematically between conditions, the observed divergence cannot be assigned to human-versus-GPT cognition. This is an internal-validity threat to the central claim, and it is not acknowledged in the Limitations section (5.1).","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a qualitative study comparing ethical concerns elicited from human expert workshops with those generated by GPT-based agent teams in the context of GAI-empowered multi-robot systems. Two human workshops (N=16 participant contributions) used a concrete domestic scenario involving two same-brand robot vacuums and a new robot arm from a different company, while a third workshop used two teams of GPT-3.5 agents prompted with the generic project description “build a team of large language model agents cooperating.” Thematic analysis identified 21 human themes and 103 GPT-agent ideas, and the authors report that humans emphasized deviance, data privacy, bias, and unethical corporate conduct, whereas GPT agents emphasized concerns already present in AI ethics guidelines. The paper also advances the MORUL framework for ethical development of multi-robot systems.","tokens_in":26356,"tokens_out":3995,"duration_ms":37166,"significance":"If the comparative claim were valid, the paper would provide early, much-needed empirical evidence on how LLM-based deliberation of ethical issues in multi-robot systems differs from human expert deliberation, and the MORUL framework is a useful structuring device for context-specific AI ethics. The study is unusually transparent: the workshop procedures are described in detail, and the GPT transcripts and judge output are linked in the paper. However, the central human-versus-GPT comparison is currently compromised by systematic differences in scenario, prompt, and procedure between the conditions, so the headline conclusion is not yet supported. The paper is best positioned as a dual case study—human expert elicitation and LLM agent deliberation—rather than a controlled comparison, unless the conditions are matched.","major_comments":[{"comment":"The central comparative claim in the Abstract—that “human experts placed greater emphasis on new themes related to deviance, data privacy, bias and unethical corporate conduct” while GPT agents “emphasized concerns present in existing AI ethics guidelines”—is confounded. Human Workshops 1 and 2 (Sections 3.3.1 and 3.3.2) used a concrete domestic scenario with two same-brand robot vacuums and a newly purchased robot arm from a different company, whereas Workshop 3 (Section 3.3.3) gave GPT-3.5 agents the generic project description “build a team of large language model agents cooperating,” with no mention of the domestic setting, the robots, or the brand mismatch. The human-emphasized themes—corporate dominance, maleficence, strategic withholding or falsifying information between brands—are plausible direct consequences of the brand-mismatch scenario itself. Because scenario, prompt content, and procedure differ systematically between conditions, the observed divergence cannot be attributed to human-versus-GPT cognition. The Limitations section (5.1) does not acknowledge this internal-validity threat.","section":"3.3.1 vs. 3.3.3"},{"comment":"There is an internal inconsistency about the number of GPT deliberation rounds. Section 3.3.3 states that Team 1 was set to 5 rounds and Team 2 to 15 rounds, while Section 4.2 reports that “two GPT agent teams” engaged in 15 rounds of deliberation. If Team 1 in fact used only 5 rounds, then the quantitative comparisons between teams—for example, Team 1 producing 6 problem-focused themes versus Team 2's 19 (Figure 9)—are confounded by the round asymmetry. Furthermore, the Judge's decision (Section 4.3) explicitly favored Team 2 for its “more specific details and suggestions,” which is at least partly a mechanical consequence of the greater number of rounds. The paper must resolve this inconsistency and either match the round counts or treat the round count as a separate experimental factor.","section":"3.3.3 vs. 4.2"},{"comment":"The conclusion that GPT agents “emphasized concerns present in existing AI ethics guidelines” rests on a classification of themes into “existing guideline concerns” versus “new concerns,” but the paper does not provide a systematic mapping of each GPT and human theme to the principle sets in Jobin et al. (2019) or to the risk taxonomy of Weidinger et al. (2022). Without an explicit coding scheme, the claim that GPT output largely reproduced existing guidelines while human output introduced new themes is not fully supported by the presented analysis. Listing central GPT themes such as “privacy and data,” “bias and discrimination,” and “accountability and transparency” is not the same as demonstrating that those themes were absent from the human workshops; the human theme list in Section 4.1 also includes data security and privacy, trustworthiness, and maleficence. The comparative coding needs to be made explicit and applied consistently to both datasets.","section":"4.2 / Figure 9"}],"minor_comments":[{"comment":"The phrase “two teams of 6 agents plus one judge” is ambiguous because each team actually had three agents; please rephrase to “six agents in two teams of three, plus one judge” to match Section 3.3.3.","section":"Abstract"},{"comment":"The citation “Kouba et al., 2023” on page 8 is inconsistent with the reference list entry “Koubaa, A. (2023)\"; please correct the in-text citation.","section":"2.4"},{"comment":"The thematic analysis section would benefit from a brief statement about how many researchers performed the coding, whether coding was independent, and how disagreements were resolved, given that the themes are the basis for all subsequent comparisons.","section":"3.4"},{"comment":"The bar charts and Venn diagrams are discussed only loosely in the text; adding explicit figure captions and referring to specific counts in the text would help readers verify the theme frequency claims.","section":"Figures 6, 7, 9, 10"}],"recommendation":"major_revision","confidential_remarks":"The paper's main contribution—the human-versus-GPT comparison—rests on a design that does not currently support the stated conclusion. In addition to the authors of this report, the reader's report and stress-test note identify the same core issue: the GPT workshop used a generic prompt while human workshops used a concrete domestic scenario, and the GPT teams had different round counts. In my view this is fixable within the manuscript's scope: the authors could either re-run the GPT workshop with the identical domestic scenario and matched round counts, or explicitly reframe the study as two complementary case studies (human expert elicitation and LLM agent deliberation) and soften the comparative language in the Abstract and Discussion. The MORUL framework contribution is independent of the comparison and can stand on its own. The paper's transparency about prompts, transcripts, and analysis steps is a strength worth preserving."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"What you should know first: this paper has one genuinely new idea—using competing teams of GPT agents plus a judge agent to elicit ethical concerns for GAI-enabled multi-robot systems—and that part is worth a look. But the headline comparison between human experts and GPT agents does not hold as stated, and the flaw is internal validity, not just sample size.\n\nThe Workshop 3 setup is a real contribution. Having three agents per team with a judge select the \"winner\" is a concrete way to observe how LLMs deliberate ethics. The report of polite, solution-oriented language is an observable behavior worth discussing. The human workshops also provide rich thematic data, and the MORUL model is a reasonable continuation of the authors' prior work.\n\nThe central problem is that the human and GPT conditions are not matched. Human workshops used a specific domestic scenario—two identical robot vacuums plus a new robot arm from a different brand (Sections 3.3.1–3.3.2). The GPT agents were given a generic project description, \"build a team of large language model agents cooperating\" (Section 3.3.3), with different round counts between teams. The themes the paper attributes to human distinctiveness—corporate dominance, maleficence, withholding or falsifying information between brands—are plausibly direct readings of the brand-mismatch scenario. A generic prompt would naturally evoke standard AI ethics principles such as privacy, bias, transparency, and accountability. The abstract's claim that humans emphasized deviance, data privacy, bias, and unethical corporate conduct while GPT agents echoed existing guidelines cannot be separated from the scenario difference. This is not acknowledged in the Limitations section (5.1), which is a miss.\n\nIs the comparison completely worthless? No. Because both sets of participants were asked about ethical concerns in GAI-enabled MRSs, the overlap in themes is still meaningful. But the divergence is not attributable to \"human versus GPT cognition\" as claimed; it is attributable to task framing. The authors could fix this by re-running Workshop 3 with the same domestic scenario, or by running human workshops on the generic prompt. Until then, treat the comparative conclusion as a hypothesis, not a result.\n\nWho gets value from this: people working on AI ethics elicitation methods, HRI researchers, and the MORUL program's followers. It deserves a serious referee, and a good referee should send it back for a matched re-run or at least for thorough revision of the claims and limitations. I would not cite the comparative result as it stands, but I would keep the GPT-team method in mind.","headline":"The GPT-team deliberation method is genuinely new, but the human-versus-GPT comparison is confounded by mismatched scenarios and should not be taken as evidence about human versus LLM cognition.","tokens_in":26910,"tokens_out":2491,"would_cite":false,"duration_ms":23909,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Human experts and GPT agents produce systematically different ethical concern profiles for generative-AI-enabled multi-robot systems, with humans adding deviance, data privacy, bias, and corporate misconduct.","keywords":["multi-robot systems","large language models","generative AI","AI ethics","ethics workshops","GPT agents","thematic analysis","human-robot interaction"],"falsifier":"Run the same protocol with matched conditions: give GPT agent teams the identical two-vacuum-plus-robot-arm household scenario that the human experts received, with equal round counts. If the GPT agents then produce deviance, corporate misconduct, and data-misuse themes at human-like rates, the paper's central difference between human and LLM ethical concern profiles disappears.","tokens_in":25924,"feed_emoji":"🤖","tokens_out":9874,"duration_ms":85281,"temperature":0.7,"pith_summary":"The paper aims to show that ethical concerns arising from generative AI in cooperating robot teams are not fully captured by the AI ethics principles that language models tend to recite. It compares two human expert workshops with a workshop in which teams of GPT agents deliberate the same broad question, using thematic analysis to identify themes. The reported result is a systematic gap: human experts put more weight on deviance, data privacy, bias, and unethical corporate conduct, while GPT agents emphasize topics already present in AI ethics guidelines such as transparency, accountability, bias, and privacy. If this gap holds, it matters because LLMs are increasingly proposed as auditors or conversational interfaces for multi-robot systems, and an audit that misses corporate misconduct and manipulation would be incomplete.","feed_headline":"GPT agents echo AI guidelines while human experts go further","feed_subtitle":"LLM teams stick to textbook AI ethics; human experts add deviance, privacy, bias, and corporate misconduct.","key_machinery":"The carrying mechanism is the three-workshop comparison analyzed through thematic analysis. The human side used two qualitative workshops built around a concrete domestic scenario, while the GPT side used two LLM agent teams prompted with the generic instruction 'build a team of large language model agents cooperating' and evaluated by a judge agent. The comparison unit is the theme: qualitative data from flip-charts, virtual sticky notes, and GPT transcripts were coded inductively into themes, counted by mentions, and positioned against technological layers and the safety, security, and societal dimensions of the emerging MORUL model for ethical development of multi-robot systems.","core_discovery":"The study's central discovery is a divergence in ethical perception between human experts and GPT agents. In the two human expert workshops, which used a domestic scenario involving two vacuum robots and a newly purchased robot arm from another brand, the themes with distinctive emphasis were communication failure and manipulation, corporate dominance, maleficence, data privacy, and bias. In the GPT-agent workshop, two teams of three gpt-3.5-turbo-16k agents discussed ethical concerns in multi-robot cooperation and produced themes that largely matched established AI ethics guidelines: privacy and data, bias and discrimination, accountability and transparency, freedom of expression and assembly, and an 'ethical considerations' theme that echoed the prompt. The GPT teams were consistently polite, solution-oriented, and prone to listing mitigations rather than dwelling on harms. The paper interprets this as evidence that LLM deliberation currently reproduces the standard ethical vocabulary, while human experts working from a concrete use scenario surface concerns about deliberate corporate and interpersonal misconduct that the guidelines and the LLMs do not foreground.","pith_inferences":["A matched re-run giving GPT agents the exact domestic scenario and equal round counts used with the human experts is the direct test of whether the observed gap reflects human-versus-LLM cognition or simply task concreteness and conversation length.","A practical consequence not drawn in the paper: if LLM-generated ethics audits for robot fleets systematically omit corporate misconduct and manipulation, then procurement and certification processes that rely on LLM outputs need human red-team review targeted at those themes.","The politeness finding could be tested adversarially by prompting GPT teams to actively hunt for malicious, deceptive, or profit-driven uses; that would show whether the absence of deviance themes is a training-data boundary or an instruction-boundary."],"forward_implications":["If LLM-generated ethical concern lists are used as a stand-in for stakeholder deliberation in multi-robot system development, they will likely reproduce standard guideline topics but under-represent corporate misconduct, deviance, and user-side privacy harms.","The observed politeness and solution-bias of LLM agents implies that conversational AI in robot teams can present ethically problematic behavior in benign language, so oversight should monitor behavior rather than relying on the surface wording.","The judge agent selected the longer, more detailed discussion, so evaluations of LLM ethics discussions are sensitive to procedural choices such as round count and discussion length.","The human themes organized by the MORUL layers provide a starting checklist for practitioners: communication, cooperation, human oversight, trustworthiness, maleficence, and legislation are categories worth probing when reviewing generative-AI-enabled multi-robot systems."],"supporting_citations":[{"why":"Provides the inventory of common AI ethics principles that the GPT agents' themes are compared against.","marker":"Jobin et al. (2019)"},{"why":"Supplies the taxonomy of LLM risks used to frame and organize the GPT-generated ethical concerns.","marker":"Weidinger et al. (2022)"},{"why":"Provides the thematic analysis method used to code the workshop and GPT discussion data.","marker":"Terry et al. (2017)"},{"why":"Earlier data collection on ethical concerns in robot-to-robot cooperation from which this study's workshop sequence extends.","marker":"Rousi et al. (2022)"}],"fun_headline_variants":["Humans flag deviance, GPT sticks to ethical script","AI ethics: GPT repeats guidelines, humans add misconduct","Multi-robot study: GPT echoes textbook ethics, humans diverge","GPT agents follow AI rules; humans spot corporate sins","Human experts outpace GPT on novel ethical concerns"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claim that humans and GPT agents differ systematically depends on the two workshops being comparable enough to attribute the difference to the reasoners, yet the human experts worked from a concrete brand-specific domestic scenario while the GPT agents received a generic prompt and different numbers of discussion rounds.","fun_headline_variants_meta":{"raw":{"variants":["Humans flag deviance, GPT sticks to ethical script","AI ethics: GPT repeats guidelines, humans add misconduct","Multi-robot study: GPT echoes textbook ethics, humans diverge","GPT agents follow AI rules; humans spot corporate sins","Human experts outpace GPT on novel ethical concerns"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000192,"raw_usage":{"total_tokens":1381,"prompt_tokens":1012,"completion_tokens":369,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":628,"completion_tokens_details":{"reasoning_tokens":290}},"tokens_in":628,"tokens_out":369,"duration_ms":4363,"temperature":1.0,"reasoning_tokens":290,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T15:37:20.424265+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same protocol with matched conditions: give GPT agent teams the identical two-vacuum-plus-robot-arm household scenario that the human experts received, with equal round counts. If the GPT agents then produce deviance, corporate misconduct, and data-misuse themes at human-like rates, the paper's central difference between human and LLM ethical concern profiles disappears.","supporting_citations":[],"review_version":1}