{"id":"0e99b650-c235-4652-b15d-e9609954a56b","arxiv_id":"2411.16723","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A two-agent ChatGPT-4 coder-reviewer setup beat both a single agent and a three-agent planner setup on code-error rate and abstract human-robot interaction tasks, with no clear overall trend from adding agents.","lead":"Researchers tested ChatGPT-4 agents that write code to control a robot dog from natural language commands, comparing one agent, a coder-plus-reviewer pair, and a planner-coder-reviewer trio. The pair produced the fewest coding errors and handled vague requests best, suggesting that collaboration helps but more agents is not automatically better.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Rerun-and-rescore protocol in §3.3 conditions observer ratings on successful code regeneration, so the claimed B advantage on abstract tasks may be an artifact of differential regeneration counts; the error-rate gap is also only 2–3 trials.","rationale":"The paper is a sincere, clearly described empirical comparison, and its negative result—that more agents do not monotonically improve performance—is plausible. The strongest positive claim is that configuration B improves error-free code generation and abstract problem solving. The error-rate evidence is weak on its own: with 21 initial executions per configuration, the reported 9–14 percentage point gaps correspond to only 2–3 trials, and no significance test is reported. The observer-rating evidence is weakened by a more design-level issue: Section 3.3 discards any trial whose first code attempt errors and regenerates the code before the observer scores it. This makes the observer scores conditional on successful regeneration, with the number of regeneration attempts unlogged and uncontrolled. Because configuration C has the highest initial error rate, its scored trials are not comparable to B's in terms of how much correction occurred before scoring. Thus the 'markedly higher performance' on abstract tasks could be an artifact of the rerun protocol rather than of the multi-agent architecture. This is a load-bearing concern for the central claim, but it is addressable with a cleaner protocol or a reanalysis of existing logs, so the verdict should remain CONDITIONAL as the reader recommended.","tokens_in":7187,"tokens_out":5451,"duration_ms":55516,"concrete_test":"Re-run Trials 3 and 5 with a no-rerun protocol: each configuration receives exactly one code generation and one execution per trial; any execution error is scored as failure (e.g., task success = 1) rather than discarded, and observers are blind to configuration with order randomized. Compare B versus A and C on these single-attempt scores. If B's advantage persists, the rerun protocol is not the source; if it disappears, the abstract-task claim fails. As a cheaper intermediate check, mine the existing logs for regeneration counts per configuration and trial and stratify Figure 9 by attempt number; if B is not higher among first-successful runs alone, the conclusion is not robust.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3.3 states that if generated code errors, the participant records nothing and the code is regenerated until a score can be recorded. Consequently every observer rating in Figures 8–10 is conditional on the code having successfully run at least once, and the number of regeneration attempts is neither logged nor controlled. Configuration C has the highest initial error rate (Figure 6), so its rated trials are drawn from a later, more heavily corrected part of its output distribution than configuration B's trials. The paper's central positive claim—that B 'shows markedly higher performance' on abstract/vague prompts (Trials 3 and 5, Figure 9)—therefore conflates architecture quality with the amount of free debugging each configuration received before being scored. If successful reruns improve scores, C's and A's scores are inflated relative to B's, and if observers notice the repeated failures, the blindness of the ratings is compromised. The same protocol also prevents the observer scores from being a valid measure of the systems' first-attempt capability, which is the quantity the abstract claims. The error-rate result itself is based on 21 initial executions per configuration; a 9–14 percentage point gap corresponds to only 2–3 trials, so without confidence intervals or a test the 'clearly least error prone' conclusion is not yet established. These are fixable design limitations, not evidence of fraud, but they are load-bearing for the headline.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper empirically compares three LLM-based agent architectures for controlling a quadrupedal robot through natural-language commands: configuration A (a single coder agent), configuration B (coder plus reviewer), and configuration C (planner, coder, and reviewer). Each configuration was tested on seven human-robot interaction prompts, with three repetitions per configuration, during which observer ratings of task success, safety, sociability, and expectation matching were collected, along with code-execution error rates and usage metrics (time and tokens). The reported findings are that there is no overall trend with the number of agents, but configuration B is the least error-prone and shows markedly higher observer-rated performance on the more abstract/vague prompts (Trials 3 and 5), while configuration C performs worst on initial error rate and never best on observer scores. A small case study with a single technical reviewer on anonymized code samples is also presented.","tokens_in":7456,"tokens_out":3423,"duration_ms":32426,"significance":"If the empirical claims were fully supported, the paper would make a useful contribution to the emerging literature on multi-agent LLM systems for embodied human-robot interaction, suggesting that adding a reviewer agent improves reliability and abstract problem-solving while adding a planner agent does not. The work has tangible strengths: the experimental setup is concrete and reproducible through the open 'spottyai' library, the prompts and trial categories are clearly tabulated, the code example in Figure 4 is informative, and the authors acknowledge several limitations in the Discussion. The core comparison is, however, small in scale (21 initial executions per configuration for error rates, three observer ratings per bar) and is affected by a protocol issue in Section 3.3 that conditions the observer scores on successful code regeneration. These concerns are load-bearing for the headline claims, so the contribution cannot yet be regarded as established.","major_comments":[{"comment":"The rerun-and-rescore protocol is a load-bearing threat to the central claim. Section 3.3 states that if the generated code errors, the participant records nothing and the code is regenerated until a score can be recorded. Since configuration C has the highest initial error rate (Figure 6), its observer-rated trials in Figures 8–10 are drawn from a later, more heavily corrected part of its output distribution than configuration B's trials. The conclusion that configuration B shows 'markedly higher performance' on the abstract Tasks 3 and 5 (Figure 9) therefore conflates architecture quality with the number of free regeneration attempts each configuration received before scoring. The number of regenerations is not logged. I recommend reporting the regeneration counts per configuration and either scoring first-attempt outputs for all configurations or including the regeneration count as a covariate in the analysis.","section":"§3.3, Figures 6 and 9"},{"comment":"The error-rate comparison is based on only 21 initial executions per configuration (three repetitions across seven trials). A 9–14 percentage point difference corresponds to just 2–3 error trials per configuration. The paper provides no confidence intervals, exact counts, or statistical test, yet states that configuration B is 'clearly least error prone' (Section 4, first paragraph). Please report the underlying contingency table and, at minimum, a two-sided Fisher exact test p-value or an exact binomial confidence interval for each proportion so that the reader can judge the strength of this conclusion.","section":"§4, Figure 6"},{"comment":"The observer-rated comparisons rest on three participants per configuration/trial, all self-reported as having low familiarity with AI and robotics. The claims of 'markedly higher performance' on Trials 3 and 5 and 'best performance' on Trial 2 are descriptive statements over n=3 per bar, with no reported inter-rater reliability, no error bars, and no inferential statistic. The paper should either show the individual participant scores, provide a measure of rating dispersion, or explicitly label these as illustrative observations rather than as evidence of a systematic performance advantage.","section":"§3.2, Figures 8–10"}],"minor_comments":[{"comment":"The statement that 'all code was pre-generated for all of the configurations and trials' is difficult to reconcile with the subsequent sentence explaining that errored code is regenerated; please clarify whether pre-generation involved multiple attempts per trial and whether observers were present during regeneration.","section":"§3.3"},{"comment":"The paper does not specify the exact ChatGPT-4 model version, sampling temperature, or other generation parameters; adding these would aid reproducibility.","section":"§3.1"},{"comment":"The protocol describes the observer feedback as blind, but it does not state whether observers knew that multiple AI configurations were being compared or whether they could infer errors from the instruction to record nothing; this should be made explicit.","section":"§3.3"},{"comment":"The case study uses two attempts per configuration across two prompts, reviewed by one technically proficient person; the quoted qualitative conclusions should be presented as illustrative, since they are based on a very small sample.","section":"§3.4"},{"comment":"The statement that multi-agent chat architecture 'seems to almost exponentially increase input token usage' is not supported by the data in Figure 7(c), which shows only three architecture points; consider softening this claim.","section":"§5"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a small empirical study with a clear, honest prose style, but the Section 3.3 regeneration protocol is the central technical problem: it makes the observer-rated results conditional on successful reruns, and the number of reruns is not recorded. This is fixable by re-analysis and more cautious wording, so major revision rather than rejection is appropriate. The error-rate evidence is also underpowered, and the authors should either add inferential statistics or soften the 'clearly least error prone' language. The paper would benefit from being framed as a preliminary pilot study; the venue should decide whether that scope fits its standards."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a small, honest empirical study of LLM agent architectures for a physical HRI task, and its most useful result is a caution: more agents is not monotonically better. The two-agent configuration (coder plus reviewer) had the lowest first-attempt error rate and scored higher on vague and abstract prompts, but the evidence for that strong claim is thinner than the prose suggests.\n\nWhat is actually new: prior multi-agent work targeted code generation, mechanics problems, and supply-chain optimization. Testing AutoGen-style collaboration on a real Spot robot with blind observers is a reasonable next step, and the non-monotonic result—the three-agent configuration with a planner was worse, not better—is a useful empirical counterweight to the enthusiasm for stacking agents. The paper is clearly written, the three configurations are described well enough to reproduce the setup, and the authors explicitly acknowledge the limited data and the possibility that perception or locomotion failures influenced scores.\n\nWhere it is soft: the headline error-rate difference is built on 21 initial executions per configuration, and a 9–14 percentage point gap is two or three trials. There are no confidence intervals or significance tests, so 'clearly least error prone' outruns the data. More load-bearing, Section 3.3 says that when code errors, the observer records nothing and the code is regenerated until a score can be recorded. That makes every observer rating conditional on a successful rerun, and the number of regeneration attempts is neither logged nor controlled. The claimed B advantage on the abstract tasks (Trials 3 and 5) could therefore be an artifact of differential regeneration counts—C, with the highest initial error rate, may have been scored on a later, more corrected output distribution than B. Blindness is also compromised if observers watch repeated failures. This is fixable by logging retries and scoring first-attempt behavior, but as it stands the central positive claim is not established. The token-saturation explanation in the discussion is post hoc speculation, though the authors flag it as such.\n\nBottom line: this is a reasonable descriptive pilot, not a confirmed result. It deserves peer review because the question is genuine and the field needs more head-to-head empirical comparisons of agent architectures, but it needs heavy revision: raw data, exact prompts, regeneration counts, and proper uncertainty quantification. I would bring it to a reading group as a methodological cautionary tale.","headline":"A useful but underpowered pilot: two-agent LLM architecture beats one and three on a physical HRI task, but the rerun-and-rescore protocol makes the main positive claim shaky.","tokens_in":7947,"tokens_out":3012,"would_cite":false,"duration_ms":26743,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A coder-plus-reviewer pair of LLM agents beats both a single agent and a three-agent team at controlling a robot.","keywords":["large language models","multi-agent systems","human-robot interaction","code generation","robot task planning","error rate","abstract reasoning","collaborative AI"],"falsifier":"Rerun the seven trials with a protocol that records observer scores for the first code attempt, including failures, with a larger panel of observers; if configuration B no longer outperforms A and C on the abstract tasks, the central claim fails. Alternatively, measure configuration C's error rate after trimming its system prompt; if errors persist, context saturation is not the primary cause.","tokens_in":6973,"feed_emoji":"🤖","tokens_out":8494,"duration_ms":63441,"temperature":0.7,"pith_summary":"This paper tests whether teams of large-language-model agents controlling a robot outperform a single agent. Across seven natural-language tasks with a quadrupedal robot, it finds no general trend: more agents do not mean better performance. The two-agent configuration, a coder paired with a reviewer, had the fewest initial code-execution errors, 9 and 14 percentage points lower than the single agent and the three-agent planner-coder-reviewer configuration. The same configuration scored markedly higher on vague and abstract tasks, while the three-agent configuration never outperformed the others. The takeaway is that a reviewer loop can add reliability, but an extra planning stage can hurt.","feed_headline":"Two-agent LLM team beats single and triple agents on robot tasks","feed_subtitle":"Adding a reviewer agent cuts initial code errors by 9–14 points and lifts performance on vague tasks.","key_machinery":"The central mechanism is a multi-agent group chat with role-separated prompts: a coder agent writes Python code, a reviewer agent critiques it for correctness, task completion, and safety, and a chat manager decides when the code is ready to execute. Configuration C inserts a planner agent that first converts the prompt into natural-language instructions for the coder. The reviewer loop is what carries the claimed benefit, while the planner stage is associated with degraded performance, possibly through longer context and constrained code generation.","core_discovery":"The paper's central claim is that in human-robot interaction, a collaborative two-agent system consisting of a coder and a reviewer is more reliable and better at abstract tasks than either a single agent or a three-agent planner-coder-reviewer system, and that there is no monotonic benefit from adding agents. The evidence is empirical: on the first code-execution attempt, configuration B failed 9 percentage points less often than configuration A and 14 percentage points less often than configuration C, and on the two most abstract prompts observers rated B markedly higher on success, safety, and sociability. The authors argue that the three-agent system's poor showing may stem from context-window saturation due to bloated token use, or from the planner's natural-language scaffold constraining the coder's flexibility.","pith_inferences":["If the context-saturation explanation is right, shortening system prompts or using retrieval-augmented generation should reduce configuration C's error rate; this is a directly testable extension.","The value of a reviewer agent may generalize beyond human-robot interaction to any embodied code-generating LLM setting, because the reviewer catches mismatches between generated code and the environment.","The observer protocol that discards failed first attempts before scoring may have compressed the measured performance gaps; scoring first attempts as-is would likely show a larger advantage for the two-agent configuration.","The results suggest an inverted-U relationship between the number of agents and task performance in embodied settings, which could be probed with four or more agents."],"forward_implications":["A two-agent coder-reviewer architecture can serve as a reliability layer for LLM-driven robot control without requiring an explicit planning stage.","Agent count is not a proxy for capability: adding a planner before the coder can make performance worse, not better.","Token usage and context length should be treated as first-order design variables when composing agent teams, since they may explain why more agents fail.","For vague or abstract commands, a reviewer agent may provide the largest gains, while for simple concrete commands a single agent may be enough.","The error-rate advantage of configuration B suggests that the reviewer loop catches coding mistakes that a lone agent lets through."],"supporting_citations":[{"why":"This reference supplies the multi-agent collaboration methodology and the premise that collaborative LLM teams outperform single agents on mechanics problems, which this paper tests in the human-robot interaction domain.","marker":"[Ni and Buehler, 2024]"},{"why":"This reference provides the group-chat framework used to combine the agents and the evidence that a dedicated checker agent improves rule-following, which motivates the reviewer role.","marker":"[Wu et al., 2023]"},{"why":"This reference establishes the code-generation approach for robot control that all three configurations use, and whose flexibility the planner stage may diminish.","marker":"[Liang et al., 2023]"},{"why":"This reference documents the lost-in-the-middle effect in long contexts, which the paper invokes to explain configuration C's higher error rate.","marker":"[Liu et al., 2024]"},{"why":"This reference suggests retrieval-augmented generation as a way to shorten system prompts and avoid context saturation, a proposed fix for configuration C's weakness.","marker":"[Zan et al., 2022]"}],"fun_headline_variants":["Two-agent LLM team beats solo and triple agents on robot tasks","Reviewer agent cuts robot code errors by 9–14 points","Coder+reviewer LLM duo excels at abstract robot tasks","More LLM agents isn't better: two-agent team wins in HRI","Collaborative LLM pair outshines single and triple agent setups"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The comparison assumes that discarding trials whose first code attempt errors does not bias observer scores; if observers infer the configuration from the rerun, or if reruns inflate ratings, the claimed advantage of the two-agent configuration on abstract tasks is unreliable.","fun_headline_variants_meta":{"raw":{"variants":["Two-agent LLM team beats solo and triple agents on robot tasks","Reviewer agent cuts robot code errors by 9–14 points","Coder+reviewer LLM duo excels at abstract robot tasks","More LLM agents isn't better: two-agent team wins in HRI","Collaborative LLM pair outshines single and triple agent setups"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000288,"raw_usage":{"total_tokens":1656,"prompt_tokens":882,"completion_tokens":774,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":498,"completion_tokens_details":{"reasoning_tokens":681}},"tokens_in":498,"tokens_out":774,"duration_ms":6855,"temperature":1.0,"reasoning_tokens":681,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T14:19:04.135689+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Rerun the seven trials with a protocol that records observer scores for the first code attempt, including failures, with a larger panel of observers; if configuration B no longer outperforms A and C on the abstract tasks, the central claim fails. Alternatively, measure configuration C's error rate after trimming its system prompt; if errors persist, context saturation is not the primary cause.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"This reference supplies the multi-agent collaboration methodology and the premise that collaborative LLM teams outperform single agents on mechanics problems, which this paper tests in the human-robot interaction domain."}],"review_version":1}