{"id":"b3a5834f-2918-456d-87b8-01c366bb5155","arxiv_id":"2604.11975","paper_version":2,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"An LLM multi-agent HRI framework with per-robot personality, long-term memory, and contextualized turn coordination improved identity clarity, preference awareness, and conversational flow in a 105-user study.","lead":"M2HRI gives each robot a personality and long-term memory, plus a coordinator that decides who speaks when. A 105-person study reports clearer identities, better preference recall, and smoother multi-robot conversation.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"Abstract-only review leaves the complementary-roles claim untestable; ecological transfer is the load-bearing premise but cannot be audited without methods.","rationale":"The Reader correctly flags that an abstract-only surface cannot support ACCEPT/CONDITIONAL/REJECT and that ecological validity of the controlled multi-agent scenario is the premise that would have to hold for the homes/hospitals framing. I agree that is the load-bearing external concern. Internally, the complementary-roles claim is also untestable without methods, baselines, and statistics; that is a structural limitation of the review surface rather than a demonstrated flaw in the work. No formal verification or code is available to offset this. Therefore the verdict remains UNVERDICTED with low confidence; no adjustment is warranted. The concrete test is the minimal audit that would either secure or undermine the claim once the full text is in hand.","tokens_in":2059,"tokens_out":507,"duration_ms":4522,"concrete_test":"Obtain the full paper (or methods/results sections) and verify: (1) presence of factorial or ablation conditions that isolate personality, memory, and coordination; (2) objective metrics or validated scales for distinguishability, preference awareness, naturalness, flow, appropriateness, and overlap; (3) reported effect sizes/CIs and multiple-comparison handling for n=105. If ablations are missing or effects are non-significant/confounded, the complementary-roles claim does not hold as stated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that personality/long-term memory and contextualized participation coordination play complementary roles in coherent multi-agent HRI, supported by an n=105 controlled study. Because only the abstract is available, the claim rests on uninspectable design choices: how personality was operationalized and measured for distinguishability, how long-term memory was stored/retrieved and scored for preference awareness/naturalness, how the coordination mechanism regulated turn-taking, what baselines or ablations isolated complementarity, and what the multi-agent scenario actually was. The reader's weakest assumption (lab scenario as proxy for homes/hospitals) is real and load-bearing for the abstract's target environments, but the more immediate soundness gap is that none of the reported gains can be checked for confounds, effect sizes, or statistical support. Without those, the complementary-roles conclusion cannot be distinguished from scenario-specific artifacts of the LLM stack.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The manuscript introduces M2HRI, a multimodal multi-agent HRI framework that treats each robot as an identity-bearing agent via personality and long-term memory, and adds a contextualized coordination mechanism to regulate multi-agent participation. In a controlled multi-agent HRI user study (n=105), the authors report that most personality contrasts were distinguishable and consistently expressed; long-term memory improved preference awareness and interaction naturalness; and contextualized coordination improved conversational flow, response appropriateness, and overlap avoidance. The central claim is that agent individuality and contextualized participation coordination play complementary roles in coherent, socially appropriate multi-agent HRI, with intended relevance to social settings such as homes and hospitals.","tokens_in":2238,"tokens_out":1171,"duration_ms":19183,"significance":"If the complementary-roles result holds under a fully specified design with appropriate controls, the work would be a useful systems and empirical contribution to multi-robot HRI: it moves beyond interchangeable functional agents toward identity-bearing multi-agent interaction and pairs that individuality with an explicit participation-coordination layer. An n=105 controlled study is a meaningful empirical asset for HRI if effect sizes, significance tests, baselines, and ablations are reported and sound. The practical framing toward homes and hospitals is relevant. Credit is due for jointly studying personality, long-term memory, and coordination rather than treating multi-robot social interaction as pure task allocation; a project website is also provided for further materials.","major_comments":[{"comment":"The complementary-roles claim is load-bearing for the paper’s contribution, but the abstract alone does not report baselines, ablations that isolate personality vs. long-term memory vs. coordination, effect sizes, or significance tests. Without those, the reported gains in distinguishability, preference awareness, naturalness, flow, appropriateness, and overlap avoidance cannot be distinguished from scenario- or stack-specific artifacts of the LLM multi-agent setup. A full methods/results section with factorial or leave-one-component-out comparisons is required before the central claim can be assessed.","section":"Abstract (user study, n=105)"},{"comment":"Personality distinguishability is a primary positive finding, yet the abstract does not specify how personality was operationalized (prompting, traits, multimodal expression), how “distinguishable” and “consistently expressed” were measured (forced-choice, Likert, behavioral coding), or against what control. These measurement choices are load-bearing for interpreting the individuality half of the complementary-roles claim and must be fully specified and justified.","section":"Abstract (personality contrasts)"},{"comment":"Long-term memory is credited with improved preference awareness and naturalness, but storage/retrieval design, what counts as a preference, and how awareness/naturalness were scored are not given. Without that operationalization and a no-memory or short-term-only control, the memory contribution cannot be audited as a separable factor complementary to coordination.","section":"Abstract (long-term memory findings)"},{"comment":"Contextualized coordination is credited with better flow, appropriateness, and overlap avoidance, but the abstract does not describe the participation policy (who speaks when, conflict resolution, multimodal cues) or the uncoordinated/baseline multi-agent condition. That mechanism is load-bearing for the coordination half of the claim and for the assertion that individuality changes multi-robot coordination requirements.","section":"Abstract (contextualized coordination)"},{"comment":"The introduction targets multi-robot social environments such as homes and hospitals, while the evidence is a controlled multi-agent HRI scenario. Ecological transfer is a load-bearing premise for the applied claim; the manuscript needs an explicit scenario description, limitations discussion, and either ecological-validity checks or a clearly scoped claim limited to the lab setting.","section":"Abstract (homes/hospitals framing vs. controlled study)"}],"minor_comments":[{"comment":"Only the abstract was available for this review; section numbering, figures, tables, and equations could not be checked. Once the full text is provided, presentation issues (figure clarity, notation for the coordination policy, completeness of related multi-agent HRI and LLM-agent citations) should be reviewed in a second pass.","section":"Manuscript availability"},{"comment":"The abstract asserts that “most” personality contrasts were distinguishable; when full results appear, report which contrasts failed and why, to avoid over-generalizing the individuality result.","section":"Abstract (personality results wording)"},{"comment":"Project website is linked; ensure the camera-ready version pins code, prompts, and study materials for reproducibility of the LLM multi-agent stack.","section":"Abstract (project website)"}],"recommendation":"uncertain","confidential_remarks":"This is an abstract-only review; full text was not available, so soundness cannot be verified and the recommendation is necessarily uncertain. The abstract’s n=105 study and complementary-roles framing look potentially interesting for cs.RO / HRI if methods hold, but I cannot assess confounds, statistics, or novelty relative to prior multi-robot and LLM-agent HRI work without the manuscript body. Please provide the full paper for a standard review cycle; I would re-evaluate with a concrete accept/revise/reject recommendation once methods, results tables, and ablations are inspectable."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"Punchline: this is an integrated multi-agent HRI systems paper arguing that per-robot personality plus long-term memory, together with contextualized participation control, improve coherence and social appropriateness. From the abstract alone that claim is clear and empirically framed; we cannot yet check whether the study actually isolates those pieces.\n\nWhat is new is the packaging, not a new foundational mechanism. Personality and memory in HRI, and multi-agent LLM dialogue stacks, already exist. M2HRI’s contribution is treating robots as identity-bearing agents and pairing that with a coordination layer that regulates who speaks when, then testing the combination in a controlled multi-agent scenario (n=105). That sample size is reasonable for HRI. The directional findings—personality contrasts mostly distinguishable, memory helping preference awareness and naturalness, coordination helping flow, appropriateness, and overlap avoidance—are the right kind of claims for this subfield, and the complementary-roles framing is useful for product design in multi-robot social settings.\n\nSoft spots, in proportion: we only have the abstract. No effect sizes, significance tests, baselines, ablations, scenario details, or operationalizations of personality/memory/coordination. So the complementary-roles conclusion cannot yet be separated from scenario-specific artifacts of the LLM stack. The homes/hospitals motivation is load-bearing; ecological transfer from a controlled lab scenario is the real generalization risk, and the abstract does not report deployment evidence. Circularity risk looks ordinary for HCI self-report plus system design, not a derivation tautology. No formal verification or shipped code/data visible from this surface.\n\nWho this is for: multi-robot HRI and social robotics people building LLM multi-agent stacks who care about identity and turn-taking. It is not a theory paper. It deserves a serious referee rather than a desk reject—the problem is real, the study size is not trivial, and the design question is worth answering carefully. I would send it to peer review and ask for methods, stats, baselines, and ecological-validity discussion. I would not cite it yet on abstract alone; I might bring it to reading group if someone is actively building multi-robot dialogue systems.","headline":"Abstract-only multi-robot HRI systems paper with a clear complementary-roles claim and n=105 study; design is sensible, evidence not yet auditable.","tokens_in":2879,"tokens_out":539,"would_cite":false,"duration_ms":10612,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Personality, long-term memory, and contextualized coordination make multi-robot teams feel like distinct social agents rather than interchangeable tools.","keywords":["multi-agent HRI","human-robot interaction","personality","long-term memory","contextualized coordination","multi-robot systems","LLM agents","social robotics"],"falsifier":"A deployment or ecological-validity study in a real home or hospital multi-robot setting in which personality contrasts become indistinct, memory fails to improve preference awareness/naturalness, or contextualized coordination fails to reduce overlap and improve flow relative to a non-individualized baseline.","tokens_in":2902,"feed_emoji":"🤖","tokens_out":768,"duration_ms":7010,"temperature":0.7,"pith_summary":"M2HRI is a multimodal multi-agent framework for human–robot interaction that treats each robot as an identity-bearing social agent instead of a replaceable function. Each agent carries a distinct personality and long-term memory, and a contextualized coordination layer decides who speaks when so the team does not talk over itself or produce mismatched replies. In a controlled user study with 105 participants, people could usually tell the personalities apart and found them consistently expressed; memory improved the system’s awareness of user preferences and made the interaction feel more natural; and the coordination mechanism improved conversational flow, response appropriateness, and overlap avoidance. The paper argues that individuality and participation control play complementary roles: one makes agents feel like distinct people, the other keeps multi-agent conversation coherent and socially appropriate. If the claim holds, multi-robot systems deployed in homes or hospitals could be designed around recognizable social identities rather than pure functional interchangeability.","feed_headline":"Robots with personality and memory talk more naturally as a team","feed_subtitle":"A 105-person study shows identity and coordination both improve multi-robot conversation.","key_machinery":"M2HRI itself: an LLM-driven multimodal multi-agent stack that instantiates each robot as an identity-bearing agent (personality model + long-term memory) and adds a contextualized coordination mechanism that regulates who participates when, so individuality does not destroy conversational coherence.","core_discovery":"Agent individuality (personality plus long-term memory) and contextualized participation coordination play complementary roles in coherent multi-agent HRI: most personality contrasts were distinguishable and consistently expressed, long-term memory improved preference awareness and naturalness, and contextualized coordination improved conversational flow, response appropriateness, and overlap avoidance in a 105-person user study.","pith_inferences":["The same identity-plus-coordination pattern may transfer to non-robot multi-agent assistants (home hubs, multi-avatar interfaces) if the social-perception mechanisms are modality-independent.","Scalability to more than a few agents will likely require richer coordination than turn-taking alone, because personality conflicts and memory load grow with team size.","Longitudinal home deployments could test whether long-term memory compounds preference accuracy over weeks rather than single sessions."],"forward_implications":["Multi-robot systems can be designed around stable, distinguishable agent identities rather than functional interchangeability.","Long-term memory is a practical lever for preference-aware, more natural multi-agent conversation.","Participation coordination is necessary when agents are individualized, otherwise conversational overlap and mismatched replies rise.","Personality, memory, and coordination should be co-designed rather than treated as separate add-ons."],"fun_headline_variants":["Personality and memory make robot teams chat more naturally","Study: Robot identities plus coordination improve multi-agent HRI","Distinct robot personalities and memory boost conversation flow","Individuality and coordination play complementary HRI roles","Long-term memory and agent identity improve multi-robot talks"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"That the lab multi-agent HRI scenario and the LLM personality/memory/coordination stack used in the study are faithful enough proxies for real multi-robot social settings (homes, hospitals) that the measured gains will transfer beyond that controlled setup.","fun_headline_variants_meta":{"raw":{"variants":["Personality and memory make robot teams chat more naturally","Study: Robot identities plus coordination improve multi-agent HRI","Distinct robot personalities and memory boost conversation flow","Individuality and coordination play complementary HRI roles","Long-term memory and agent identity improve multi-robot talks"]},"model":"grok-4.5","effort":"low","cost_usd":0.002652,"raw_usage":{"total_tokens":992,"prompt_tokens":729,"num_sources_used":0,"completion_tokens":59,"cost_in_usd_ticks":26520000,"prompt_tokens_details":{"text_tokens":729,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":204,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":729,"tokens_out":59,"duration_ms":2652,"temperature":1.0,"reasoning_tokens":204,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-12T21:38:54.958653+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"A deployment or ecological-validity study in a real home or hospital multi-robot setting in which personality contrasts become indistinct, memory fails to improve preference awareness/naturalness, or contextualized coordination fails to reduce overlap and improve flow relative to a non-individualized baseline.","supporting_citations":[],"review_version":2}