{"id":"28a5309f-9824-4ad6-9b70-58343331bb21","arxiv_id":"2412.17954","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"In a new partially observable cooking game, training waiters with human or autonomous chef partners did not improve their later performance with a new human chef, though the study demonstrates a behavior-clustering experimental design for evaluating asynchronous team training.","lead":"This paper introduces and tests a method for training human teammates asynchronously, using an autonomous chef agent instead of a human teammate during practice for a partially observable cooperative cooking game. The experiment found no significant benefit of any training condition on later performance with a new human partner, but it offers a reusable experimental design and clustering approach for future human-AI team training studies.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Load-bearing concern: the in-sample/out-of-sample evaluation contrast assumes behavior clusters learned from full-observability self-play persist when chefs play with waiters under partial observability; the paper neither enforces nor validates this, so the cluster-based design's central mechanism…","rationale":"The reader's weakest-assumption analysis identifies the same load-bearing concern: cluster assignments from self-play with full observability may not persist when the same human chef plays with a waiter under partial observability. I agree that this is the central vulnerability of the paper's methodological claim. The paper provides modest internal support for the clustering itself (e.g., high silhouette score, per-subject cluster concentration in Table 1), but it offers no evidence about temporal or contextual stability of those clusters. The authors themselves flag the distribution shift and the lack of compliance enforcement, but they do not treat this as a validity threat to the in-sample/out-of-sample contrast — which is exactly where the threat is most dangerous. The concern is empirical and addressable: if the proposed test shows cluster assignments are stable in evaluation sessions, the design's central mechanism survives; if not, the condition-reduction argument loses its foundation. Given that the paper is otherwise honest about its null training result and openly lists limitations, a conditional accept remains appropriate. No critical red flag invalidates the central claim as stated; the missing validation should be a required revision rather than grounds for rejection.","tokens_in":17526,"tokens_out":3198,"duration_ms":31441,"concrete_test":"Using the logged chef-role trajectories from the evaluation sessions (which the protocol must have recorded, since both players' actions were captured), re-run the same sequential encoder and K-medoids pipeline (or a classifier trained on self-play cluster labels) on each human chef's evaluation-session trajectories. Compute the assignment of each evaluation session to the three clusters and compare it with the chef's training-time cluster assignment. If, say, more than 30% of evaluation sessions are assigned to a different cluster than the chef's training-time cluster, or if the concordance is not much better than chance, then the in-sample/out-of-sample manipulation is not valid and the claimed condition-reduction benefit is unsupported. This test directly settles whether the cluster labels used to form training conditions describe behavior during evaluation.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that behavior clustering makes the study tractable by letting clusters serve as conditions (Section 4.1). For that claim to hold, the cluster assigned to each established chef during self-play must reliably describe that chef's behavior during evaluation sessions, when the chef plays with a new waiter under partial observability. The paper explicitly states 'we did not enforce established chefs' compliance with their previous behavior' (Section 5) and acknowledges in Section 4.1 that self-play has 'complete information' and 'may have downstream effects (e.g., distribution shift).' Nothing in the paper verifies that the unsupervised clusters learned from full-observability self-play trajectories (sequential encoder + K-medoids, Section 4.1) transfer to the partially observable, mixed-role evaluation setting. If a chef assigned to Cluster 1 exhibits Cluster 2 behavior when paired with an unfamiliar waiter, then the in-sample versus out-of-sample evaluation contrast collapses: waiters trained with a 'similar' cluster chef may actually be evaluated with a chef whose behavior differs, and vice versa. The design's mechanism for reducing conditions would then be invalid, because the labels on which the condition reduction is based do not describe the evaluated behavior. This is not an internal inconsistency; it is an unvalidated empirical assumption that is load-bearing for the methodological contribution.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents a cooperative-asynchronous-training evaluation paradigm applied to a new two-role partial-observability game (HYBS) derived from Overcooked-AI. Established chefs play self-play games; their decision sequences are embedded with a sequential encoder and clustered via K-medoids; waiters are trained with a human, an apprentice agent, a heuristic agent, or no training, and are evaluated with human chefs from the assigned cluster (in-sample) and from a different cluster (out-of-sample). A 52-participant between-subjects study found significant training-tip differences across conditions but no significant differences in evaluation tip, which the authors report as a null training effect. The paper's main contribution is the clustered experimental design that reduces the number of conditions needed for the evaluation.","tokens_in":17754,"tokens_out":4796,"duration_ms":46658,"significance":"The proposed design is potentially valuable for reducing scheduling and participant burden in team-training studies, and the paper is careful and honest in its statistical reporting (normality checks, CIs, Tukey corrections, explicit acknowledgement of the null result). The game environment and clustering methodology are interesting, and the paper ships a real human-subjects demonstration rather than only a simulation. However, the design claim depends on an unvalidated assumption about cluster persistence across settings, and the current manuscript contains at least one internal statistical inconsistency. As a methodological contribution, the paper is promising but needs revision.","major_comments":[{"comment":"The central claim that behavior clustering enables a tractable experimental design rests on the assumption that cluster assignments learned from full-observability self-play continue to describe chefs' behavior during the partially observable evaluation with waiters. The paper explicitly states \"we did not enforce established chefs' compliance with their previous behavior\" (Section 5) and acknowledges \"distribution shift\" from full-observability training data (Section 4.1), yet no validation is provided that a chef assigned to a given cluster in self-play exhibits the corresponding cluster behavior during evaluation. If a chef's behavior changes when paired with an unfamiliar waiter, the in-sample/out-of-sample contrast collapses and the condition-reduction mechanism is invalid. Please add an empirical check of cluster persistence, such as re-encoding evaluation trajectories and comparing cluster assignments, or a behavioral consistency metric collected during evaluation.","section":"Section 4.1 and Section 5"},{"comment":"The NULL condition was conducted as a post-hoc follow-on study rather than as part of the original randomized between-subjects design. This makes the comparison between NULL and the other conditions vulnerable to temporal confounds, including changes in the recruitment pool, experimenter experience, or any fixes to game scenarios over time. The Q3 null result (no training benefit) is therefore not a clean causal finding. The paper should either justify that these temporal factors are unlikely to affect the comparison, or explicitly restrict the causal interpretation of Q3.","section":"Section 6.2 and footnote 4"},{"comment":"The reported evaluation-tip statistic \"F(3, 29) = 2.184; p > 0.100\" in Section 6.2 matches the Training Quality row of Supplement Table 5, not the evaluation-tip rows. The supplement reports Evaluation Tip (Seen): F(3, 29) = 0.698, p = 0.561 and Evaluation Tip (Unseen): F(3, 29) = 0.714, p = 0.551. The main text should report the correct test statistics; the null conclusion is unaffected, but the internal inconsistency must be corrected.","section":"Section 6.2 and Supplement Table 5"}],"minor_comments":[{"comment":"The typo \"Bonerroni\" should be \"Bonferroni\".","section":"Supplement Section 5"},{"comment":"The p-value formatting is inconsistent (e.g., \"5 .82 × 10−2\" contains stray spaces); please standardize the scientific notation.","section":"Supplement Table 4"},{"comment":"The heading \"Independent Variables\" is followed by \"We have one independent variable in our experiment\"; the plural \"Variables\" should be \"Variable\" for consistency.","section":"Section 5, first paragraph"},{"comment":"The silhouette score that determined the number of clusters K is mentioned but no numeric values for K = 2, 3, 4, 5 are reported; providing those scores would make the clustering choice more transparent.","section":"Section 4.1"},{"comment":"The claim of being \"the first evaluation of a cooperative asynchronous training system in a partially observable, multi-agent setting\" is strong; consider adding a qualifying phrase such as \"to our knowledge\" and briefly checking the recency of related literature.","section":"Abstract and Section 2.2"}],"recommendation":"major_revision","confidential_remarks":"The cluster-persistence issue is the central risk to the paper's main contribution; a focused validation would substantially strengthen the design claim. The post-hoc NULL condition and the misreported F-statistic also need attention, but they are local. The authors' honest reporting of the null training effect is a positive aspect."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know two things about this paper. First, it is a real methodological contribution: clustering demonstrator behaviors to compress the condition space in a cooperative training study is a sensible idea, and the new HYBS game is a well-designed partially observable testbed with distinct roles. Second, the paper's central design claim is shakier than the authors let on, because the behavior clusters that justify the reduced conditions are never shown to persist when chefs leave self-play and enter evaluation with a waiter.\n\nThe strongest part is the honesty. The training effect is null, and they report it without spin. The statistical analysis is careful—normality checks, Tukey corrections, confidence intervals—and the subjective metrics are interpretable. The negative result on Q3 actually strengthens the paper as a design demonstration rather than a hype-driven claim.\n\nNow the soft spots, in order of severity. The stress-test concern lands: Section 5 says they did not enforce established chefs' compliance with their previous behavior, and Section 4.1 acknowledges distribution shift from full-observability self-play. If a chef assigned to Cluster 1 behaves like Cluster 2 during evaluation, then the in-sample/out-of-sample contrast that the whole condition-reduction design depends on collapses. Nothing in the paper checks this. That is not a fatal flaw in the idea, but it is load-bearing and currently unsupported. A revision could validate cluster stability with a small follow-up or by reporting per-evaluation-session behavior labels.\n\nSecond, the NULL condition was added post-hoc. That weakens the strongest negative claim (no benefit of training) because the control group was not randomized concurrently. Third, the sample is small (n=52 total, roughly 13 per condition), so the null result is low-powered. Fourth, no code or data are released, which hurts reproducibility for a paper whose main offer is a reusable experimental method.\n\nWho is this for? Researchers designing human-subject studies with asynchronous training or surrogate teammates. They will get value from the clustering idea and the candid lessons learned. The paper deserves a serious referee, but the referee should push for evidence that the clusters used to assign conditions actually describe behavior during evaluation.\n\nMy recommendation: engage with it, send it to review, and ask for a revision that either validates cluster stability or softens the methodological claim accordingly.","headline":"A genuinely useful experimental design with an honest null result, but its condition-reduction mechanism rests on an unvalidated assumption about cluster stability—worth reviewing, not desk-rejecting.","tokens_in":18337,"tokens_out":1416,"would_cite":false,"duration_ms":16584,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Behavior clustering of human demonstrations makes asynchronous cooperative-training experiments tractable.","keywords":["asynchronous training","cooperative training","human-machine teaming","behavior clustering","partially observable environment","human-subject experiment","Overcooked-AI","autonomous teammate"],"falsifier":"Re-encode each established chef's evaluation-session trajectories with the same encoder used in the study and compare their cluster assignments to the clusters assigned from self-play: if a large fraction of evaluation trajectories falls in a different cluster than assigned, or if one chef's trajectories scatter across clusters, then the in-sample versus out-of-sample contrast is not measuring what the design claims and the condition-reduction argument loses its foundation.","tokens_in":17288,"feed_emoji":"🎮","tokens_out":5444,"duration_ms":51456,"temperature":0.7,"pith_summary":"This paper claims that an asynchronous cooperative-training experiment can be made feasible by clustering human demonstrators into a few behavior types rather than treating each person as a separate training condition. The authors build a partially observable two-role game, Have You Been Served?, in which waiters coordinate with chefs who hold different information, and they train waiters with a human chef, an imitation-learning apprentice chef, a heuristic chef, or no training. A sequential encoder maps chef decision trajectories into a low-dimensional space, and K-medoids clustering groups them into three behavior clusters, after which waiters are evaluated with a chef from their training cluster and with a chef from a different cluster. The study's own training intervention did not significantly change evaluation outcomes, but the paper argues that the clustering-based design is what makes such a study tractable and offers design recommendations for future asynchronous team training.","feed_headline":"Behavior clustering shrinks cooperative-training study conditions","feed_subtitle":"Grouping chef playstyles cuts the number of teams needed to test agent training partners.","key_machinery":"The central object is the behavior cluster: a group of chef demonstrations that share a style, found by training a decoder-free sequential encoder on state-action sequences, projecting each game into a two-dimensional representation, and applying K-medoids clustering with silhouette-score selection. The clusters do two jobs: they let the experiment define in-sample versus out-of-sample evaluation partners, and they condition the apprentice agent's personalized embeddings so that an autonomous chef can exhibit a cluster's style. The apprentice agent is trained with per-cluster embeddings using a personalized apprenticeship-learning objective, while the heuristic chef combines Monte Carlo Tree Search over macro-actions with a Goal-Oriented Action Planner and is programmed to obey waiter recommendations.","core_discovery":"On its own terms, the paper's central claim is that behavior clustering is the load-bearing experimental device: segmenting the design by learned behaviors instead of by individual demonstrators drastically reduces the number of conditions and hence participants required for evaluation. The discovery is methodological, a way to turn a factorial pairing problem into a small set of behavior-conditioned pairings, and it is demonstrated in a real human-subjects study. Within that demonstration, the training manipulation produced a null result on evaluation tips, and waiters rated human chefs significantly higher than both agent chefs on trust, reliance, adaptability, and predictability despite the two agents differing significantly in training performance.","pith_inferences":["If chef behavior clusters are stable across partners and game contexts, the same clustering pipeline could be reused to design human-subjects experiments in any cooperative domain where a large population of demonstrators supplies the training partners.","A natural next experiment would vary the enforcement of cluster identity during evaluation, making some chefs deliberately switch styles, to measure how much of the in-sample versus out-of-sample effect is due to the waiter's adaptation rather than the chef's consistency.","The equal subjective ratings of the two agents despite their score gap suggest that perceived humanness, such as responsiveness to waiter recommendations and predictable staging habits, may be a better design target than task performance when building robot training partners."],"forward_implications":["Future cooperative-training studies can stratify evaluation by behavior clusters rather than by individual demonstrators, reducing the factorial design from many waiter-chef pairs to a handful of behavior-conditioned pairings.","Because the two agent chefs differed in training tips but were rated similarly by waiters, subsequent surrogate agents should be tuned for expressed behavior, not raw score, to improve the subjective training experience.","The null evaluation-tip result implies that training benefits may be masked unless training and evaluation are matched in observability and scenario structure; the paper recommends deliberate scenario design and proficiency checks.","Autonomous chefs can be generated per behavior cluster, with the apprentice conditioned on per-cluster embeddings and the heuristic given staged-ingredient variants, so waiters can practice with a chosen style."],"supporting_citations":[{"why":"Supplies the decoder-free sequential encoder architecture used to embed chef trajectories into two dimensions.","marker":"[26]"},{"why":"Provides the unsupervised behavior-inference and clustering procedure that turns demonstrations into behavior clusters.","marker":"[11]"},{"why":"Gives the personalized apprenticeship-learning algorithm used to train the apprentice chef on per-cluster embeddings.","marker":"[9]"},{"why":"Defines Goal-Oriented Action Planning, used for low-level action planning in both agent chefs.","marker":"[20]"},{"why":"Provides the Overcooked-AI environment that HYBS adapts with waiter and chef roles and partial observability.","marker":"[25]"},{"why":"Supplies Monte Carlo Tree Search, used by both agents for high-level macro-action planning.","marker":"[29]"},{"why":"Motivates the need for shared mental models of teammates in partially observable teaming, the training goal the study targets.","marker":"[6]"},{"why":"Supports the claim that humans judge AI teammates on behavior rather than only objective performance.","marker":"[24]"}],"fun_headline_variants":["Behavior clustering shrinks study conditions for cooperative training","Clustering behaviors reduces conditions in training experiments","Autonomous teammates lose trust vs human partners in training","Agent training partners rated lower on trust and adaptability"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the behavior clusters learned from chefs playing alone with full information still describe those chefs' behavior when they play with a waiter under partial observability, even though the experiment never forced chefs to stick to their assigned style.","fun_headline_variants_meta":{"raw":{"variants":["Behavior clustering shrinks study conditions for cooperative training","Clustering behaviors reduces conditions in training experiments","Autonomous teammates lose trust vs human partners in training","Agent training partners rated lower on trust and adaptability"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000263,"raw_usage":{"total_tokens":1555,"prompt_tokens":854,"completion_tokens":701,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":470,"completion_tokens_details":{"reasoning_tokens":642}},"tokens_in":470,"tokens_out":701,"duration_ms":7899,"temperature":1.0,"reasoning_tokens":642,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T05:07:49.334761+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-encode each established chef's evaluation-session trajectories with the same encoder used in the study and compare their cluster assignments to the clusters assigned from self-play: if a large fraction of evaluation trajectories falls in a different cluster than assigned, or if one chef's trajectories scatter across clusters, then the in-sample versus out-of-sample contrast is not measuring what the design claims and the condition-reduction argument loses its foundation.","supporting_citations":[{"cited_title":"Unsupervised scalable representation learning for multivariate time series","cited_arxiv_id":null,"evidence_quote":"Supplies the decoder-free sequential encoder architecture used to embed chef trajectories into two dimensions."},{"cited_title":"Unsupervised behavior inference from human action sequences","cited_arxiv_id":null,"evidence_quote":"Provides the unsupervised behavior-inference and clustering procedure that turns demonstrations into behavior clusters."},{"cited_title":"Interpretable and personalized apprenticeship scheduling: Learning interpretable scheduling policies from heterogeneous user demonstrations","cited_arxiv_id":null,"evidence_quote":"Gives the personalized apprenticeship-learning algorithm used to train the apprentice chef on per-cluster embeddings."},{"cited_title":"Symbolic representation of game world state: Toward real-time planning in games","cited_arxiv_id":null,"evidence_quote":"Defines Goal-Oriented Action Planning, used for low-level action planning in both agent chefs."},{"cited_title":"On the utility of learning about humans for human- ai coordination","cited_arxiv_id":null,"evidence_quote":"Provides the Overcooked-AI environment that HYBS adapts with waiter and chef roles and partial observability."},{"cited_title":"Bandit based M onte- C arlo planning","cited_arxiv_id":null,"evidence_quote":"Supplies Monte Carlo Tree Search, used by both agents for high-level macro-action planning."},{"cited_title":"The role of shared mental models in human- AI teams: a theoretical review","cited_arxiv_id":null,"evidence_quote":"Motivates the need for shared mental models of teammates in partially observable teaming, the training goal the study targets."},{"cited_title":"Evaluation of human- AI teams for learned and rule-based agents in H anabi","cited_arxiv_id":null,"evidence_quote":"Supports the claim that humans judge AI teammates on behavior rather than only objective performance."}],"review_version":1}