{"id":"43589abb-5fdf-4a65-a0fd-a1bd7e2f72de","arxiv_id":"2509.04876","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":4,"one_line_summary":"OSC uses learned Collaborator Knowledge Models and RL-trained communication policies to make LLM agents communicate adaptively, claiming gains on AlpacaEval 2.0 and MT-Bench.","lead":"This paper presents OSC, a framework that trains large language model agents to model each other's knowledge states and adapt their messages through reinforcement learning. It reports higher AlpacaEval and MT-Bench scores than some multi-agent baselines, but the evaluation has significant flaws.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 81.4% AlpacaEval headline is an in-sample number: the evaluation set includes the 160-instruction PPO training subset, and no reasoning-benchmark results are reported despite the abstract's claim. The central claim is currently unsupported.","rationale":"The reader's weakest_assumption identifies the same load-bearing concern: the evaluation set is not cleanly separated from the training set, and the abstract's reasoning-benchmark claim is not backed by experiments. This is the single most important issue because the entire paper's acceptance hinges on the 81.4% AlpacaEval result and the claimed reasoning-benchmark superiority. If the headline number is contaminated by in-sample optimization, the main comparison to KABB and MoA collapses; if no reasoning benchmark is actually run, the abstract's central validation claim is unsupported. The concern is concrete and textually grounded: §4.4, §5, §9, and §10 all state or imply that a subset of AlpacaEval is used for PPO training, while Table 1 reports performance on AlpacaEval 2.0 without excluding that subset. The paper also contains multiple internal inconsistencies (e.g., §4.1 says 6 experts and Qwen2-72B aggregator, table 1 OSC-Single-LLaMa3 uses one model; §11 describes a different 13B configuration with 78.6% result), but the evaluation-contamination issue is the most load-bearing because it directly undermines the central claim and the comparison to baselines. The proposed test—evaluating on the 645 held-out AlpacaEval instructions and a held-out reasoning benchmark—would settle whether the reported advantage is real generalization or an artifact of training on the test benchmark. I therefore agree with the reader's verdict: the paper should be rejected or, at minimum, substantially revised with a clean evaluation.","tokens_in":20754,"tokens_out":4186,"duration_ms":44905,"concrete_test":"Run a clean evaluation protocol: train OSC exactly as in §5 on the 160-instruction development subset, then compute the LC win rate on the remaining 645 AlpacaEval instructions (excluding all training and validation instances) and compare against KABB, MoA, and DeepSeek-R1 reproduced under the same protocol. Also evaluate the same trained system on a held-out reasoning benchmark such as MATH-500 or GSM8K, which the abstract claims to validate. If the 645-instance LC win rate no longer exceeds KABB by a meaningful margin, or the reasoning-benchmark improvement does not replicate, the central claim is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central empirical claim—that OSC significantly improves task performance and communication efficiency on complex reasoning benchmarks—rests on AlpacaEval 2.0 and MT-Bench results, with the headline 81.4% LC win rate. That number is not a clean out-of-sample measurement. Section 4.1 states evaluation is on AlpacaEval 2.0 (805 instructions), and Table 1 reports 81.4% for OSC. But §4.4 says the study 'utiliz[es] 805 instructions for training and evaluation, with specific subsets of 160 instructions reserved for development and validation respectively.' Section 5 explicitly says fine-tuning uses '160 for fine-tuning, 160 for validation' on AlpacaEval 2.0; §9 and §10 confirm training on the development subset. The 805-instruction benchmark therefore contains the 160 instructions used for PPO training, so the headline result is partly in-sample and not comparable to public-leaderboard baselines that were not trained on the benchmark. Even if the 81.4% only corresponds to the 160-instruction validation subset, that would invalidate the comparison with baselines evaluated on the full 805-instruction set. Separately, the abstract and §1 claim validation on 'complex reasoning and problem-solving benchmarks' and specifically MATH, yet no MATH or other reasoning result table appears anywhere; only AlpacaEval (instruction following) and MT-Bench (chat) are reported. The framework is not inherently implausible, but the central claim of significantly improved performance on complex reasoning is not supported by the evidence as presented.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes OSC, a multi-agent LLM collaboration framework with trainable Collaborator Knowledge Models (CKM), a learnable cognitive-gap function fgap, and a reinforcement-learned communication policy pi_comm. The authors claim that OSC significantly improves task performance and communication efficiency, reporting an 81.4% LC win rate on AlpacaEval 2.0, a 9.94 average on MT-Bench, and a Pareto-optimal cost/performance curve. The abstract and introduction also claim validation on complex reasoning benchmarks such as MATH. The main evidence is Tables 1-4 and several appendix experiments, with the headline result trained via PPO on a subset of the evaluation benchmark.","tokens_in":21169,"tokens_out":3485,"duration_ms":35331,"significance":"If the central claims were cleanly supported, this would be a useful contribution: it addresses an underexplored intermediate layer between expert selection and answer aggregation, and its components are explicitly trainable with RL, moving beyond static role- or debate-based communication. The idea of modeling collaborators' cognitive states and learning communication objectives is plausible and potentially impactful. However, the paper's current evidence does not support the advertised claims: the headline evaluation is contaminated by training on the same benchmark, the promised reasoning-benchmark results are absent, and no uncertainty quantification is provided for the reported improvements. These are fixable in a revision, but they are load-bearing.","major_comments":[{"comment":"The headline 81.4% LC win rate is not a clean out-of-sample comparison. §4.4 states that AlpacaEval 2.0 uses 805 instructions for training and evaluation with 160 reserved for development and 160 for validation; §5 says fine-tuning uses 160 of the 805 instructions; §10 confirms that the development set is used for training and the validation set for evaluation. Thus the OSC result reported in Table 1 appears to be measured on a 160-instruction validation subset, while the baseline values (KABB, MoA, DeepSeek-R1, etc.) are either public-leaderboard numbers on the full 805-instruction set or reproduced on the full set. This makes the comparison invalid. The authors must either evaluate on the full 805-instruction set without training on it, or evaluate all baselines on exactly the same held-out subset.","section":"§4.4, §5, §10, Table 1"},{"comment":"The abstract and §1 claim validation on 'complex reasoning and problem-solving benchmarks' and specifically MATH, and §1 lists MATH as a contribution. §7.6 states that training environments are constructed from MATH and GSM8K. However, no MATH, GSM8K, or any other reasoning benchmark result appears anywhere in the paper. The only evaluations reported are AlpacaEval 2.0 (instruction following) and MT-Bench (chat). The authors should either provide the missing reasoning-benchmark results or substantially revise the claims to match the actual experiments.","section":"Abstract, §1, §4.1, §7.6"},{"comment":"The paper repeatedly uses the word 'significantly' (e.g., §1, §4.1, §4.3) but reports no error bars, confidence intervals, or significance tests. §4.4 says results are averaged over 3 runs, yet Table 4 contains only point estimates. The differences between OSC and KABB on AlpacaEval (81.4 vs. 77.9) and between OSC and DeepSeek-R1 (81.4 vs. 80.1) are small; without variance or statistical testing, the claimed significance is unsupported. Add intervals or tests, at least for the headline metric.","section":"Tables 1-4"}],"minor_comments":[{"comment":"The baselines TalkHier, REMALIS, DyLAN, and MAC are not defined or referenced. Their experimental setup, model pool, and metric definitions are missing, so the communication-efficiency comparison cannot be reproduced or assessed.","section":"Table 2"},{"comment":"The phrase 'To address reviewer requests 6' should be removed; it is not appropriate in a submitted manuscript.","section":"§4.6"},{"comment":"There are incomplete sentences. For example, §4.1 ends with 'Additional assessments include MT-Bench(...)' and §4.2 contains the repeated heading 'Experimental Setup Experimental Results' with no setup description.","section":"§4.1, §4.2"},{"comment":"Stray footnote markers or placeholder characters appear as '4' at the end of sentences (e.g., '...dynamic collaboration benefits 4.' and '...reasonable rounds and token count 4.'). These should be cleaned up.","section":"§5, §9"},{"comment":"The reference to 'Equation 6 in Section 3.4' is off: the state representation is Eq. (5) in Section 3.4. Equation (6) is the message-generation equation.","section":"§7.3"}],"recommendation":"major_revision","confidential_remarks":"The in-sample evaluation issue is the most serious problem. If the authors cannot provide a clean evaluation on AlpacaEval 2.0 or on a properly matched held-out subset, the paper's central performance claim collapses. The missing reasoning-benchmark results are also a direct mismatch with the abstract. Given that both are fixable with additional experiments, I recommend major revision rather than outright rejection, but the revision must substantively address these points."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Hi,\n\nThe short version: OSC is a real architectural idea, but the empirical case for it is not credible as reported. The headline 81.4% AlpacaEval win rate is not an out-of-sample number, and the paper never reports the reasoning-benchmark results its abstract promises. I would not trust the central claim without a clean redo.\n\nWhat's new: the specific combination — collaborator knowledge models, a learned cognitive gap function, and an RL-trained communication policy — is a new recipe, and the authors describe the components in enough detail (appendices 7–11) that someone could build a version of it. The scalability and price-performance plots are also the right kind of analysis to run for a method like this. That is real work, and it should be credited.\n\nThe soft spots are serious. Section 4.4 says the AlpacaEval 2.0 study used 805 instructions for training and evaluation, with 160 for development and 160 for validation. Sections 5, 9, and 10 all say training is done on the development subset and evaluation on the validation subset. So Table 1's 81.4% is either computed on the 160-instruction validation subset (which makes it non-comparable to baselines scored on the full 805) or it includes the training set. Either way, the comparison to public leaderboard numbers is invalid. The abstract and Section 1 say the method is validated on complex reasoning and problem-solving benchmarks and cite MATH, but no reasoning benchmark appears anywhere in the results. The communication-efficiency metrics (redundancy, conflict resolution, information density) are asserted but never formally defined, and the tables lack error bars even where the text says results were averaged over three runs.\n\nNone of this means the framework is impossible. With a clean evaluation — held-out reasoning benchmarks, a development set that is not part of the reported test set, and defined communication metrics — the method could easily justify a paper. But as it stands, the evidence for the central claim is not there.\n\nWho gets value from this: someone building a multi-agent communication layer and looking for architectural ideas, and anyone teaching a class on evaluation pitfalls in LLM systems. It deserves a serious referee only on the condition that the referee sees revision as mandatory. I'd send it to review, expecting a reject-and-resubmit outcome rather than acceptance.\n\n— [name]","headline":"The OSC framework is a real architectural idea, but the headline AlpacaEval result is in-sample and the promised reasoning benchmarks are missing, so the central claim is unsupported as reported.","tokens_in":21629,"tokens_out":2605,"would_cite":false,"duration_ms":26698,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A learned collaboration layer that lets each LLM agent model what its teammates know lifts AlpacaEval 2.0 win rate to 81.4% and cuts communication overhead.","keywords":["multi-agent LLM systems","collaborator knowledge model","cognitive gap analysis","adaptive communication policy","reinforcement learning","PPO","AlpacaEval 2.0","agent collaboration"],"falsifier":"Take the 805 AlpacaEval instructions, remove the 160-instruction development subset from all PPO training and hyperparameter tuning, then evaluate on the remaining held-out instructions plus MATH and GSM8K; if the length-controlled win-rate margin over KABB shrinks to noise or reverses, the cognitive-alignment explanation for OSC's advantage is not supported.","tokens_in":1687,"feed_emoji":"🧠","tokens_out":1604,"duration_ms":72618,"temperature":0.7,"pith_summary":"The paper claims that the main bottleneck in multi-agent LLM systems is neither which experts are chosen nor how their outputs are merged, but how they talk to one another. It introduces OSC, an intermediate collaboration layer in which each agent maintains a learned Collaborator Knowledge Model (CKM) of every teammate's current knowledge, confidence, and task understanding. A learned cognitive-gap function compares that model with the agent's own state, and a PPO-trained communication policy decides whom to address, which gap to close, what objective to pursue, and in what style. The headline result is an 81.4% length-controlled win rate on AlpacaEval 2.0, above KABB (77.9%) and MoA (68.1%), with fewer rounds and tokens than comparison systems. A sympathetic reader would care because it reframes multi-agent performance as a communication-design problem and offers a trainable mechanism for making agents' messages contingent on collaborators' inferred mental states.","feed_headline":"Teammate-aware LLM agents hit 81.4% on AlpacaEval","feed_subtitle":"Learned collaborator models and gap-aware messaging beat fixed multi-agent baselines with fewer tokens.","key_machinery":"The Collaborator Knowledge Model (CKM) is the load-bearing mechanism: a per-agent-pair latent state, produced by a Transformer encoder and updated by a GRU, that encodes agent i's evolving belief about agent j's knowledge, confidence, and task understanding. It carries the argument because every downstream decision—gap analysis, target selection, communication objective, and style choice—is a function of these learned beliefs, and because the CKM, the gap function, and the policy are fine-tuned end-to-end from the same task reward, the representations are shaped by whether they lead to successful collaboration.","core_discovery":"OSC claims that the way expert agents communicate, rather than the choice of experts or the aggregation of their answers, is what limits deep multi-agent LLM collaboration. To address this, each agent maintains a Collaborator Knowledge Model (CKM): a learned latent vector, updated by a GRU as dialogue proceeds, representing what that agent believes another agent knows, believes, or misunderstands about the task. A learnable gap function compares the agent's own cognitive state with the CKM-derived state of each collaborator, and a Proximal Policy Optimization (PPO) trained policy picks the communication action: which collaborator to address, which gap to close, what objective to pursue, and","pith_inferences":["The headline comparison is partly in-sample: the number of communication rounds and the communication-cost weight were tuned on a 160-instruction development subset of AlpacaEval 2.0, the same benchmark where the reported 81.4% win rate is measured; an independent hold-out test could yield a smaller margin.","The abstract says OSC was validated on complex reasoning and problem-solving benchmarks, but the reported experiments use AlpacaEval 2.0 and MT-Bench; MATH appears only as a stated training environment. Testing OSC on a clean reasoning benchmark such as MATH or GSM8K with no overlap between training and test tasks is the natural next check.","The same learned communication mechanism could transfer beyond LLM agents—to agents that are code modules, robots, or other learned communicators—whenever a state model and a language-realization layer exist.","The CKM's latent vectors could be probed against human judgments of confusion or confidence; if they do not track those states, OSC's advantage may come from extra computation or more conversational turns rather than from genuine cognitive modeling."],"forward_implications":["If OSC's mechanism works as claimed, teams of off-the-shelf LLM experts can be made to collaborate more effectively by inserting a learned communication layer, without redesigning expert selection or aggregation.","Adaptive, gap-targeted messaging allows the same task quality with fewer dialogue rounds and fewer tokens, lowering inference cost at a fixed performance level.","Because the CKM, gap function, and policy are trained end-to-end from a composite reward, the same machinery can be optimized for other objectives, including explicit cost constraints or communication budgets.","The reported optimal team size is 6 agents; with 8 or 10 agents, coordination overhead—more rounds, more tokens, lower conflict resolution—begins to erode the gains.","A single-model variant of OSC outperforms the same model without the collaboration layer, suggesting the benefit does not depend on combining different model families."],"supporting_citations":[{"why":"Supplies AlpacaEval 2.0, the instruction-following benchmark and GPT-4-based evaluator used for the headline LC win rate.","marker":"Li et al., 2023b"},{"why":"Supplies KABB, the main multi-agent baseline whose expert pool and aggregator configuration OSC reuses.","marker":"Zhang et al., 2025d"},{"why":"Supplies MoA, the aggregation-based multi-agent system that OSC must beat on AlpacaEval 2.0 and MT-Bench.","marker":"Wang et al., 2024a"},{"why":"Provides the PPO algorithm used to train OSC's adaptive communication policy and to fine-tune the CKM and gap modules.","marker":"Schulman et al., 2017"},{"why":"Supplies MT-Bench, the multi-turn dialogue benchmark used as the second reported evaluation.","marker":"Zheng et al., 2023"}],"fun_headline_variants":["Agents that model teammates' knowledge gaps beat fixed multi-agent baselines","LLM agents that adapt to each other's cognitive states gain reasoning edge","OSC: Learn when to explain, ask, or simplify in multi-agent LLM teams","With CKM, LLM agents outperform by understanding what peers know","Collaborator knowledge models let LLM agents hit 81.4% on AlpacaEval"],"cache_read_input_tokens":23296,"weakest_assumption_plain":"The result stands or falls on the evaluation being a fair, out-of-sample test of OSC's learned communication layer: the reported 81.4% win rate is measured on AlpacaEval 2.0 after tuning on its development subset, and the promised validation on complex reasoning benchmarks is not present in the reported experiments.","fun_headline_variants_meta":{"raw":{"variants":["Agents that model teammates' knowledge gaps beat fixed multi-agent baselines","LLM agents that adapt to each other's cognitive states gain reasoning edge","OSC: Learn when to explain, ask, or simplify in multi-agent LLM teams","With CKM, LLM agents outperform by understanding what peers know","Collaborator knowledge models let LLM agents hit 81.4% on AlpacaEval"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00036,"raw_usage":{"total_tokens":1753,"prompt_tokens":683,"completion_tokens":1070,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":427,"completion_tokens_details":{"reasoning_tokens":965}},"tokens_in":427,"tokens_out":1070,"duration_ms":10645,"temperature":1.0,"reasoning_tokens":965,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T05:47:28.357660+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the 805 AlpacaEval instructions, remove the 160-instruction development subset from all PPO training and hyperparameter tuning, then evaluate on the remaining held-out instructions plus MATH and GSM8K; if the length-controlled win-rate margin over KABB shrinks to noise or reverses, the cognitive-alignment explanation for OSC's advantage is not supported.","supporting_citations":[],"review_version":1}