{"id":"32ca8d32-1730-4c3d-b8af-c19c6472a962","arxiv_id":"2608.07583","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"Routing gain among LLM advisors can be certified with a finite-sample bracket and a matching minimax lower bound, and certification fails on uninformative gates and statistically redundant advisor pools.","lead":"RouteGuard is a statistical guardrail for LLM multi-agent routing: from a batch of evaluation examples it either certifies a lower bound on the deployable accuracy gain over the best single model, or refuses to certify. The authors show that a gate's AUC and advisor complementarity do not determine whether routing helps, and that on the benchmarks tested most gains are too small, too fragile, or too redundant to certify.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Plug-in variance in the certification bracket invalidates the literal 1−δ guarantee; Appendix K's empirical-Bernstein fix is required but changes m⋆.","rationale":"Read in good faith, the paper's theoretical core is strong: T1 and Prop. 1 are correct, AUC-insufficiency witnesses are valid, the ITV ceiling and redundancy analysis are internally consistent, and the authors supply extensive numerical verification. The single load-bearing concern is the plug-in variance in the certification bracket. The manuscript itself flags this in §3.6 and Appendix K: 'the plug-in σ̂² remains, bounded by the empirical-Bernstein bracket below' and 'an empirical-Bernstein bound that assumes no σ² needs 810.' So the reader's weakest assumption is correct and is self-acknowledged. This concern does not undermine the empirical verdicts, which rely on large m or unambiguous refusals, but it does invalidate the literal claim of a 1−δ guarantee at m=312. The appropriate disposition remains CONDITIONAL: the protocol is fixable by replacing the plug-in bracket with the empirical-Bernstein bound, which the paper already computes. The stress-test therefore agrees with the reader and sees no reason to move the verdict.","tokens_in":36851,"tokens_out":3358,"duration_ms":32010,"concrete_test":"Monte Carlo coverage test on the audited law P(Z=+1)=0.04, P(Z=0)=0.96, P(Z=−1)=0 (G=0.04, σ²=0.0384). For m=312 (the paper's m*) and m=810, draw 100,000 samples, compute σ̂² from each sample, form γ̂_m = bG_m − B(m,0.05;σ̂), and record whether γ̂_m ≤ G (true population gain). If the empirical coverage of the one-sided lower bound is below 0.95 at m=312, the plug-in bracket's 1−δ claim fails. Repeat with the empirical-Bernstein bracket of Appendix K; if that holds coverage ≥0.95 at m=810, the fix is confirmed and only the headline m* changes.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The certification claim in §4.4 is that the protocol 'certifies G_μ(R̂) ≥ γ̂_m ≥ 0 w.p. ≥ 1−δ'. The bracket B(m,δ) in §3.6 is a Bernstein bound for the true per-sample variance σ². In the protocol, σ² is replaced by σ̂² estimated from the same m samples. Bennett's inequality provides a valid tail bound only for the true variance; replacing σ by σ̂ can arbitrarily shrink the bracket when σ̂ underestimates σ, so the 1−δ coverage statement is not implied. Appendix K concedes that σ̂² is the only remaining estimated quantity and reports that a valid empirical-Bernstein bound requires m*=810 instead of 312. The paper treats Appendix K as a robustness check, but it actually shows the headline m*=312 is not a valid sample size for the claimed guarantee. The reported verdicts may survive (Bank refuses with m⋆>135 under any bracket, and RouterBench has m=18230), but the central finite-sample coverage theorem lacks a proof for the plug-in protocol as written. This is a correctness risk in the paper's main deliverable.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes RouteGuard, a finite-sample certification framework for routing gain in LLM multi-agent systems. It proves that the largest achievable routing gain equals the gate's informativeness functional Phi = E[max_j eta_j(T) - eta_1(T)] (Theorem 1), that the gain decomposes as G = pi * Delta_E (Proposition 1), and that AUC is not a sufficient statistic for Phi. It also proves an informativeness ceiling Phi <= (1/2) ITV(T), establishes that complementarity is necessary but not sufficient for positive gain, derives a Bernstein-type certification bracket with a matching Le Cam lower bound, and evaluates the protocol on RouterBench and OpenRCA. On RouterBench the verdict depends on the sampling unit; on OpenRCA the protocol refuses to certify. A pre-registered semi-synthetic control is used to show calibration. The paper includes extensive appendices with proofs, numerical verification tables, and frozen artifact specifications.","tokens_in":1275,"tokens_out":1346,"duration_ms":56819,"significance":"If the results hold, the paper makes a substantive contribution: it identifies the right functional for routing gain, argues convincingly against AUC as the design objective, and connects the certification problem to a Le Cam lower bound with a constant-sharp class-level statement. The strengths are real: Theorem 1 and the characterization of equality in the informativity ceiling are proven rigorously; the lower-bound constants in Appendix E are derived carefully and independently verified; the empirical claims are supported by frozen artifacts and a table of numerical checks; and the pre-registered positive control addresses calibration in a way that is rare in this literature. These strengths make the paper worth serious consideration. The main weakness is that the headline finite-sample certificate as implemented plugs in an estimated variance, so the literal 1-delta coverage claim is not implied by the stated Bernstein bound; this is acknowledged in Appendix K but the presentation still puts the invalid bracket at center stage.","major_comments":[{"comment":"The certification protocol as written does not have the claimed 1-delta coverage, because the Bernstein bracket B(m, delta) in §3.6 is valid for the true variance sigma^2, while the pseudocode in §4.4 replaces sigma^2 by the plug-in sigma-hat^2 computed from the same m samples. If sigma-hat^2 underestimates sigma^2, the bracket shrinks and the stated guarantee \"certifies G_mu(R-hat) >= gamma-hat_m >= 0 w.p. >= 1-delta\" can fail at finite m. Appendix K concedes that sigma-hat^2 is the only remaining estimated quantity and reports that a valid empirical-Bernstein bound requires m* = 810 instead of the headline m* = 312. Thus the main deliverable lacks a proof for the plug-in protocol as written, and the sample size m* = 312 cannot be claimed for the literal guarantee. The empirical verdicts may survive (Bank still refuses at m=135 and RouterBench has m=18,230), but the paper should either adopt an empirical-Bernstein certificate in the main protocol, provide a valid joint bound for the plug-in procedure, or restate precisely what m* = 312 certifies.","section":"§4.4 and Appendix K"},{"comment":"The leading-order sharpness theorem (Theorem 4) and the class-level optimality claim (Corollary 3) are proved for a certificate that either knows sigma^2 or uses the class variance cap V* = pi, not for the protocol's plug-in sigma-hat. As a result, the sentence in §3.6 that the protocol is \"asymptotically minimax-optimal\" over the fixed-activity class applies to the variance-capped certificate, not to the implemented bracket. The relationship between the known-variance theory and the implemented plug-in protocol should be made explicit, or the optimality claim should be limited to the capped certificate.","section":"§3.6, §4.4, and Appendix E.1"}],"minor_comments":[{"comment":"The text \"cb-2 apost-hoc third distribution\" should read \"a post-hoc third distribution\".","section":"§5.2"},{"comment":"The comparison \"m* = 312 > 136\" would be cleaner as \"m* = 312 > 135\", since the parseable Bank count is 135.","section":"§5.1"},{"comment":"The cluster-robust adjustment in Protocol step 4 uses a design effect with an ICC estimated from the same data; the paper should state whether this adjustment is an asymptotic approximation or comes with its own finite-sample coverage guarantee.","section":"§4.4"},{"comment":"The statement \"With bG=0 there is no gain to certify at any m\" in §5.1 is too strong as a population statement: bG=0 is an estimate, and a larger sample could in principle reveal a positive gain; the surrounding argument works because the bracket refuses, but the phrasing should be softened.","section":"§3.6"}],"recommendation":"major_revision","confidential_remarks":"The conceptual core of the paper is sound and unusually well verified, and the plug-in variance issue is fixable within the manuscript's scope by adopting the empirical-Bernstein certificate as the primary protocol or by proving a valid plug-in bound. I therefore recommend major revision rather than rejection. The revision should also ensure that the optimality claims in §3.6 and Appendix E.1 are attributed to the certificate for which they are actually proved."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read it. The mathematical core is mostly right. The clean new thing is Theorem 1: max routing gain equals the gate's informativeness Φ, not its AUC, with the two AUC=1/2 witnesses (Φ=0 vs 1/4) being genuinely sharp and easy to verify. The ITV ceiling, the no-Boolean-NSC construction, and the Le Cam/Bretagnolle-Huber lower bounds are all careful, and the numerical verification against frozen golden files is a real strength. The empirical work is unusually honest: the cluster-resampling flip on RouterBench, the post-hoc cb-2 disclosure, and the refusal on redundant OpenRCA pools are reported plainly rather than buried. The E≤0 screen across 221 pools is a useful cautionary finding. Now the soft spot, and it is real. The protocol in §4.4 plugs σ̂², estimated from the same m samples, into a Bernstein bound that is proven for the true σ². That does not give the claimed finite-m 1−δ coverage. The stress-test is correct: Bennett holds for the true variance, and a plug-in estimate can shrink the bracket when σ̂ underestimates σ. The paper essentially concedes this in Appendix K, where a proper empirical-Bernstein bound needs m*=810 instead of the headline 312. This does not destroy the empirical verdicts—RouterBench has m=18,230 and Bank refuses under any bracket—but it does break the headline claim that the protocol certifies at m*=312 with the stated guarantee. The fix is straightforward: use a valid empirical-Bernstein bound, or cap the variance conservatively, and state the theorem for the protocol actually executed. Two smaller issues: code and frozen artifacts are promised but not shipped, so the numbers are not independently checkable yet; and the pre-registration was amended after the fact (M=1+|G| changed to fixed M=2), which is disclosed but still a small transparency dent. The theoretical results stand on their own and are worth a serious referee. Who gets value: deployers of LLM routing and anyone working on statistical guarantees for routing. Recommendation: send to peer review, conditional on the authors fixing the variance proof and releasing artifacts. The core contribution is solid; the certificate needs to be stated and proved for what is actually run.","headline":"A mostly sound theory paper whose headline finite-sample certificate has a real plug-in-variance gap; the empirical screen is honest and the work deserves referee time.","tokens_in":37634,"tokens_out":2007,"would_cite":true,"duration_ms":22741,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The largest attainable gain of any LLM advisor router is exactly the gate's informativeness $\\Phi = \\mathbb{E}[\\max_j \\eta_j(T) - \\eta_1(T)]$, not its AUC, and a finite-sample certificate can either certify this gain before deployment or…","keywords":["LLM multi-agent routing","routing-gain certification","gate informativeness","AUC insufficiency","finite-sample concentration bound","minimax lower bound","error diversity","independence baseline"],"falsifier":"Simulate a true-null law on $\\{-1,0,1\\}$ with known variance, draw many $m = 312$ evaluation samples, and run the protocol's exact bracket with plug-in $\\hat\\sigma^2$ at $\\delta = 0.05$: if the false-certification rate exceeds 5% whenever the estimator understates the variance, the plug-in coverage claim fails.","tokens_in":36554,"feed_emoji":"🚦","tokens_out":6966,"duration_ms":59253,"temperature":0.7,"pith_summary":"This paper asks a deployer's question: before shipping an LLM multi-agent router, how can one know whether selecting among advisors will actually beat the strongest single advisor? The central claim is that the answer is governed by a single functional of the gating signal, its informativeness $\\Phi$, and not by anything so coarse as AUC. The paper proves that the largest gain any router can achieve equals $\\Phi$ exactly, attained by the Bayes selector, and factors every router's gain as the product of how often it routes away and how much better it does on those cases. It supplies a finite-sample certificate that either reports a high-confidence lower bound on the deployed gain or refuses to certify, together with a matching lower bound on how many samples any such test can need. On two benchmarks the protocol acts as a guardrail: it certifies a small real gain only under the sampling unit where that gain is robust, and it refuses to certify gains from statistically redundant advisor pools and uninformative gates.","feed_headline":"Routing gain is capped by gate informativeness, not AUC","feed_subtitle":"New protocol certifies a router's real gain before shipping and refuses when advisors are redundant.","key_machinery":"The load-bearing object is the conditional-regret functional $\\Phi(L) = \\mathbb{E}[\\max_j \\eta_j(T) - \\eta_1(T)]$, called the gate's informativeness: the expected per-instance improvement of the best posterior-correct advisor over the primary. It works by making the selection problem pointwise: any router's expected accuracy is $\\mathbb{E}[\\eta_{R(T)}(T)]$, which is bounded above by $\\mathbb{E}[\\max_j \\eta_j(T)]$, so the gap to the primary is exactly the best achievable gain. Around $\\Phi$ the paper builds the certification bracket $B(m, \\delta) = \\sigma\\sqrt{2\\log(1/\\delta)/m} + 2M\\log(1/\\delta)/(3m)$ on the bounded increments $Z_i = C_{R(T_i),i} - C_{1,i}$, with fixed $M = 2$, giving a lower confidence value $\\hat\\gamma_m = \\hat G_m - B(m, \\delta)$; the second support is the independence-baseline excess $E = A^\\star - (1 - \\prod_j (1-p_j))$, which diagnoses whether co-failure beyond chance lowers the routing ceiling. The same machinery yields the robustness phase transition at $\\rho^\\star = \\pi\\Delta_E/2$ and the minimax sample-size constant.","core_discovery":"On the paper's own terms, the discovery is that routing gain obeys an exact decomposition and an exact ceiling. For any router $R$, $G(R) = \\pi \\Delta_E$, where $\\pi$ is the probability of routing away from the primary and $\\Delta_E$ is the conditional accuracy edge on that routed set. Maximizing over all measurable routers gives $\\max_R G(R) = \\Phi$, where $\\Phi = \\mathbb{E}[\\max_j \\eta_j(T) - \\eta_1(T)]$ is the expected per-instance advantage of the best posterior-correct advisor over the primary, and the Bayes selector attains it. The gate's AUC is not a sufficient statistic for $\\Phi$: two laws with identical $\\mathrm{AUC} = 1/2$ can have $\\Phi = 0$ or $1/4$. The finite-sample consequence is a one-sided Bernstein bracket $B(m, \\delta)$ such that the certified gain $\\hat\\gamma_m = \\hat G_m - B(m, \\delta)$ lower-bounds the population gain with probability at least $1-\\delta$, with a two-point lower bound showing that any test needs $\\Omega(\\sigma^2 \\log(1/\\delta)/G^2)$ samples, constant-sharp over the fixed-activity class. Empirically, the protocol certifies on RouterBench only under prompt-level exchangeability and refuses on OpenRCA because the advisors co-fail beyond the independence baseline and the gate is uninformative.","pith_inferences":["The paper leaves implicit that the same decomposition should carry over to cost-aware routing: replace correctness increments with utility increments, and the bracket machinery still applies because the routing increment on $\\{-1,0,1\\}$ remains bounded.","The RouterBench verdict flip under clustering suggests that the sampling unit itself is a deployment decision: if future queries can come from new workload types, the correct certificate is the cluster-robust one, and any in-workload certification should be labeled as such.","A cheap pre-screen implied by the results is to compute $E$ on a candidate advisor pool before running expensive LLM evaluation: pools with $E \\le 0$ and uninformative gates are unlikely to yield certifiable gains, and only pools with positive $E$ deserve the full bracket.","A testable extension would construct pools with opposite-direction competence curves, genuine specialization, and verify that the protocol certifies positive gain only when both $E > 0$ and the gate's posterior separates advisors; this would check whether $E$ is a sufficient screen in practice."],"forward_implications":["If the paper is right, deployers can replace AUC-tuning with estimation of $\\Phi$ (or equivalently of $\\Delta_E$ on routed rows) and know the ceiling they are aiming at before training a router.","Any router that routes away on a set where it is not better than the primary, with recoveries offset by destructions, contributes zero gain, so gate quality must be measured by the direction of its information, not by advisor diversity alone.","A pre-deployment certificate can refuse to ship a router: on the tested OpenRCA pools the protocol withholds certification, and on RouterBench it certifies only when the sampling unit of evaluation matches the sampling unit of deployment.","The matching two-point lower bound means the required evaluation size $m^\\star \\approx \\sigma^2\\log(1/\\delta)/G^2$ is not an artifact of a particular bound; no other test can certify the gain with substantially fewer samples.","The independence-baseline screen $E \\le 0$ on all 221 RouterBench pools and three OpenRCA distributions predicts that many seemingly diverse advisor pools are statistically redundant, capping routing gains below what marginal accuracies suggest."],"supporting_citations":[{"why":"Supplies the RouterBench benchmark and 11-model scoring matrix on which the real-data certificate is exercised.","marker":"[10]"},{"why":"Supplies the OpenRCA Bank dataset with three Gemini advisors whose redundancy leads to the protocol's refusal.","marker":"[4]"},{"why":"Provides the Bennett inequality that underlies the one-sided concentration bracket.","marker":"[23]"},{"why":"Gives the standard Bernstein/Bennett concentration form used to define the bracket.","marker":"[24]"},{"why":"Provides the two-point lower-bound method for the matching sample-size guarantee.","marker":"[26]"},{"why":"Supplies the Bretagnolle-Huber refinement that recovers the $\\log(1/\\delta)$ dependence in the lower bound.","marker":"[27]"},{"why":"Supplies the empirical-Bernstein bound used in the robustness check on the plug-in variance.","marker":"[25]"},{"why":"Introduces the ensemble diversity measures reframed as the independence-baseline screen $E$.","marker":"[16]"}],"fun_headline_variants":["Routing gain is not AUC: RouteGuard certifies only real edges","RouteGuard: certifies routing gain only when it exists","Gain = π Δ_E: RouteGuard certifies it before you ship","Complementarity isn't enough: RouteGuard checks the real gain","RouteGuard: refuses when advisors are redundant"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The certificate treats the sample variance $\\hat\\sigma^2$ computed from the same $m$ evaluation rows as the true variance inside the concentration bound, and if that estimate runs low at finite $m$, the claimed $1-\\delta$ coverage is not guaranteed by the headline bracket alone.","fun_headline_variants_meta":{"raw":{"variants":["Routing gain is not AUC: RouteGuard certifies only real edges","RouteGuard: certifies routing gain only when it exists","Gain = π Δ_E: RouteGuard certifies it before you ship","Complementarity isn't enough: RouteGuard checks the real gain","RouteGuard: refuses when advisors are redundant"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000659,"raw_usage":{"total_tokens":3100,"prompt_tokens":1115,"completion_tokens":1985,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":731,"completion_tokens_details":{"reasoning_tokens":1900}},"tokens_in":731,"tokens_out":1985,"duration_ms":13379,"temperature":1.0,"reasoning_tokens":1900,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T00:33:12.266655+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate a true-null law on $\\{-1,0,1\\}$ with known variance, draw many $m = 312$ evaluation samples, and run the protocol's exact bracket with plug-in $\\hat\\sigma^2$ at $\\delta = 0.05$: if the false-certification rate exceeds 5% whenever the estimator understates the variance, the plug-in coverage claim fails.","supporting_citations":[{"cited_title":"OpenRCA: Can large language models locate the root cause of software failures?","cited_arxiv_id":null,"evidence_quote":"Supplies the OpenRCA Bank dataset with three Gemini advisors whose redundancy leads to the protocol's refusal."},{"cited_title":"Probability inequalities for the sum of independent random variables,","cited_arxiv_id":null,"evidence_quote":"Provides the Bennett inequality that underlies the one-sided concentration bracket."},{"cited_title":"Boucheron, G","cited_arxiv_id":null,"evidence_quote":"Gives the standard Bernstein/Bennett concentration form used to define the bracket."},{"cited_title":"Convergence of estimates under dimension- ality restrictions,","cited_arxiv_id":null,"evidence_quote":"Provides the two-point lower-bound method for the matching sample-size guarantee."},{"cited_title":"Estimation des densit ´es: risque minimax,","cited_arxiv_id":null,"evidence_quote":"Supplies the Bretagnolle-Huber refinement that recovers the $\\log(1/\\delta)$ dependence in the lower bound."},{"cited_title":"Empirical Bernstein bounds and sample variance penalization,","cited_arxiv_id":null,"evidence_quote":"Supplies the empirical-Bernstein bound used in the robustness check on the plug-in variance."},{"cited_title":"Measures of diver- sity in classifier ensembles and their relationship with the ensemble accuracy,","cited_arxiv_id":null,"evidence_quote":"Introduces the ensemble diversity measures reframed as the independence-baseline screen $E$."}],"review_version":1}