{"id":"98ea0b8f-9a36-42f3-9a73-7c6ea017dd1e","arxiv_id":"2607.26052","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"CARE replaces fixed top-k expert selection with a confidence-adaptive nucleus rule, improving accuracy at matched compute and OOD detection in MoE-LoRA models.","lead":"A new routing rule for mixture-of-experts LoRA adapters decides per token how many experts to activate, spending more compute on tokens the router is unsure about. The method claims better accuracy at the same average compute and a free out-of-distribution signal from the same forward pass.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Budget-optimality proof's final equivalence is unsupported: Eq. (3)'s single threshold on cumulative router mass need not implement greedy marginal-gain allocation unless a global monotone relationship between marginal gain and C_k holds; Fig. 4(c) never tests it.","rationale":"The reader's weakest assumption is precisely the monotonicity premise behind Props. 2 and 3. I agree that this is the most load-bearing weakness: it is the hinge between the empirical observation that concentrating router mass tracks correctness and the advertised claims of a Bayes-optimal confidence score and a budget-optimal thermostat. My stress-test sharpens that concern by pointing out that Prop. 3's proof contains an unproved equivalence step, not merely an untested empirical premise. Even granting concavity, a single threshold on cumulative mass does not follow from greedy marginal-gain optimality unless there is a token-independent mapping between marginal gain and C_k. The manuscript's own evidence (Fig. 4a, 4c) does not address this mapping. I do not regard this as a fatal flaw in the empirical method: CARE could still improve accuracy at matched compute on the evaluated benchmarks even if Prop. 3 is false, and the paper's single-forward-pass, parameter-free design is a genuine contribution. However, the theoretical claims are over-stated, and the empirical centerpiece (Tables 1-3, Fig. 3) is not independently verifiable without error bars and code. The conditional-accept verdict is therefore appropriate: the authors should either prove the missing equivalence under stated assumptions, or empirically validate the cross-token monotonicity and clearly relabel Prop. 3 as a heuristic justification. No change to the reader's verdict is needed.","tokens_in":15380,"tokens_out":13401,"duration_ms":194790,"concrete_test":"On a large held-out labeled split of the same MoE-LoRA backbone, compute for each token j and each k the model's prediction correctness under fixed top-k routing and under fixed top-(k+1) routing, using the same trained router and experts. Define marginal gain Δg_j(k) as the indicator difference. Bin all (j, k) pairs by cumulative router mass C_k(p_j) and check whether the binned mean marginal gain is monotone in C_k across tokens (e.g., by a rank-correlation or isotonic-regression fit). If it is not monotone, Prop. 3's premise fails and the thermostat's optimality claim is unsupported. As a second, decision-relevant check, run an oracle greedy allocation that spends the budget B on the tokens with the largest true marginal gains and compare its accuracy to CARE's at the same B; if the oracle beats CARE by more than CARE beats fixed top-k, the threshold rule is not budget-optimal even if","verdict_should_be":"UNCHANGED","load_bearing_attack":"Proposition 3 is the paper's advertised budget-optimality guarantee, and it is the piece that turns the thermostat from a curve-fitting device into a principled allocation rule. The proof correctly states that, for concave per-token accuracy curves g_j(k), greedy allocation by largest marginal gain is optimal under a cardinality budget. But the final step—'if marginal gain is monotone in covered mass, then the threshold on marginal gain corresponds to a single threshold τ on cumulative mass'—is asserted, not proved. Even if each token's marginal gain is decreasing in its own k (and hence in its own C_k), a single global τ on C_k does not equalize marginal gains across tokens. A peaked token can stop at k=1 with C_1=0.9, while a flat token can stop at k=5 with C_5 just above τ; the marginal gains at these stopping points can be wildly different. To make the threshold rule optimal you need a stronger, token-independent relation: the marginal gain of the last admitted expert must be approximately the same function of C_k for all tokens. The paper provides no direct evidence for this. Figure 4(a) only shows that correct predictions have higher top-1 mass; Figure 4(c) plots marginal gain against expert disagreement, not against C_k. Thus the central theoretical support for CARE's allocation is conditional on an untested cross-token monotonicity assumption. If the assumption fails on a target distribution, the thermostat is no longer budget-optimal, and the accuracy-compute advantage rests entirely on the empirical tables—which lack error bars and a reproducible artifact. This is load-bearing because the paper explicitly advertises 'budget optimality' as a contribution, and the method's practical appeal is the claim that one calibrated threshold is not merely convenient but optimal.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes CARE, a test-time routing rule for MoE-LoRA that replaces fixed top-k expert selection with per-token nucleus admission: experts are added in decreasing router weight until cumulative mass exceeds a threshold tau, with a disagreement-based extension, and tau is calibrated so that the average number of active experts matches a target budget. The same router signals (concentration and expert disagreement) are used as an uncertainty score for OOD detection and abstention. The paper claims that CARE improves accuracy over fixed top-k at matched compute (e.g., +0.5 on LLaMA commonsense and +0.9 on math/code/knowledge), matches fixed-k=4 accuracy with 12% fewer experts, and improves OOD-AUROC from 0.640 to 0.668 in a single forward pass. Theoretical support is provided via four propositions: nucleus fidelity, confidence ranking, budget optimality, and epistemic disagreement.","tokens_in":15823,"tokens_out":4517,"duration_ms":64320,"significance":"The contribution is potentially significant: if the empirical claims hold, CARE is a zero-retraining, single-pass improvement applicable to any MoE-LoRA model, and it yields a useful uncertainty signal as a by-product. The router-as-uncertainty idea is simple and well motivated, and the experimental suite spans two backbones and four task families. The reported gains are small but consistent in direction. However, the strength of the claims is undercut by an incomplete budget-optimality proof and by a lack of statistical detail in the empirical comparisons; these are fixable, so the paper merits revision rather than rejection.","major_comments":[{"comment":"The proof of budget optimality does not establish that the nucleus threshold rule of Eq. (3) implements the optimal allocation. The greedy marginal-gain argument for a separable concave objective is standard, but the final step asserts that a threshold on marginal gain corresponds to a single threshold on cumulative mass C_k. This requires a cross-token homogeneity condition: the marginal gain of the last admitted expert must be a function of C_k that is comparable across tokens. The manuscript states this only as a modeling assumption and provides no direct evidence; Fig. 4(c) plots marginal gain against expert disagreement, not against C_k. Without a precise sufficient condition or a direct empirical test, the advertised 'budget optimality' is not supported. Please either prove a concrete condition under which the equivalence holds, add a test of monotonicity in C_k, or weaken the clai","section":"A.3 / Prop. 3"},{"comment":"All headline numbers are three-seed means with no confidence intervals, error bars, or per-seed results. The key differences are small (+0.5 on LLaMA commonsense, +0.9 on math/code/knowledge) and the text describes the top fixed-k baselines as a 'tight band.' Without variance estimates it is impossible to judge whether the improvements are statistically distinguishable from noise. Since the central empirical claim is that CARE improves accuracy at matched compute, this is load-bearing. Please report per-seed values or error bars, and state whether the differences are significant.","section":"Tables 1-4 / App. B"},{"comment":"The OOD split is described only as 'a harder-and-flatter distribution shift of the same domain,' which is too vague to reproduce or to rule out selection bias. Moreover, the 'MC-dropout (proxy)' and 'Deep ensemble (proxy)' rows report AUROC 0.636/0.638, below the simple MSP baseline of 0.640. This is anomalous and suggests the proxy implementations are not faithful to the methods they approximate, undermining the claim that CARE beats multi-pass uncertainty methods. Specify exactly how the OOD split is constructed and how the proxies are computed, or remove the comparison to multi-pass methods.","section":"§6.5 / Table 3"}],"minor_comments":[{"comment":"CARE is described as 'no extra parameters' and 'parameter-free,' but Algorithm 1 depends on hyperparameters gamma, delta, w, kmin, kmax, and the calibrated tau*. These are not learned parameters, but the phrasing is overstated and could mislead readers about tuning burden.","section":"Abstract / §4"},{"comment":"The x-axis label 'token difficulty' is not defined. It should be made explicit how difficulty is measured (e.g., empirical error rate, router entropy, or something else).","section":"Figure 4(b)"},{"comment":"The column header 'SV AMP' should be 'SVAMP'.","section":"Table 2"},{"comment":"The notation for the sequence-level uncertainty score is ambiguous: H(p) and D are introduced as per-token quantities, but the equation appears to define a sequence-level average. Clarify the averaging.","section":"Eq. (6)"}],"recommendation":"major_revision","confidential_remarks":"The paper has a promising core idea and a substantial experimental suite, but the budget-optimality proof is incomplete and the empirical reporting lacks error bars and OOD details. The anomalous proxy baselines in Table 3 should be investigated. The paper's own Limitations section does not acknowledge the proof gap; it should be updated. I recommend major revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a solid, modest contribution. The core idea — admit experts per token by nucleus-style cumulative router mass, with a budget thermostat and a disagreement extension — is genuinely new to the MoE-LoRA literature, where fixed top-k is universal. The evaluation is broad: two backbones, four task families, matched compute, ablations, an accuracy-compute frontier, and an OOD-detection read-out. Gains are small (+0.5 points, 12% expert savings) but consistent across benchmarks, and the single-pass uncertainty signal is a useful by-product. The component ablation is clean: the confidence/nucleus term drives accuracy, the disagreement term drives OOD-AUROC, and they don't hurt each other.\n\nNow the soft spots, in proportion. The theory is the weakest part. Prop. 2 is explicitly conditional on correctness being monotone in p(1), which Figure 4(a) supports only as a pooled density plot. Prop. 3's advertised budget optimality has a real gap: the greedy marginal-gain argument is standard, but the final equivalence — threshold on marginal gain equals threshold on cumulative mass C_k — needs a cross-token monotonicity condition that is asserted, not proved. The stress-test note is right that Fig. 4(c) plots marginal gain against disagreement, not against C_k, so the central theoretical support for the thermostat is conditional on an untested assumption. This doesn't sink the empirical claim — CARE could work fine even with the proof gaps — but the paper should stop advertising 'budget optimality' as an unconditional guarantee.\n\nThe empirics have fixable but annoying gaps: three-seed means with no error bars, an underspecified OOD split, and 'proxy' MC-dropout/deep-ensemble baselines that behave anomalously (0.636/0.638 vs. MSP 0.640). The code release is announced but there's no artifact, hash, or reproduction checklist. All of this is addressable in revision. Also worth noting: the paper's own limitations section is honest about degenerate routers and batching issues, which I appreciate.\n\nBottom line: this deserves a serious referee. The mechanism is cheap, drop-in, and plausibly useful for anyone deploying MoE-LoRA; the empirical comparisons target the right baselines; the gaps are tightening, not fatal. I'd accept for peer review, ask for qualified theory and reproducibility details, and would cite this if the artifact checks out.","headline":"Solid, modest contribution: adaptive expert-count routing via router-mass thresholding is genuinely new to MoE-LoRA and broadly evaluated, but the optimality proofs rely on unproven monotonicity and the empirics need error bars and a real artifact.","tokens_in":16345,"tokens_out":1304,"would_cite":true,"duration_ms":21283,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Replacing fixed top-k routing in MoE-LoRA with a confidence-adaptive nucleus rule improves accuracy at matched compute and yields a free uncertainty signal.","keywords":["MoE-LoRA","adaptive routing","uncertainty estimation","nucleus sampling","out-of-distribution detection","conditional computation","low-rank adaptation","mixture of experts"],"falsifier":"Measure the empirical correctness rate of tokens as a function of top-1 router mass p(1) on a given MoE-LoRA checkpoint. If the curve is non-monotonic (e.g., tokens with p(1) near 0.9 are less accurate than those near 0.6), then the confidence ranking guarantee (Prop. 2) fails and CARE's allocation may not be optimal. Similarly, plot marginal accuracy gain from the (k+1)-th expert against cumulative mass C_k; if the gain is not non-decreasing in C_k, the budget thermostat may not implement the claimed optimal allocation.","tokens_in":15267,"feed_emoji":"🎯","tokens_out":5527,"duration_ms":62403,"temperature":0.7,"pith_summary":"MoE-LoRA adapters route every token to a fixed number of experts, wasting compute on easy tokens and starving hard ones. CARE exploits the router's own softmax distribution as a per-token confidence signal: peaked mass means confident, flat means ambiguous. It admits experts in nucleus fashion until cumulative router mass reaches a budget-calibrated threshold, with an extension when admitted experts disagree. Across eight commonsense benchmarks and math, code, and knowledge tasks on two backbones, CARE improves accuracy over fixed top-k at matched average compute and matches fixed-k=4 accuracy while activating 12% fewer experts. The same signals give single-pass OOD detection that beats MSP, entropy, and multi-pass proxies.","feed_headline":"Adaptive expert routing beats fixed top-k at matched compute","feed_subtitle":"CARE's confidence-based admission saves 12% experts and improves OOD detection in one pass.","key_machinery":"The carrying mechanism is the nucleus admission rule: sort router weights, accumulate them, and stop at the first k whose cumulative mass C_k reaches threshold τ; optionally extend by up to γ experts scaled by disagreement D(h;S) beyond δ. The budget thermostat (Eq. 5) calibrates τ by bisection on a held-out set to match an average budget B, so the average number of active experts is controlled. Disagreement D is the weighted per-coordinate variance of admitted expert outputs normalized by mean squared magnitude. Proposition 3 connects the thermostat to an optimal budget allocation under concavity and monotonicity of marginal gains.","core_discovery":"The central claim is that the router distribution in a trained MoE-LoRA is already a per-token uncertainty signal, and that using it to adapt the number of active experts—admitting experts in decreasing router weight until cumulative mass reaches a threshold, plus a small extension when admitted experts disagree—strictly improves the accuracy–compute frontier over fixed top-k routing. CARE is a drop-in, parameter-free, single-forward-pass rule that replaces only the gate. At matched average compute it improves accuracy (e.g., +0.5 on LLaMA commonsense), and at matched accuracy it needs 3.5 experts on average versus 4.0 for fixed-k. The same concentration and disagreement signals, blended as","pith_inferences":["The optimality guarantees (Props 2 and 3) rely on monotonicity of correctness probability in top-1 mass and monotonicity of marginal gain in cumulative mass; these are only partially validated. If they fail on a target distribution, CARE's allocation may not be budget-optimal, though the empirical gains could persist.","CARE's benefit should scale with per-token difficulty heterogeneity; on homogeneous tasks it converges to fixed-k. A practitioner can pre-screen by measuring the variance of routing entropy across a small sample.","The disagreement signal is defined over classification-style outputs; extending it to free-form generation (e.g., via semantic equivalence) is a natural next step that the paper leaves open.","Per-layer thermostats, mentioned as an extension, could improve allocation further; a testable hypothesis is that layerwise calibration with per-layer budgets yields additional accuracy gains at the same total compute."],"forward_implications":["Any trained MoE-LoRA checkpoint can be wrapped with CARE at inference time by replacing the gate; no retraining, no extra parameters, no extra forward passes.","The accuracy–compute frontier shifts upward: at any average budget CARE is more accurate than fixed top-k, and it reaches the fixed-k=4 accuracy with 12% fewer active experts.","CARE's confidence and disagreement signals yield a single-pass OOD detection score (AUROC 0.668) that outperforms max-softmax, entropy, and multi-pass proxies like MC-dropout and deep ensembles.","Under distribution shift, CARE retains accuracy better than fixed top-k at the same average compute (53.1% vs 50.3%), because shifted inputs flatten the router and trigger more expert admissions.","The thermostat provides a single knob τ that trades compute for accuracy along the frontier, letting deployments select an operating point."],"fun_headline_variants":["Router confidence sets expert count","Let model uncertainty pick experts","Adaptive expert budget from router entropy","CARE: cut experts when model is sure"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The central premise is that a token's chance of being answered correctly increases as the router's top-1 mass increases, and that each additional expert helps in proportion to the router mass it covers; if either monotonicity fails, CARE's optimality guarantees and potentially its accuracy gains weaken.","fun_headline_variants_meta":{"raw":{"variants":["Router confidence sets expert count","Let model uncertainty pick experts","Adaptive expert budget from router entropy","CARE: cut experts when model is sure"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000461,"raw_usage":{"total_tokens":2166,"prompt_tokens":790,"completion_tokens":1376,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":534,"completion_tokens_details":{"reasoning_tokens":1327}},"tokens_in":534,"tokens_out":1376,"duration_ms":14278,"temperature":1.0,"reasoning_tokens":1327,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T03:15:49.684025+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Measure the empirical correctness rate of tokens as a function of top-1 router mass p(1) on a given MoE-LoRA checkpoint. If the curve is non-monotonic (e.g., tokens with p(1) near 0.9 are less accurate than those near 0.6), then the confidence ranking guarantee (Prop. 2) fails and CARE's allocation may not be optimal. Similarly, plot marginal accuracy gain from the (k+1)-th expert against cumulative mass C_k; if the gain is not non-decreasing in C_k, the budget thermostat may not implement the claimed optimal allocation.","supporting_citations":[],"review_version":2}