{"id":"c4e7bc49-e083-437e-97d7-71be8aabb92b","arxiv_id":"2603.14324","paper_version":5,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"high","formal_verification":"none","parameter_count":2,"one_line_summary":"Separated routing-and-advice surrogates for Learning-to-Defer are inconsistent; a composite expert–advice surrogate is H-consistent and recovers the Bayes-optimal policy in the limit.","lead":"The paper claims that when Learning-to-Defer also chooses advice for the selected expert, common two-head surrogates are inconsistent, and a joint expert–advice surrogate restores Bayes optimality. This matters for systems that retrieve documents, call tools, or escalate only after routing.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"Manuscript mismatch: provided full text is CSI compression (2603.14325), not L2D-with-advice (2603.14324), so the central consistency claims cannot be audited.","rationale":"The reader correctly identified that the full text does not match the abstract/title under review. That is the load-bearing blocker: every formal claim (inconsistency of separated heads, H-consistency of the augmented surrogate, excess-risk transfer, Bayes recovery) lives in proofs and experiments that are simply not present. No internal mathematical soft spot can be isolated because the relevant sections are absent. The appropriate action is not to invent a critique of a CSI paper as if it were L2D, nor to upgrade the verdict on abstract claims alone. Keep UNVERDICTED / LOW confidence until the correct manuscript is supplied; the concrete test is exactly that substitution and re-audit.","tokens_in":21447,"tokens_out":462,"duration_ms":5019,"concrete_test":"Replace the CACHEABLE prefix with the actual PDF/source of arXiv:2603.14324 (or the authors’ L2D-with-advice manuscript). Re-run the audit on the composite-action surrogate theorem and the synthetic separated-surrogate failure experiment; only then can consistency and excess-risk transfer be accepted or rejected.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that separated routing/advice surrogates are inconsistent even in the smallest non-trivial setting, while an augmented surrogate on the composite expert–advice action space is H-consistent with an excess-risk transfer bound that recovers the Bayes-optimal policy. The CACHEABLE full text is a different paper (Fundamental Limits of CSI Compression / GMTC, self-labeled arXiv:2603.14325). No definitions of the composite action space, no statement of the separated-surrogate family, no H-consistency theorem, no excess-risk transfer proof, and no L2D experiments appear. Without those objects, the strongest claim is not checkable from the supplied manuscript; the reader’s abstract-only UNVERDICTED stance is forced by a document identity failure, not by a soft spot inside a correct L2D argument.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The manuscript (as supplied under the title Learning-to-Defer with Expert-Conditional Advice) claims that standard Learning-to-Defer is incomplete when experts can also receive selectable advice (documents, tools, escalation context). It asserts that a broad family of natural separated surrogates with distinct routing and advice heads is inconsistent even in the smallest non-trivial setting, and that an augmented surrogate on the composite expert–advice action space admits an H-consistency guarantee and excess-risk transfer bound that recovers the Bayes-optimal policy in the limit. Experiments on tabular, language, and multi-modal tasks are claimed to improve over standard L2D and to adapt advice acquisition to the cost regime, with a synthetic check of the separated-surrogate failure mode. The body text actually provided, however, is a different paper on Gaussian-mixture transform coding (GMTC) for CSI compression in FDD massive MIMO (self-identified as arXiv:2603.14325), with RD bounds, reverse-waterfilling, and COST2100 experiments.","tokens_in":21678,"tokens_out":991,"duration_ms":8253,"significance":"If the L2D-with-advice claims were substantiated, they would be a useful extension of the surrogate-consistency literature to joint routing-and-information-acquisition decisions that arise in modern retrieval- and tool-augmented systems. The abstract’s structure (inconsistency of separated heads; H-consistency of a joint surrogate; excess-risk transfer) is the standard and valuable pattern in that literature. The supplied body, by contrast, is a coherent CSI-compression paper with a clean multi-modal reverse-waterfilling result and strong empirical RD gains, but that is not the paper under review. Because the two documents do not match, the significance of the claimed L2D contribution cannot be assessed from the materials provided.","major_comments":[{"comment":"Document identity failure: the abstract and title describe Learning-to-Defer with Expert-Conditional Advice (stat.ML, 2603.14324), but the full manuscript text is “Fundamental Limits of CSI Compression in FDD Massive MIMO” (self-labeled arXiv:2603.14325). No definition of the composite expert–advice action space, no statement of the separated-surrogate family, no H-consistency theorem, no excess-risk transfer bound, and no L2D experiments appear. The central claims are therefore not checkable from the supplied manuscript.","section":null},{"comment":"Because the body never states the surrogate losses, hypothesis class H, or cost structure for advice, the abstract’s claim that “a broad family of natural separated surrogates is inconsistent even in the smallest non-trivial setting” cannot be verified or refuted. A referee cannot assess whether the inconsistency is load-bearing or an artifact of a particular reduction.","section":null},{"comment":"The claimed H-consistency guarantee and excess-risk transfer bound that “yield recovery of the Bayes-optimal policy in the limit” are not present in the provided text. Without the theorem statements, assumptions on H, and proof sketches, the recovery claim cannot be audited.","section":null},{"comment":"Experimental claims (tabular/language/multi-modal gains; synthetic confirmation of the separated-surrogate failure mode; cost-regime adaptation of advice acquisition) have no corresponding figures, tables, or protocols in the supplied manuscript, which instead reports COST2100 CSI NMSE and FLOPs comparisons. The empirical support for the L2D contribution is therefore absent.","section":null}],"minor_comments":[{"comment":"The supplied CSI manuscript itself has presentation issues (e.g., incomplete citation “Proof. The formal proof is provided in [?], [36]” in Lemma 1; dense Fig. 1 encoding; heavy use of special characters that render poorly), but these are irrelevant to the L2D paper under review.","section":null},{"comment":"If the correct L2D manuscript is resubmitted, the authors should ensure that the abstract, arXiv id, and full text refer to the same work and that all theorem numbers and experiment sections are present.","section":null}],"recommendation":"reject","confidential_remarks":"This appears to be a hard manuscript-mismatch / wrong-PDF error rather than a scientific disagreement with an L2D argument. The CSI paper (2603.14325) looks potentially reviewable on its own merits in an information-theory or wireless venue, but it is not the paper announced by the title and abstract. I recommend desk rejection or return for correct manuscript upload; I would be willing to re-review the actual L2D paper if it is provided."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The one thing you need to know: the abstract is for Learning-to-Defer with expert-conditional advice (2603.14324), but the full manuscript we were given is Fundamental Limits of CSI Compression / GMTC (self-labeled 2603.14325). So I cannot verify the inconsistency proof, the H-consistency theorem, the excess-risk transfer bound, or the experiments. That is a document identity failure, not a soft spot inside a correct L2D argument.\n\nFrom the abstract alone, the contribution is clear and non-trivial if true. Standard L2D assumes each expert’s information is fixed when you route. Here, after you pick an expert you also choose what advice they get (retrieval, tools, escalation). The paper claims a broad family of natural separated surrogates—routing head plus advice head—is inconsistent even in the smallest non-trivial setting, and that an augmented surrogate on the composite expert–advice action space is H-consistent with an excess-risk transfer bound that recovers the Bayes policy in the limit. Experiments are said to beat standard L2D and to adapt advice spend to cost, with a synthetic check of the separated-surrogate failure mode. That is a real problem formulation for hybrid decision systems and tool-use pipelines, not a cosmetic rename.\n\nWhat I cannot do is confirm any of it. No definitions of the composite action space, no statement of the separated family, no theorem statements, no proofs, no tables. The CSI paper in the cache is a different, self-contained RD/transform-coding story; it does not substitute. Free parameters (advice costs, hypothesis class H) and the modeling reduction to a finite composite action space are exactly the places a referee would pressure—and we have no text to pressure.\n\nWho it is for: people who train deferral / routing with post-routing information acquisition. Value is high if the negative result on separated heads is clean and the joint surrogate is practical. Right now I would not bring it to reading group or cite it until we have the correct PDF. A serious editor should still send the real paper to peer review on the strength of the abstract’s claims; the topic and the consistency angle deserve referee time. My recommendation: hold judgment, pull 2603.14324, then re-read. Do not treat the CSI manuscript as evidence for or against the L2D results.","headline":"We cannot audit the L2D-with-advice claims: the supplied full text is a different paper (CSI/GMTC, 2603.14325), so the consistency theorems are not checkable from what we have.","tokens_in":22306,"tokens_out":599,"would_cite":false,"duration_ms":12451,"reading_group":"no","serious_thinker":"unclear","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Separated heads for routing and advice are inconsistent; a joint expert–advice surrogate recovers the Bayes-optimal deferral policy.","keywords":["learning to defer","expert advice","surrogate consistency","H-consistency","excess risk","Bayes-optimal policy","composite action space"],"falsifier":"On the paper’s synthetic benchmark, train a separated two-head surrogate in the smallest nontrivial expert–advice setting; if it still converges to the Bayes-optimal joint policy (rather than the predicted inconsistent limit), the inconsistency claim fails. On the real tasks, if the joint method does not improve over standard Learning-to-Defer or fails to adapt advice acquisition to cost, the practical claim fails.","tokens_in":22325,"feed_emoji":"🔀","tokens_out":830,"duration_ms":11231,"temperature":0.7,"pith_summary":"Standard Learning-to-Defer picks which expert should handle each input, but treats the information each expert sees as fixed. Many real systems also choose what extra context that expert gets—retrieved documents, tool outputs, or escalation notes—after the routing decision. The paper shows that a natural family of surrogates that train routing and advice with separate heads is inconsistent even in the smallest nontrivial case, so they need not recover the optimal joint policy. It then defines an augmented surrogate over the combined expert–advice action space, proves an H-consistency guarantee and an excess-risk transfer bound, and shows that the Bayes-optimal policy is recovered in the limit. Empirically the method beats ordinary Learning-to-Defer on tabular, language, and multi-modal tasks and adjusts how much advice it buys to the cost regime; a synthetic check reproduces the predicted failure of separated surrogates.","feed_headline":"Separate routing and advice heads fail; joint surrogate works","feed_subtitle":"An augmented expert–advice surrogate is consistent and recovers the Bayes-optimal deferral policy","key_machinery":"The augmented surrogate on the composite expert–advice action space: it replaces separate routing and advice heads with a joint surrogate whose H-consistency and excess-risk transfer bound force recovery of the Bayes-optimal joint policy.","core_discovery":"A broad class of separated surrogates that learn routing and advice with distinct heads is inconsistent even in the smallest nontrivial setting. An augmented surrogate that treats the composite expert–advice pair as the action recovers Bayes optimality in the limit via an H-consistency guarantee and an excess-risk transfer bound.","pith_inferences":["The same composite-action construction may apply to other two-stage decisions (who acts, then what context they receive) beyond classical Learning-to-Defer.","Inconsistency of separated heads suggests that multi-head architectures for sequential decisions need joint consistency proofs, not only separate calibration of each head.","If advice is continuous or combinatorial (large document sets), the finite composite space may need approximation theory the paper leaves open."],"forward_implications":["Routing and advice should be optimized jointly; separate heads are not a safe default even for simple instances.","Systems that retrieve documents, call tools, or escalate can treat those choices as part of the deferral action and still target Bayes optimality.","As the joint surrogate is driven to zero excess risk, the induced policy approaches the cost-optimal expert–advice map.","Advice acquisition can be made cost-aware: the method spends more on context only when the cost regime justifies it."],"fun_headline_variants":["Separated routing-advice heads fail even in simplest cases","Joint expert-advice surrogate restores Bayes-optimal deferral","Augmented composite-action loss yields H-consistency","Distinct heads inconsistent; treat expert-advice as one action","Advice-aware deferral needs joint surrogate, not separated ones"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"That treating practical advice (documents, tools, escalation context) as a finite composite action space, and that the H-consistency conditions hold for the hypothesis classes and costs used in the experiments, is enough for the consistency theory to carry over from the abstract setting.","fun_headline_variants_meta":{"raw":{"variants":["Separated routing-advice heads fail even in simplest cases","Joint expert-advice surrogate restores Bayes-optimal deferral","Augmented composite-action loss yields H-consistency","Distinct heads inconsistent; treat expert-advice as one action","Advice-aware deferral needs joint surrogate, not separated ones"]},"model":"grok-4.5","effort":"low","cost_usd":0.005444,"raw_usage":{"total_tokens":1382,"prompt_tokens":712,"num_sources_used":0,"completion_tokens":82,"cost_in_usd_ticks":54440000,"prompt_tokens_details":{"text_tokens":712,"audio_tokens":0,"image_tokens":0,"cached_tokens":0},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":588,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":712,"tokens_out":82,"duration_ms":4702,"temperature":1.0,"reasoning_tokens":588,"cache_read_input_tokens":0,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-14T21:22:05.293106+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"On the paper’s synthetic benchmark, train a separated two-head surrogate in the smallest nontrivial expert–advice setting; if it still converges to the Bayes-optimal joint policy (rather than the predicted inconsistent limit), the inconsistency claim fails. On the real tasks, if the joint method does not improve over standard Learning-to-Defer or fails to adapt advice acquisition to cost, the practical claim fails.","supporting_citations":[],"review_version":2}