{"id":"c42c9559-7d7f-4b66-8e18-603f6e8596c9","arxiv_id":"2607.23524","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"End-to-end deep-search accuracy masks distinct failures in search decision-making and evidence synthesis; a controllable reverse-engineered benchmark exposes those failures.","lead":"The paper defines Delegation Intelligence—knowing when and how to search—and builds a reverse-engineered document pipeline plus DelegSearchBench to measure it separately from end-to-end accuracy. It shows that strong final answers can hide premature answering, ineffective search, and position-sensitive evidence use.","discovery_kind":"new_method","skeptic_critique":{"model":"moonshotai/kimi-k3","headline":"The headline behavioral finding of Setting II — two opposing delegation failure profiles across named models — rests on 33 items (~16–17 per model per mode) with N=3 samples and no reported uncertainty, so the per-model delegation profiles driving the paper's central conclusion may not be statistica","rationale":"The reader's weakest assumption targeted external validity: whether LLM-generated, closed-pool tasks transfer to open-web search. That is a real concern, but it attacks the benchmark's downstream usefulness more than the internal truth of the central claim, which is deliberately framed around controlled-setting diagnosis (\"capability boundaries masked by standard end-to-end evaluations\"). I agree the synthetic-construction and judge-dependence stack is a genuine correctness risk (all correctness labels come from a single held-out judge, deepseek-v4-flash, with no human agreement study on judged outputs), but the more load-bearing internal weakness is the statistical basis of the Setting II narrative: the paper's signature finding — opposing delegation failure profiles attributed to named models — is drawn from 33 items, ~16–17 per model-mode cell, N=3 samples, no uncertainty quantification, and a closed-pool retrieval target of ~9 documents that makes Hit@FirstSearch nearly a keyword-matching test. This concern and the reader's overlap on the small Setting II and closed-pool points (hence partial agreement), but I locate the risk in the internal support for the central claim rather than in transfer. My concrete test does not change the verdict: the reader's CONDITIONAL already prices in exactly this class of replication risk, the Setting I results (position sensitivity, mode dependence, Pass@1/Pass@3 gaps on 429 items) independently support the weaker reading of the central claim, and the honest framing of the pipeline's limitations (residual-defect audit in Appendix B, acknowledgment that Setting II is not open-world) counts in the paper's favor. CONDITIONAL with the expanded-Setting-II check as a named condition remains the right posture.","tokens_in":25878,"tokens_out":2105,"duration_ms":48352,"concrete_test":"Expand Setting II from the 33 hand-picked items to all structurally eligible multi_hop and sufficiency items in the 429-item pool (176 items per Table 6), applying the same withheld-golden construction mechanically rather than by manual curation, and recompute Table 2 with per-model 95% bootstrap CIs over items. If the named failure profiles (premature answerers vs ineffective searchers) and their ordering survive with non-overlapping CIs, the central delegation claim holds; if the profiles merge or reorder, the paper's Setting II conclusions reduce to a small-sample anecdote and the central claim should be scoped to Setting I phenomena only.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's strongest claim is that endpoint accuracy conceals distinct delegation and synthesis failures. The Setting I evidence (429 items, 6 modes, 3 position variants) is reasonably powered for the mode-dependence and position-sensitivity conclusions, and those sub-claims look internally sound. The more exposed part of the central claim is the Setting II narrative (§5.2, Table 2): the paper names specific models as exemplars of two opposing failure modes — doubao-seed-2.0-pro and deepseek-v4-pro as premature answerers (high Direct rate), hunyuan-3-preview as an aggressive-but-ineffective searcher (high Search rate, low Correct@Search), kimi-k2.6 as a high Hit@FirstSearch model, claude-sonnet-4.6 as the calibrated profile. These named behavioral profiles are what the paper's central insight (\"calibrated coordination of uncertainty awareness, evidence acquisition, and faithful synthesis\") is built on. But Setting II has only 33 items total, split across two modes, so each model's Direct/Search fraction, Hit@FirstSearch, and Correct@Search in each mode is estimated from roughly 13–20 items. Differences of 0.2–0.3 in these rates correspond to 3–5 items changing outcome; e.g., hunyuan's Correct@Search of 0.206 vs kimi's 0.933 in multi_hop is a handful of items' difference, and no confidence intervals, per-item breakdowns, or significance tests are reported. Hit@FirstSearch is additionally conditioned on the model having searched at all, shrinking the effective denominator further and making cross-model comparisons composition-dependent. A secondary concern compounds this: the search backend retrieves top-k from the item's own closed pool of ~9 documents (§5.2), so \"search competence\" here is measured against a near-trivial retrieval target — Hit@FirstSearch largely reduces to whether the model's query string lexically matches the withheld golden doc, which is much weaker evidence about \"expressing the missing information need\" than the paper's framing implies. Neithe","agreement_with_reader":"partial"},"referee_report":null,"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The useful core here is not another end-to-end agent leaderboard. It’s a clean separation of Search Decision-Making from Information Synthesis & Verification, plus a reverse-engineering pipeline that fixes answer and evidence structure before task construction. That control is what lets them run full-context position ablations and partial-context action metrics without the usual retrieval confounds.\n\nSetting I is the part I’d trust. 429 items, six shaping modes, three evidence orders, Pass@1/Pass@3 across a wide model panel. Mode-dependent rankings and lost-in-the-middle effects with retrieval removed are real and useful. The Pass@1–Pass@3 gaps make the stability point concrete rather than hand-wavy. The construction recipe (Stages A–C, evidence-anchored distractors, multi-view rejection) is more interesting as a reusable method than as a one-shot dataset, and the 100-item human audit is at least honest about residual defect rates.\n\nSetting II is thinner than the prose implies. Thirty-three curated items, split across two modes, N=3, no intervals. Naming doubao/deepseek as premature answerers, hunyuan as aggressive-but-ineffective, and claude as calibrated is the narrative the abstract leans on, but those fractions move on a handful of items, and Hit@FirstSearch is conditioned on searching at all. Worse, “search” is top-k over the item’s own ~9-doc closed pool, so query quality mostly means lexical hit on a withheld golden doc—not open-web information need expression. That doesn’t kill the premature-vs-ineffective contrast as a qualitative observation; it does mean the central “calibrated coordination” claim is oversold relative to the n and the retrieval setup.\n\nOther soft spots in proportion: code/data still “after internal review,” heavy LLM generate/filter/judge stack, and transfer to live web unproven. Citations to multi-hop, BrowseComp-Plus, GAIA, τ-bench, lost-in-the-middle are appropriate; novelty is incremental disentanglement, not a new paradigm.\n\nWho it’s for: people building or auditing search-agent benchmarks and training signals. Worth a serious referee. I’d bring the pipeline and Setting I tables to reading group; treat Table 2 model portraits as provisional. Engage the method; don’t overfit the named failure profiles until the set is larger and the search backend is less toy.","headline":"Solid disentangled eval recipe with real Setting I signal; the named Setting II “delegation profiles” are underpowered and the closed-pool search is a weak stand-in for open web.","tokens_in":27288,"tokens_out":608,"would_cite":true,"duration_ms":19106,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"End-to-end answer accuracy alone cannot diagnose whether models know when and how to search.","keywords":["Delegation Intelligence","deep search","agent evaluation","search decision-making","information synthesis","evidence verification","controllable synthesis","DelegSearchBench"],"falsifier":"Re-run the same models on live open-web deep-search tasks with matched queries, or expand Setting II to a large open corpus: if mode rankings, position effects, and premature-vs-ineffective-search profiles collapse or reverse while end-to-end accuracy stays high, the claim that these controlled gaps diagnose real Delegation Intelligence fails.","tokens_in":26937,"feed_emoji":"🔍","tokens_out":1006,"duration_ms":19933,"temperature":0.7,"pith_summary":"Deep-search agents are usually scored only on whether the final answer is right. That single number mixes retrieval luck, long-context reading, source judgment, and the decision to search at all, so failures stay hard to attribute. This paper names the missing meta-skill Delegation Intelligence: knowing when evidence is insufficient, whether and how to search, and how to verify and fuse what comes back under noise and adversarial distractors. It offers a controllable, document-grounded reverse-engineering recipe that builds tasks with known answers, labeled evidence, natural noise, and targeted distractors, then instantiates DelegSearchBench and a protocol that isolates synthesis versus search decisions by changing document composition and tool access. Across models, full-context results show mode-specific weaknesses and lost-in-the-middle sensitivity even when all evidence is present; partial-context results show premature answering on one side and aggressive but ineffective searching on the other. The sympathetic takeaway is that reliable deep search needs calibrated coordination of uncertainty awareness, acquisition, and faithful synthesis—not merely calling a search tool.","feed_headline":"Answer accuracy hides when models fail at search","feed_subtitle":"Controlled bench splits search decisions from evidence fusion and exposes opposing failure modes","key_machinery":"Delegation Intelligence, decomposed into Search Decision-Making and Information Synthesis & Verification, measured via a controllable synthesis pipeline of document-grounded reverse engineering (evidence-first query/answer construction, evidence-anchored distractors, multi-view rejection filtering) and a disentangled protocol that varies document composition and tool access (Setting I full-context, search off; Setting II partial golden evidence, search on over a closed item pool).","core_discovery":"Deep-search competence cannot be characterized by final-answer accuracy alone. Under controlled full-context and partial-context protocols on DelegSearchBench, models show mode-dependent synthesis and verification failures, position sensitivity to supporting evidence even when the full pool is available, unstable reasoning across repeated attempts (Pass@1 vs Pass@3 gaps), and two opposing delegation failures—answering too early despite missing evidence, or searching without retrieving and integrating what is needed—so reliable agents must coordinate uncertainty recognition, evidence acquisition, and verification, not only invoke search.","pith_inferences":["Position-robust evidence localization may need architectural or training fixes separate from better retrievers, since the paper’s full-context setting already removes retrieval noise.","Judge-model and construction-model separation reduces self-preference, but residual alignment between synthetic distractors and particular model families could still skew comparative rankings—worth an independent cross-generator audit.","A natural next stress test is multi-turn, budgeted search where cost of each call is explicit, to see whether “knowing when not to search” survives under real latency and quota pressure."],"forward_implications":["Agent benchmarks should report decoupled scores for search decision-making and synthesis/verification, not only task success rate.","Full-context evaluation with fixed document order variants can expose lost-in-the-middle failures independent of retrieval quality.","Pass@1 versus Pass@3 (and related stability metrics) should be standard for agents, because stochastic success is not the same as reliable evidence use.","Training and product design should target calibrated under- and over-search, not only higher tool-call rates or longer contexts.","The reverse-engineering recipe can be reused to build new controlled deep-search suites as domains and tools change."],"fun_headline_variants":["Answer accuracy masks deep-search delegation failures","Models fail search decisions even with full evidence","Bench splits search timing from evidence fusion errors","Agents answer too early or search without integrating","Final scores hide opposing deep-search failure modes"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"If tasks built by reverse-engineering real documents with model generators and filters, plus a small closed search pool in the partial-context setting, do not stand in for open-web search, then the measured “delegation” gaps may not transfer beyond the synthetic bench.","fun_headline_variants_meta":{"raw":{"variants":["Answer accuracy masks deep-search delegation failures","Models fail search decisions even with full evidence","Bench splits search timing from evidence fusion errors","Agents answer too early or search without integrating","Final scores hide opposing deep-search failure modes"]},"model":"grok-4.5","effort":"low","cost_usd":0.002513,"raw_usage":{"total_tokens":1026,"prompt_tokens":802,"num_sources_used":0,"completion_tokens":50,"cost_in_usd_ticks":25128000,"prompt_tokens_details":{"text_tokens":802,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":174,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":802,"tokens_out":50,"duration_ms":4028,"temperature":1.0,"reasoning_tokens":174,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-30T20:12:16.384398+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Re-run the same models on live open-web deep-search tasks with matched queries, or expand Setting II to a large open corpus: if mode rankings, position effects, and premature-vs-ineffective-search profiles collapse or reverse while end-to-end accuracy stays high, the claim that these controlled gaps diagnose real Delegation Intelligence fails.","supporting_citations":[],"review_version":1}