{"id":"5f98efdb-8971-4782-9e66-ad310b6df675","arxiv_id":"2607.10198","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.5,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Equal semantic accuracy across Brave, Tavily, and Firecrawl (25–26/100) hides sharply different pre-fetch support, rank-1 concentration, contradiction ratios, and agent exploration regimes under one frozen agent.","lead":"Three commercial search APIs give nearly identical answer accuracy on hard questions, yet expose agents to very different snippet evidence, rank concentration, contradictions, and fetch behavior. Provider choice is therefore a retrieval-budget and policy decision, not just a recall decision.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified beyond the paper's own stated single-agent/observational boundary.","rationale":"The paper's strongest claim is not that Brave/Tavily/Firecrawl are universally different products, but that equal final accuracy can mask different pre-fetch evidence economies for the same progressive-disclosure agent. That claim is supported by the controlled isolation (only search provider varies), the per-URL oracle (6,869 valid judgments), the support split and decision partition, length-normalized surface checks (Table 4), and paired-bootstrap intervals that separate accuracy parity from structural differences (Table 5). The main soft spot is exactly what the reader and §7 already flag: observational single-agent/judge analysis with n=100 and live APIs. That soft spot limits how far the regimes can be generalized, which is why CONDITIONAL is appropriate, but it does not falsify the reported contrast under the frozen protocol. I therefore leave the reader's CONDITIONAL verdict and medium correctness-risk assessment unchanged; the concrete multi-agent re-run is the cleanest next check rather than a reason to reject or re-verdict the present manuscript.","tokens_in":16820,"tokens_out":536,"duration_ms":7244,"concrete_test":"Re-run the identical 100-query sample with one alternate frozen agent (e.g., Claude or Gemini) under the same tools, prompt contract, and Kimi oracle; if Brave's pre-fetch-support advantage and Tavily's rank-1 concentration both reverse or collapse while correctness remains near-parity, the policy-transfer claim weakens; if the directional regimes persist, the scoped claim is reinforced.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is carefully scoped: under one frozen GPT-5.4 agent, shared fetch backend, and audited semantic_match, accuracy is near-parity (25/25/26) while pre-fetch support, rank-1 concentration, contradiction ratio, and decision-partition mass differ (Tables 2–5, Fig. 3). The paper already labels these as provider-associated regimes under a fixed policy, not universal rankings (§1, §6–7), and reports bootstrap intervals that keep correctness differences including zero while structural contrasts stay positive. The reader's weakest_assumption (single agent/judge transfer) is real but is already the paper's explicit limitation rather than a hidden load-bearing flaw that overturns the scoped claim. No internal inconsistency, metric circularity, or unacknowledged confound appears strong enough to reverse the evidence-economy contrast under the stated protocol.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper argues that commercial search APIs for tool-using agents should be evaluated as decision surfaces—the ranked URLs, titles, snippets, and metadata that shape whether an agent answers, re-searches, or spends tokens on page fetches—rather than solely by final answer accuracy. Under a frozen GPT-5.4 agent, shared tools and Jina fetch backend, and 100 SEALQA-HARD questions, Brave, Tavily, and Firecrawl yield near-parity semantic correctness (25, 25, 26/100) while differing sharply in pre-fetch gold support (30 vs 16 vs 16), rank-1 concentration among gold-supporting pre-fetch rows (13% vs 50% vs 13%), surface contradiction-to-gold ratio rc:g (0.92–2.59), and decision-partition mass (SMART/MISSED/BLIND/NO-OP). A Kimi-K2.6 per-URL oracle (6,869 valid judgments; 94% human agreement on clear cases) separates pre-fetch from post-fetch support and grounds the structural claims.","tokens_in":17042,"tokens_out":1125,"duration_ms":11383,"significance":"If the result holds under the stated protocol, it is a useful and timely contribution for agent systems research: it shows that interchangeable-looking answer accuracy can hide different retrieval-budget, contamination, and policy regimes, and it supplies concrete, trace-computable diagnostics (pre-fetch support, rc:g, decision partition) that practitioners can use without treating providers as a pure leaderboard. Strengths include a carefully controlled freeze of model/prompt/tools/fetch backend, an audited semantic_match answer label, a large per-URL oracle with human validation, paired-bootstrap intervals that keep correctness differences including zero while structural contrasts stay positive, and a released reproducibility package (code, configs, sampling, evaluation scripts, aggregates) with clear non-redistribution boundaries for raw provider content. The single-agent/observational scope is explicit and does not overclaim universality.","major_comments":[{"comment":"§5.1–5.3 and §7: The central claim is scoped to one frozen GPT-5.4 policy and one Kimi-K2.6 judge, and the paper correctly labels regimes as provider-associated rather than causal. That scope is load-bearing for transfer: SMART/MISSED/BLIND/NO-OP mass and fetch appetite could shift under a different answer model or fetch policy. The manuscript should either (a) add a small second-agent or policy-sensitivity check on a subset of queries, or (b) strengthen the abstract/conclusion wording so that “provider choice is a retrieval-budget and policy decision” is always read as under a fixed agent policy, not as a provider ranking independent of policy. Without one of these, the practical recommendation in §6 is slightly stronger than the evidence boundary.","section":null},{"comment":"Table 4 and §7 (snippet-surface asymmetry): Brave’s pre-fetch support advantage is partly volume-driven (more provider-native extra snippets). The length-normalized rows help, but the paper still treats Brave’s richer surface as a product difference while aggregating extra snippets into the generic snippet channel. For the decision-surface claim, this is acceptable only if the main text more clearly separates “more pre-fetch text” from “higher per-token evidence density” when stating Brave’s gold-answer-rich-snippet regime; otherwise readers may over-attribute the 30 vs 16 support gap to ranking quality alone.","section":null}],"minor_comments":[{"comment":"Table 3 rank-1 row: the denominator is gold-supporting pre-fetch rows (101/34/30), not queries; a one-line reminder in the caption would prevent misreading the 13%/50%/13% figures as query-level rates.","section":null},{"comment":"Figure 3 cell counts are small (e.g., SMART 3/3); Appendix F Wilson intervals are good, but the main-text discussion of SMART rarity should briefly flag small-cell uncertainty.","section":null},{"comment":"§3 run window (2026-05-17) and live-API non-bit-identical reruns are stated; a short note near Table 2 that absolute counts are time-stamped snapshots would help practitioners interpreting the numbers.","section":null},{"comment":"Related Work could more explicitly contrast decision-surface metrics with classical NDCG/BEIR and with over-searching cost metrics, so the complementary evaluation target is sharper for IR readers.","section":null},{"comment":"Minor consistency: abstract says “semantic match” while body uses semantic_match; unify the label name.","section":null}],"recommendation":"minor_revision","confidential_remarks":"Fit is solid for a systems/evaluation venue in agentic IR or tool-using LLMs. The contribution is diagnostic rather than a new model; that is fine if the journal values measurement papers. I would not treat the single-agent limitation as grounds for reject—the paper already owns it—but I would insist the abstract and §6 stay tightly scoped so the piece is not misread as a provider leaderboard."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The punchline is simple: with one frozen GPT-5.4 agent, shared Jina fetch backend, and audited semantic_match, Brave/Tavily/Firecrawl land at 25/25/26 correct on 100 SEALQA-HARD questions, yet the pre-fetch surfaces are not interchangeable. Brave shows more pre-fetch gold support, Tavily packs support at rank 1, Firecrawl is associated with more blind exploration, and the surface contradiction-to-gold ratio runs from 0.92 to 2.59. That is the paper’s real contribution.\n\nWhat is new is not “agents use search.” It is the decision-surface framing plus a usable measurement stack: per-URL oracle over everything the agent saw, pre-fetch vs post-fetch support, the SMART/MISSED/BLIND/NO-OP join to observed fetches, and rc:g. The protocol is careful. Model, prompt, tools, iteration budget, and page backend stay fixed; only the search provider changes. They report 6,869 valid judgments, a 180-case human audit at 94% agreement on clear cases, semantic answer audit, and paired bootstraps that correctly keep accuracy differences including zero while structural contrasts stay positive. The non-leaderboard framing is honest and matches the evidence.\n\nSoft spots are real but mostly the ones the paper already owns. Everything is under one agent and one judge; the partition is observational, not a randomized intervention on ranks or snippets; n=100 is diagnostic; Brave’s richer native snippets partly inflate pre-fetch volume (they normalize for this and the advantage shrinks but does not vanish). Those limits bound transfer, not the scoped claim that equal accuracy can hide different evidence economies under a fixed policy.\n\nMath and metrics look clean: no circular scoring, correctness separated from the oracle, rc:g independent of the model answer. Citations are the right neighborhood (RAG, ReAct/WebGPT, BEIR, SEALQA, RAGAS). Code and aggregate artifacts are released; raw provider content is not, for the usual rights reasons.\n\nThis is for people building or evaluating tool-using agents and progressive-disclosure retrieval, not for classical IR theory. I would send it to peer review. Engage with it if you care about agent retrieval budgets; treat the regimes as provider-associated under this policy, not universal rankings.","headline":"Accuracy parity is real under a frozen agent; the useful result is that search providers still create different pre-fetch evidence economies and fetch regimes.","tokens_in":17713,"tokens_out":579,"would_cite":true,"duration_ms":7526,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Search APIs that look interchangeable by answer accuracy actually hand agents different decision surfaces that change fetch cost, exploration, and contradiction risk.","keywords":["search APIs","tool-using agents","decision surfaces","progressive disclosure","retrieval-augmented generation","pre-fetch evidence","contradiction-to-gold ratio","agent evaluation"],"falsifier":"Rerun the same 100 questions with a second answer model or a deliberately different fetch policy while still freezing the provider adapters and oracle; if pre-fetch support, rank-1 concentration, contradiction ratios, and decision-cell shares collapse to the same profile across providers, the claim that equal accuracy hides distinct decision surfaces fails.","tokens_in":17676,"feed_emoji":"🔎","tokens_out":1007,"duration_ms":12207,"temperature":0.7,"pith_summary":"This paper argues that commercial web search APIs should not be judged mainly by final answer accuracy for tool-using agents. Because agents usually see ranked titles, URLs, and snippets first and only then decide whether to open pages, the pre-fetch result list is a decision surface: it shapes whether the agent answers immediately, searches again, or spends tokens on fetches. Holding one agent, prompt, tools, and fetch backend fixed while swapping only the search provider, the authors find nearly identical audited correctness on a hard 100-question set, yet sharply different evidence economies. One provider surfaces more gold answers in snippets, another concentrates support at rank 1, and a third is tied to broader page exploration; a new surface contradiction-to-gold ratio also swings widely. A sympathetic reader should care because provider choice then becomes a retrieval-budget and policy choice, not a simple recall or leaderboard swap.","feed_headline":"Equal agent accuracy, unequal search evidence","feed_subtitle":"Same answers hide different snippet wealth, rank focus, and fetch costs across search APIs.","key_machinery":"The decision surface: the ranked snippets, URLs, and metadata visible before any page fetch. The paper measures it with a per-URL oracle that labels every element the agent saw, splits pre-fetch support from post-fetch discovery, assigns each query to a four-way action partition (SMART, MISSED, BLIND, NO-OP), and computes a surface contradiction-to-gold URL ratio over snippet-only rows.","core_discovery":"Under a frozen agent and an audited semantic-match correctness label, three commercial search providers reach nearly the same accuracy (25, 25, and 26 of 100 hard questions), but their pre-fetch surfaces differ sharply: gold-answer-rich snippets, rank-1 concentration of supporting URLs, and broader exploration regimes, with surface contradiction-to-gold ratios from 0.92 to 2.59. Equal accuracy therefore masks unequal evidence economies, so a search API is better understood as a decision surface than as a static ranked-list retriever.","pith_inferences":["Agent scaffolding that hard-codes “always fetch top-1” or “never fetch if the snippet looks answerable” will systematically favor some providers and punish others even when accuracy looks tied.","Cost and latency SLOs for production agents may move more from swapping models than from swapping search providers, once surface-aware fetch policies are tuned.","A natural next stress test is whether the same surface signatures appear on fresher or multi-hop tasks where contradiction density and rank instability are higher.","If providers begin optimizing for agent actionability, classical IR leaderboards may diverge further from what progressive-disclosure agents actually experience."],"forward_implications":["Provider selection should be paired with a provider-aware fetch policy rather than a single universal top-k heuristic.","Agent evaluation needs decision-surface metrics—pre-fetch support, rank concentration, contradiction contamination, and fetch budget—not only final accuracy or gold-URL hit rate.","Practitioners can score candidate search APIs from traces by judging visible URLs and classifying snippet-rich, rank-concentrated, or exploration-heavy surfaces.","Agent-ready search products should optimize actionability: calibrated snippets, low contradiction-to-gold contamination, and rankings aligned with common fetch policies.","Because providers solve overlapping but different question sets, routing or multi-provider strategies can matter more than crowning a single winner."],"fun_headline_variants":["Equal accuracy, unequal search evidence for agents","Same answers, different decision surfaces across APIs","Search APIs match scores but not evidence economies","Frozen agent: equal accuracy, divergent snippet surfaces","Gold-rich snippets vs rank focus under equal scores"],"cache_read_input_tokens":128,"weakest_assumption_plain":"That the provider-linked fetch, support, and exploration patterns seen under one fixed agent policy and one judge will still describe how other agents and prompts behave on the same surfaces.","fun_headline_variants_meta":{"raw":{"variants":["Equal accuracy, unequal search evidence for agents","Same answers, different decision surfaces across APIs","Search APIs match scores but not evidence economies","Frozen agent: equal accuracy, divergent snippet surfaces","Gold-rich snippets vs rank focus under equal scores"]},"model":"grok-4.5","effort":"low","cost_usd":0.005716,"raw_usage":{"total_tokens":1622,"prompt_tokens":908,"num_sources_used":0,"completion_tokens":54,"cost_in_usd_ticks":57160000,"prompt_tokens_details":{"text_tokens":908,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":660,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":908,"tokens_out":54,"duration_ms":5568,"temperature":1.0,"reasoning_tokens":660,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-14T13:33:50.125695+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Rerun the same 100 questions with a second answer model or a deliberately different fetch policy while still freezing the provider adapters and oracle; if pre-fetch support, rank-1 concentration, contradiction ratios, and decision-cell shares collapse to the same profile across providers, the claim that equal accuracy hides distinct decision surfaces fails.","supporting_citations":[],"review_version":1}