{"id":"54b5a034-a5e6-4279-9fb2-d404e10e9253","arxiv_id":"2506.02865","paper_version":2,"verdict":"REJECT","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":5,"one_line_summary":"The paper claims a state-of-the-art, low-cost open-weight web agent, but its WebVoyager result is inflated by training on the evaluation benchmark.","lead":"H Company presents Surfer-H, a web agent built around new open-weight vision-language models (Holo1) that they say reaches 92.2% accuracy on the WebVoyager benchmark at a fraction of the cost of commercial agents. The headline number is misleading, because the models were trained on WebVoyager tasks and the paper's own ablation shows that in-domain training adds 4.5 points.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 92.2% WebVoyager score is not a held-out result: WebVoyager tasks are in the Holo1 training mixture (Sec. 3.5), and Table 4 attributes 4.5 points to exactly this in-domain training.","rationale":"The reader's weakest assumption is exactly the load-bearing point: the WebVoyager evaluation is not held-out because WebVoyager traces are in the training mixture. Table 4 provides direct internal evidence that this choice contributes 4.5 points to the headline number, and the paper's 'Performance on Unseen Tasks' section does not actually evaluate on an unseen task set. The central claim of state-of-the-art WebVoyager accuracy therefore conflates in-domain fitting with generalization. This is not a disagreement with current consensus; it is a mismatch between the claim and the experimental protocol. The cost-efficiency claim is secondary and also has issues (internal versus external cost estimates, baselines evaluated at different times), but the accuracy contamination is sufficient on its own to invalidate the Pareto-optimal SOTA claim. I credit the open release of Holo1 and WebClick and the stronger localization results on Screenspot and GroundUI, but those do not rescue the headline claim. The reader's REJECT verdict is supported; my independent read leaves it unchanged.","tokens_in":12503,"tokens_out":6756,"duration_ms":73320,"concrete_test":"Construct a fresh WebVoyager-style evaluation set: 100 tasks on 10 websites not present in WebVoyager or WebVoyagerExtended, annotated using the same majority-vote GPT-4o protocol. Run Surfer-H+Holo1-7B and Surfer-H+Holo1-7B-WVE with 10 attempts and GPT-4o as validator on this held-out set. If Holo1-7B's accuracy is not materially above Holo1-7B-WVE (e.g., within 2 points) and remains near 92.2%, the training-on-WebVoyager explanation for the SOTA score is ruled out; if it drops toward or below the 87.7% WVE level, the reported 92.2% is an artifact of test-set training.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim of state-of-the-art WebVoyager performance depends on treating the 92.2% accuracy as a measure of generalization. The paper itself breaks this condition. Section 3.5 states that the policy training mixture includes agent traces generated on WebVoyager tasks, and Section 5.2's Table 4 quantifies the consequence: removing WebVoyager traces (Holo1-7B-WVE) drops WebVoyager accuracy by 4.5 points, from 92.2% to 87.7%. The paper labels this 'in-domain experience,' but since the evaluation set is exactly the WebVoyager task corpus, this is training on the test set. The subsection titled 'Performance on Unseen Tasks' is therefore misnamed: it compares a model trained on WebVoyagerExtended with one additionally trained on WebVoyager, and evaluates both on WebVoyager; it never evaluates on tasks absent from training. Because the 4.5-point contribution is larger than the 3.1-point margin over BrowserUse, the claimed SOTA ordering is not robust to this contamination. Additionally, WebClick's agent subset (Section 4.2) is derived from agent attempts on WebVoyager tasks, so it cannot serve as an independent generalization check for a model trained on the same traces. The localization results on Screenspot and GroundUI, and the open release of weights and WebClick, remain useful contributions, but they do not support the headline WebVoyager claim as stated.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces Surfer-H, a modular web agent composed of a policy, a localizer, and a validator, and Holo1, a family of open-weight VLMs fine-tuned from Qwen2.5-VL. It also introduces WebClick, a new web-localization benchmark. The headline claim is that Surfer-H powered by Holo1-7B reaches 92.2% accuracy on WebVoyager at an average cost of $0.13 per task, presented as state-of-the-art and Pareto-optimal. However, the WebVoyager task corpus is explicitly included in the policy training mixture (Section 3.5), and the paper's own ablation (Table 4) shows that removing those traces lowers accuracy by 4.5 points to 87.7%. The 92.2% result is therefore not a held-out measurement. The localization results on public benchmarks and the open release of model weights and WebClick are useful contributions, but they do not support the WebVoyager SOTA claim as stated.","tokens_in":12733,"tokens_out":5197,"duration_ms":48244,"significance":"If the WebVoyager result were a clean held-out evaluation, demonstrating that a 7B open-weights model can outperform proprietary agents such as OpenAI Operator and BrowserUse at a fraction of the cost would be an important result. The paper also provides publicly released weights and a new benchmark, and the localization improvements on Screenspot and GroundUI are credible evidence of progress on UI grounding. However, the central SOTA claim is compromised by in-domain training on the evaluation corpus, and the paper's own ablation quantifies the contamination. The remaining contributions are genuine but incremental; the headline result as framed is not supported.","major_comments":[{"comment":"The headline 92.2% WebVoyager score is not a held-out result. Section 3.5 states that the policy training mixture includes agent traces generated on the WebVoyager task corpus, and Table 4 shows that removing those traces (Holo1-7B-WVE) lowers WebVoyager accuracy from 92.2% to 87.7%. Since the evaluation is run on exactly the WebVoyager task set, this 4.5-point difference constitutes training on the test distribution. Because this difference is larger than the 3.1-point margin over the strongest reported baseline (BrowserUse at 89.1%), the abstract's 'state-of-the-art' and Pareto-optimality claims are not robust once the contamination is accounted for.","section":"Section 3.5, Table 4"},{"comment":"This subsection is misnamed. It compares a model trained on WebVoyagerExtended only (Holo1-7B-WVE) with one additionally trained on WebVoyager traces (Holo1-7B), and evaluates both on WebVoyager. There is no evaluation on a task set absent from both training mixtures, so the experiment does not demonstrate generalization to unseen tasks; it demonstrates the effect of adding the evaluation corpus to training.","section":"Section 5.2, 'Performance on Unseen Tasks'"},{"comment":"The WebClick benchmark's agent subset is built from agent attempts on WebVoyager tasks. Since Holo1 is trained on WebVoyager traces (Section 3.5), the WebClick agent-subset scores (e.g., Holo1-7B at 89.77% vs Qwen2.5-VL-7B at 78.47%) are not an independent localization check. The human and calendar subsets are less exposed to this issue, but the average scores in Table 2 and Figure 2 include the contaminated agent subset, so the claimed localization advantage is also partly confounded.","section":"Section 4.2, Table 2"},{"comment":"The external baselines (BrowserUse, Operator, Mariner) are taken from reported numbers, computed at a different time, with different websites and evaluation functions, as the paper itself notes in Section 5.1. Combined with the in-domain training issue, this does not support the claim that Surfer-H+Holo1 is Pareto-optimal against these systems. A controlled re-evaluation under the same harness would be needed to make such a comparison.","section":"Section 5.1, Table 5"}],"minor_comments":[{"comment":"Section 3.5 describes WebVoyager as comprising 643 tasks on 15 common websites, while Section 5.1 says 'all 643 tasks from 10 different websites.' Please reconcile these numbers.","section":"Section 3.5 vs Section 5.1"},{"comment":"The reported average localization scores (73.55% for Holo1-3B and 76.16% for Holo1-7B) do not match the arithmetic means of the seven columns in Table 2 (73.00% and 76.19%, respectively). Please correct the text or the table.","section":"Section 4.3 vs Table 2"},{"comment":"The cost comparison relies on internal estimates for Holo1 and Qwen2.5-VL inference costs; it would be helpful to state the sensitivity of the Pareto-front conclusions to these estimates.","section":"Section 5.1, Table 3"},{"comment":"The plot legend is dense and some points are not individually readable; consider labeling the Surfer-H/Holo1 points directly or adding a companion table with exact values.","section":"Figure 3"}],"recommendation":"reject","confidential_remarks":"The central WebVoyager claim is not held-out, and the paper's own ablation quantifies the contamination. The 'unseen tasks' experiment does not actually test unseen tasks. I do not see how a revision can salvage the headline SOTA claim without retraining and re-evaluating on a truly held-out benchmark; the WVE-only model scores 87.7%, which would no longer be SOTA. The localization and open-source contributions are genuine, but they are not enough to offset the core evaluation problem."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The reader and stress-test got the main thing right: the central WebVoyager claim is not a clean generalization result. Section 3.5 says the policy training mixture includes agent traces from WebVoyager itself, and Table 4 shows that removing those traces costs 4.5 points (87.7% vs 92.2%). Since the margin over BrowserUse is only 3.1 points, the claimed SOTA ordering is not robust. The paper is transparent about the mixture, which is good, but the abstract and Section 5 still sell the number as if it measured generalization. That is an overclaim, not a fabrication.\n\nWhat is genuinely new and valuable: the Holo1 open-weight models, the WebClick localization benchmark, and the careful modular design of Surfer-H (policy, localizer, validator). The localization results on independent benchmarks — Screenspot v1/v2/Pro and GroundUI — are credible and competitive for 3B and 7B models, and the open release of weights and benchmark is a concrete contribution that the community can build on. The WebClick benchmark itself is also useful, even if the agent-sourced portion overlaps with WebVoyager-derived training data.\n\nSoft spots beyond the main one: WebClick's agent subset is built from agent attempts on WebVoyager tasks, so it cannot independently validate a model trained on those same traces. The external baselines (Operator, Mariner, BrowserUse) are all reported numbers from different times, websites, and evaluation functions, so the Pareto-optimal cost comparison is not apples-to-apples. These are real caveats but they don't sink the localization contributions; they just mean the paper's headline should be reframed.\n\nWho is this for? Researchers working on GUI agents, web navigation, and UI grounding. It's a solid engineering-plus-data paper with an evaluation design flaw in its flagship claim. I'd bring it to a reading group to discuss train/eval contamination and what counts as a benchmark. I would cite it for WebClick and the Holo1 release. It deserves a serious referee: the artifacts are real and the main claim is fixable by either reporting a held-out evaluation or honestly labeling the WebVoyager result as in-domain performance. Send it to peer review with a request for major revision.","headline":"The 92.2% WebVoyager headline is not a held-out result — the paper's own ablation gives 4.5 points to in-domain training — but the open Holo1 weights and the WebClick benchmark are real, useful contributions that deserve a close look.","tokens_in":13588,"tokens_out":1295,"would_cite":true,"duration_ms":15502,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A small open-weights model can power a state-of-the-art web agent at a fraction of the cost, if the benchmark measures what it claims.","keywords":["web agent","vision-language model","GUI grounding","UI localization","web benchmark","open-weight model","cost efficiency","behavioral cloning"],"falsifier":"Run Surfer-H with Holo1-7B (trained with the full mixture) on WebVoyager-style tasks drawn from websites the agent never saw in training, or on an independent benchmark such as VisualWebArena; if accuracy falls to the ~87.7% level of the model trained without WebVoyager traces, the 92.2% headline is substantially a result of in-domain training rather than general web competence.","tokens_in":12195,"feed_emoji":"🤖","tokens_out":11155,"duration_ms":98284,"temperature":0.7,"pith_summary":"The paper claims that a web agent can reach state-of-the-art performance on WebVoyager using only screenshots and an open-weights 7-billion-parameter vision-language model, at an average cost of about $0.13 per task. The mechanism is Surfer-H, a three-module agent (policy, localizer, validator) paired with Holo1, a VLM family fine-tuned from Qwen2.5-VL on a large mixture of web-grounding, synthetic, and agent-trajectory data. If true, this would push the Pareto frontier of accuracy versus cost for computer-use agents, beating reported numbers from proprietary counterparts. The paper also introduces WebClick, a web-specific localization benchmark, and releases both the model weights and the benchmark. Its own ablation shows that part of the gain comes from training on WebVoyager traces, so the headline number is not a clean held-out evaluation.","feed_headline":"Open-weights 7B web agent hits 92.2% at 13 cents a task","feed_subtitle":"A screenshot-only agent with an open 7B model beats costlier proprietary rivals on a standard web task.","key_machinery":"The load-bearing mechanism is the Surfer-H loop: at each step the policy VLM predicts thought, optional note, and next action from the task, memory, and recent screenshots; if the action is a click or a write, a localizer VLM turns the element description into pixel coordinates; when the policy emits an answer, a validator VLM scores it against the task and screenshots and, on rejection, feeds the feedback into memory for another attempt. Holo1 is a single VLM family trained to serve all three roles, and its training mixture—particularly the filtered behavioral cloning of successful traces and the 5M-triplet coordinate-validation set—is what the paper credits for state-of-the-art localization and policy behavior.","core_discovery":"Surfer-H is a screenshot-only web agent whose policy emits thoughts and actions, whose localizer converts element descriptions into coordinates, and whose validator checks final answers, with all modules able to share a single VLM. Holo1 is that VLM family: starting from Qwen2.5-VL-Instruct, it is fine-tuned on a 31.5B-token mixture that is 50.8% GUI-grounding data, 32.3% complex-visual-understanding data (including a novel 5M-sample coordinate-validation task and 7M-page UI extraction), and 16.9% behavior data from successful agent traces, plus a 1M-pair validator corpus. The paper's central result is that Surfer-H with Holo1-7B as policy and localizer and GPT-4o as validator attains 92.2% accuracy on all 643 WebVoyager tasks after up to 10 attempts, at an estimated $0.13 per task, outperforming the reported Operator (87.0%), Mariner (83.5%), and BrowserUse (89.1%) figures and matching GPT-4.1-driven Surfer-H (92.0%) at a quarter of the cost. It also reports that Holo1 tops average accuracy on localization benchmarks including the new WebClick set, and that a fully self-hosted Holo1-7B agent reaches 80.4% at $0.06 per task. The authors attribute the results to the training mixture, especially filtered behavioral cloning of successful agent traces, and release WebClick and the Holo1 weights.","pith_inferences":["The 4.5-point gap between Holo1-7B and Holo1-7B-WVE is a direct estimate of how much of the 92.2% rests on having trained on the evaluation benchmark's own tasks; I would expect a genuinely unseen set of comparable web tasks to land closer to the 87.7% figure.","The drop observed when Holo1 replaces GPT-4o as validator suggests answer verification is harder than localization or action selection for small models, and might respond to larger validation-specific training sets or a stronger validator model.","The Pareto-curve methodology (accuracy against average cost per task as attempts vary) is transferable to other agent frameworks as a practical way to pick a deployment point given a budget.","A direct testable extension: applying the same training mixture and three-module design to other agent benchmarks such as VisualWebArena or WebArena would show whether the GUI-grounding gains generalize beyond WebVoyager-style tasks."],"forward_implications":["If accurate, the result means a 7B-parameter open model can substitute for frontier API models as the policy of a computer-use agent, cutting per-task cost from $0.54 to $0.13 at matched WebVoyager accuracy.","A screenshot-only interface without DOM or accessibility trees is enough for near-state-of-the-art web navigation, which removes a dependency on site-specific integrations and makes the agent portable to any graphical interface.","Filtered behavioral cloning of successful traces is a transferable training strategy: it contributes a 9.5-point gain over the base model, and adding in-domain WebVoyager traces contributes another 4.5 points.","The WebClick benchmark offers a compact, web-specific measure of localization skill, on which Holo1 models lead their size class; this makes localization an isolatable and optimizable component of agent performance.","Fully self-hosted deployment is feasible: Surfer-H with Holo1 in all three roles stays on the Pareto front at $0.06 per task, though with a 12-point accuracy drop, which locates validation as the current bottleneck."],"supporting_citations":[{"why":"This is the WebVoyager benchmark on which the 92.2% claim is measured, and it is also the source of the in-domain agent traces used in Holo1's training mixture.","marker":"[11]"},{"why":"Qwen2.5-VL is the base model whose weights Holo1 fine-tunes, and it provides the internal baseline that Holo1 is compared against.","marker":"[2]"},{"why":"Operator is a proprietary agent whose reported 87.0% WebVoyager score is the main external state-of-the-art point Surfer-H claims to exceed.","marker":"[23]"},{"why":"BrowserUse is an external web-agent baseline whose reported 89.1% is the closest number Surfer-H+Holo1-7B (92.2%) is compared against.","marker":"[3]"},{"why":"Project Mariner is a proprietary agent baseline whose reported 83.5% appears in the Pareto accuracy-cost comparison.","marker":"[8]"},{"why":"OS-Atlas supplies open-source GUI-grounding data that complements the WebCrawl mixture and the Screenspot-v2 localization benchmark numbers.","marker":"[34]"},{"why":"Cauldron is the VQA dataset, remapped and filtered, used in the complex-visual-understanding portion of Holo1's training mixture.","marker":"[14]"},{"why":"UI-TARS is a baseline VLM evaluated on the localization benchmarks, including the new WebClick, against which Holo1's localization results are positioned.","marker":"[27]"}],"fun_headline_variants":["Open 7B web agent hits 92.2% on WebVoyager for 13 cents","Surfer-H with Holo1: 92.2% at $0.13 per task","Open-weights 7B agent: Pareto-optimal balance at 92.2%","Screenshot-only agent: 92.2% accuracy, 13 cents per task","Holo1-7B powers 92.2% web agent at a fraction of cost"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The result stands on the assumption that a model trained partly on traces of WebVoyager tasks can be measured on WebVoyager as a test of generalization; the paper's own ablation suggests removing those traces costs 4.5 points of accuracy.","fun_headline_variants_meta":{"raw":{"variants":["Open 7B web agent hits 92.2% on WebVoyager for 13 cents","Surfer-H with Holo1: 92.2% at $0.13 per task","Open-weights 7B agent: Pareto-optimal balance at 92.2%","Screenshot-only agent: 92.2% accuracy, 13 cents per task","Holo1-7B powers 92.2% web agent at a fraction of cost"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001063,"raw_usage":{"total_tokens":4500,"prompt_tokens":1034,"completion_tokens":3466,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":650,"completion_tokens_details":{"reasoning_tokens":3345}},"tokens_in":650,"tokens_out":3466,"duration_ms":22947,"temperature":1.0,"reasoning_tokens":3345,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T11:14:29.809204+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run Surfer-H with Holo1-7B (trained with the full mixture) on WebVoyager-style tasks drawn from websites the agent never saw in training, or on an independent benchmark such as VisualWebArena; if accuracy falls to the ~87.7% level of the model trained without WebVoyager traces, the 92.2% headline is substantially a result of in-domain training rather than general web competence.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"This is the WebVoyager benchmark on which the 92.2% claim is measured, and it is also the source of the in-domain agent traces used in Holo1's training mixture."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Qwen2.5-VL is the base model whose weights Holo1 fine-tunes, and it provides the internal baseline that Holo1 is compared against."},{"cited_title":"https://openai.com/index/introducing-operator/","cited_arxiv_id":null,"evidence_quote":"Operator is a proprietary agent whose reported 87.0% WebVoyager score is the main external state-of-the-art point Surfer-H claims to exceed."},{"cited_title":"Browser use: Sota technical report.https://browser-use.com/posts/ sota-technical-report, 2024","cited_arxiv_id":null,"evidence_quote":"BrowserUse is an external web-agent baseline whose reported 89.1% is the closest number Surfer-H+Holo1-7B (92.2%) is compared against."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Project Mariner is a proprietary agent baseline whose reported 83.5% appears in the Pareto accuracy-cost comparison."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"OS-Atlas supplies open-source GUI-grounding data that complements the WebCrawl mixture and the Screenspot-v2 localization benchmark numbers."}],"review_version":1}