REVIEW 2 major objections 6 minor 131 references
Routing Is Least Learnable Where It Is Most Valuable: Bounds on Representation Routing for Web Agents
T0 review · 2 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read The paper argues that representation routing for web agents is currently bounded by supervision scarcity, not by the routing estimator, and that the surviving upper bound is a no-learning cost saving.
desk verdict Careful, honest measurement of representation routing for web agents that deserves peer review, though the headline slogan overreaches the data. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The instrument that carries every comparison is the same-condition rerun band: rerunning one mode on the same task set changes the outcome of 12–14% of tasks on the replicated cell, with a set-level movement of 4.91–7.59pp and a resolution threshold of roughly 3.8–4.2pp under an exchangeability null, and every mode-to-mode or policy-to-policy difference is read against this band rather than against zero. The formal object underneath the argument is the solve matrix — which tasks each mode solved in each cell — because label supply (rows some mode solved), the routable set (rows more than one mode solved), and the instability rows are all functionals of that matrix, which is why the paper can show that two apparently separate obstructions are one: both track the best single mode's success rate at $\rho = 0.952$.
What would settle it
Take one of the six cells that carried no replicate, such as red·B2 or WA·B1, rerun one mode on the same scored task set, and compare the flip rate with the imported 12–14% band and the 3.8–4.2pp resolution threshold; a materially different flip rate in that cell would invalidate every comparison that imported the floor. A complementary check is to rerun the whole measurement on an agent with success above 50%: the paper predicts label supply and the contested set grow together, and if a routing policy then still fails to beat a fixed mode, the supervision-scarcity diagnosis would not hold.
Extended reading notes
Core claim
The central claim is that representation routing for web agents is currently bounded by supervision scarcity rather than by the routing estimator. Six observation modes were run across eight site–backbone cells; 44 of 48 mode–cell pairs solved at least one task no other mode solved, and the winning mode reversed between the classifieds and WebArena task sets. The apparent prize, an oracle that picks a winning mode per task, is a six-arm union quoted against one arm: rerunning one condition changes 12–14% of task outcomes, and adding the best distinct arm buys +1.97 to +7.14pp while rerunning an arm already in hand buys 4.91–7.59pp where replicates exist. The upper bound that survives the rerun control is the cost ceiling, which adds no arm: keeping the best mode everywhere and sending never-solved tasks to the cheapest arm leaves success unchanged by construction and cuts cost 9.5–30.6% in 8 of 8 cells. Five routing policies — mode selection, learned spend triage, a zero-token rule read off the task text, a confidence cascade, and pooled cost tiers — all fail to robustly beat a fixed mode, each for a different reason, and the common obstruction is that the rows a router can learn from exist only where some mode succeeded, while the rows where a per-task choice exists at all are the same contested rows that flip between identical reruns. Across the eight cells, label supply and routing opportunity track the best single mode's success rate at $\rho = 0.952$, so the negative result is indexed to today's 2–36% success regime and predicts its own reversal if a stronger agent solves most tasks.
Load-bearing premise
The load-bearing premise is that the rerun noise floor measured in two cells transfers to the six cells that carried no replicate, so every claim that a mode difference clears the noise floor in those cells assumes the flip rate is the same there.
Editorial extensions
If this is right
- A builder should price a rerun before pricing a representation: until a same-condition rerun band is measured, a gain from adding a representation cannot be attributed to the representation rather than to resampling.
- The cost ceiling is reachable with one bit per task and no learned policy: sending only never-solved tasks to the cheapest arm cuts cost by 9.5–30.6% at unchanged success in all eight cells.
- None of the five tested routing policies robustly beats simply fixing one well-chosen mode; the single exception sits in the sparsest cell and its triage signal is below chance.
- Graded evaluators would attack the supply obstruction directly, because a graded per-task score would make every episode a training signal regardless of success.
- A stronger agent can overturn the result: since label supply and routing opportunity rise together, routing becomes more learnable as success rates rise.
Reading between the lines
- The paper leaves untested richer router classes such as LLM-based routers, reinforcement learning, and contextual bandits; if the supply obstruction is truly estimator-independent, those are unlikely to help on today's benchmarks, but that remains an open question.
- The near-linear growth of the routable share in per-mode success (where independent modes would give near-quadratic growth) suggests task difficulty dominates mode–task matching on these workloads; a multi-family replication could test whether that dominance is a general property or an artefact of the two shared benchmarks.
- The cost-ceiling result is the immediately deployable finding: a static rule that requires no training could be adopted by production web agents today even while learned routing remains closed.
- A natural next experiment is to measure whether soft or graded labels on the contested rows reduce the flip-rate problem, since the paper shows the contested rows are exactly the rows that flip between identical reruns.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper measures six observation representations of a browser page (accessibility-tree text, vision, set-of-marks, and three screenshot-free ablations) under one scaffold, one prompt budget and one action space, across eight site-by-backbone cells on VisualWebArena and WebArena, totalling 7,686 scored episodes. Three empirical claims are made. (i) The modes are complementary: 44 of 48 mode–cell pairs solve at least one task no other mode in that cell solved, the losing channel fails in structurally different ways, and the winning deployment class reverses between task sets (§3). (ii) The apparent per-task oracle ceiling (7.1–51.9% against a best single mode of 2.2–35.6%) is largely an artefact of run-to-run noise: rerunning one condition changes 12–14% of task outcomes on the fully replicated cell, and adding a genuinely distinct arm (+1.97 to +7.14pp) buys about as much as a rerun (4.91–7.59pp); the bound that survives the rerun control is a cost ceiling of 9.5–30.6% savings at unchanged success in all 8 cells (§5). (iii) Five routing policies (which-mode selection, learned triage, a zero-cost regex rule, a confidence cascade, and pooled cost tiers) fail to robustly beat a fixed, well-chosen mode, each for a distinct reason (§6).
Significance. If the measurements hold, this is a valuable and unusually rigorous empirical contribution to web-agent evaluation. The same-condition rerun floors, the careful separation of observed band from resolution threshold, the set-difference versus mean-difference estimand distinction, the paired bootstrap intervals, the earned-versus-leaked success audit, the machine-checked provenance pipeline described in §8, and the explicit falsifiability statement in §7 are exemplary practices that the field should adopt; the cost ceiling (Table 16, 'same tasks, lower cost') is a clean, machine-checkable bound, and the five-policy negative is informative precisely because the five policies fail for different reasons (label supply, label ambiguity, and value).
major comments (2)
- [Abstract; §7 (Falsifiability); Tables 15–16] The claim that routing is 'least learnable exactly where it would be most valuable' is not supported by the paper's own measurements because the sense of 'valuable' is never defined, and the operationalizations the paper's tables provide point in the opposite direction. Under absolute headroom (Table 16, 'headroom' column), value is largest in the strongest cells (WA·B0 +16.35pp, cls·B0 +16.07pp) and smallest in the weakest (red·B2 +3.45pp, cls·B2 +4.91pp). Under Table 15's own label for routing value — the routable set of tasks with more than one solver — value tracks the best single mode's success rate at Spearman rho = 0.952, i.e. positively, which the caption itself summarizes as 'the two obstructions are one wall, whose height is set by how much the agent can do.' Under the only definition that produces an inverse pattern, the ceiling-to-best ratio, the pattern is driven by the B2 cells (cls·B2 at 3.20x, red·B2 at 1.88x), which the Limitations explicitly disqualify ('should not carry a comparison'); across the six B0/B1 cells the ratio is flat (1.46–1.88x) and shows no monotone inverse relationship with success rate. The Abstract contains the same tension in adjacent sentences, asserting both 'exactly where routing would be most valuable' and 'label supply and routing opportunity rise together (correlation 0.95 across cells).' The empirical content — supervision scarcity and the five-policy negative — does not require the inverse-value gloss. The authors should either define a value metric that survives their own caveats and establish the inverse relationship with it, or reframe the contribution as: routing supervision and routing opportunity are both scarce in the weak-agent regime, so the current regime affords neither the labels nor the headroom to train representation routers.
- [Abstract; §5; Tables 16 and 19] The statement that 'a second run of a mode already in hand gains about as much as adding a new one' is established in only one of the two replicated cells. On cls·B0 the comparison is genuinely close: +7.14pp for the best distinct arm versus a 4.91–7.59pp measured rerun band (Table 16). On wa_red·B1 the best distinct arm (DOM+stext, +4.81pp) exceeds that cell's ten-task rerun draw under the pooled band (2.00–4.00pp) that the paper itself reports, and the decision to mark that row 'inside the rerun band' (Table 19) relies on the cell's own 0.00–10.00pp band, which is one task wide and therefore nearly uninformative. The unqualified Abstract sentence and the Table 19 verdict thus overstate what the replication data support; I recommend restricting the 'as much as a rerun' claim to cells with full-arm replicates and reporting wa_red·B1 as the one replicated cell where the distinct-arm gain sits above the own-cell floor.
minor comments (6)
- [Appendix A; Table 4 caption] The appendix text still quotes a '≥83% consistency bar' for the repeated-extrema tally, while Appendix Table 5's caution states that the correct threshold is ≥7/8 = 87.5% and that the 83% figure was wrong; the stale value should be corrected wherever it appears.
- [Appendix Table 5] The caution admits the 7-of-8 bar 'was chosen after the cell count changed' and is 'load-bearing and disclosed rather than defended'; because the choice is post hoc, please add a sensitivity analysis (for instance the ≥5/8 tally that the caution says would let DOM+stext clear two metrics) so a reader can confirm the deployment-class grouping of §2 is unaffected, since the table is currently not allowed to license that grouping on its own.
- [Abstract; §5] The statement that 'rerunning the same mode on the same tasks changes 12–14% of outcomes' is measured on the three full cls·B0 arms; the wa_red·B1 ten-task draw shows 6.0% (2.00–4.00pp in each direction). Please qualify the abstract figure as the fully replicated cell's rate.
- [Abstract] The phrase 'correlation 0.95 across cells' should identify the statistic as the Spearman rho = 0.952 of Table 15 between the routable set and the best single mode's success rate, and should carry the caption's caveat that part of the association is structural because both columns are functionals of one solve matrix.
- [§3 (No portable winner)] The section applies the §4 resolution threshold (3.8–4.2pp) to all eight cells even though the threshold is derived from the cls·B0 replicates and the Limitations state it is imported into six cells; one pointer at the point of use would connect the section-3 claims to that limitation.
- [§7 (Falsifiability)] The 'mean gap 1.65pp' is the mean absolute difference between the best single mode's success rate and the routable share, in percentage points of the cell; please state the unit and the estimand at first use, since the immediately adjacent rho = 0.952 is on a different scale.
Circularity Check
No significant circularity: the central bounds are measured against external benchmarks, and the two potentially circular-looking identities are explicitly labeled as constructions.
full rationale
The paper's derivation chain is self-contained. The success-rate ceiling is a six-arm union and is explicitly corrected with rerun controls rather than presented as a prediction. The surviving cost ceiling (9.5-30.6%) is an accounting identity: rerouting only never-solved tasks leaves success unchanged by construction, and the paper says so in Section 5 and Table 16; an identity presented as a bound is not a circular derivation. The rho=0.952 coupling between label supply and routable share is acknowledged in Table 15 as partly structural, with both columns being functionals of one solve matrix, and the non-structural near-linearity is separated out; no hidden equivalence is smuggled in. The five routing negatives are empirical: the triage policy uses nested cross-validation, the regex rule is ex-ante, and the cascade is compared against a random-escalation control. There are no load-bearing self-citations, no imported uniqueness theorem, and no fitted parameter renamed as a prediction. The slogan that routing is least learnable where it is most valuable is interpretive rather than derived, since the paper never defines 'valuable' and the B2 cells that would support the relative-gain reading are disqualified by the paper's own Limitations; but an unsupported interpretive claim is a correctness risk, not circularity.
Assumptions & free parameters
assumptions (5)
- domain assumption The rerun noise band measured on cls.B0 and wa_red.B1 transfers to the six cells that carry no replicate.
- domain assumption VisualWebArena and WebArena evaluators' binary scores are the accepted success ground truth.
- standard math The exchangeability null for rerun differences, where each discordant task flips with probability one half, is the correct model for the resolution threshold.
- ad hoc to paper The threshold of 7 of 8 cells for Table 5's repeated-extrema bar was selected after the cell count changed, and this choice is load-bearing for the non-separability of the four text modes.
- domain assumption Cascade outcomes can be approximated by an offline splice of standalone rich runs.
Cite this review
Pith. "Pith review of Routing Is Least Learnable Where It Is Most Valuable: Bounds on Representation Routing for Web Agents." pith.science (2026). https://pith.science/paper/AIPQUGS4
@misc{pith2026260806171,
author = {Pith},
title = {Pith review of: Routing Is Least Learnable Where It Is Most Valuable: Bounds on Representation Routing for Web Agents},
year = {2026},
howpublished = {\url{https://pith.science/paper/AIPQUGS4}},
note = {Machine review of arXiv:2608.06171}
}
read the original abstract
Web agents observe a browser through text, pixels, or both, and the choice is usually fixed once for all tasks. We measure six observation modes across eight site-model combinations (cells) on VisualWebArena and WebArena and ask what choosing per task would buy. The modes are complementary: each solves tasks the others miss, they fail in structurally different ways, and the best choice reverses between task sets. The obvious prize, an oracle that picks a winning mode for every task, looks large but is inflated by run-to-run noise: rerunning the same mode on the same tasks changes 12-14% of outcomes, so a second run of a mode already in hand gains about as much as adding a new one. What survives is a cost bound: sending only the tasks no mode solves to the cheapest mode cuts cost by 9.5-30.6% in 8 of 8 cells at unchanged success. We then test five routing policies (picking the mode, deciding when to spend on the strong mode, a zero-cost rule read off the task text, a confidence cascade, and pooled cost tiers), and none robustly beats simply fixing one well-chosen mode; the one exception is a fragile result in our sparsest cell. The central obstruction is that routing supervision is produced at the agent's success rate: the weaker the agent, the fewer labels a router gets, exactly where routing would be most valuable. This limit belongs to today's agents rather than to routing itself. Label supply and routing opportunity rise together (correlation 0.95 across cells), so a stronger agent can overturn the result, and we report the rerun noise bands and the full measurement protocol.
Figures
Reference graph
Works this paper leans on
-
[2]
Defeating Nondeterminism in
He, Horace and. Defeating Nondeterminism in. 2025 , month = sep, howpublished =
2025
-
[6]
2025 , howpublished =
2025
-
[7]
Proposer-Agent-Evaluator (
Zhou, Yifei and Yang, Qianlan and Lin, Kaixiang and Bai, Min and Zhou, Xiong and Wang, Yu-Xiong and Levine, Sergey and Li, Erran , booktitle =. Proposer-Agent-Evaluator (. 2025 , eprint =
2025
-
[10]
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , year =
Magma: A Foundation Model for Multimodal AI Agents , author =. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , year =. doi:10.48550/arXiv.2502.13130 , url =. 2502.13130 , archivePrefix =
-
[11]
International Conference on Learning Representations , year =
Ferret-UI 2: Mastering Universal User Interface Understanding Across Platforms , author =. International Conference on Learning Representations , year =. doi:10.48550/arXiv.2410.18967 , url =. 2410.18967 , archivePrefix =
-
[13]
Findings of the Association for Computational Linguistics: ACL 2022 , pages =
Reframing Instructional Prompts to GPTk's Language , author =. Findings of the Association for Computational Linguistics: ACL 2022 , pages =. 2022 , doi =
2022
-
[16]
FocusAgent: Simple Yet Effective Ways of Trimming the Large Context of Web Agents , author =. 2025 , eprint =. doi:10.48550/arXiv.2510.03204 , url =
-
[17]
Avenir-Web: Human-Experience-Imitating Multimodal Web Agents with Mixture of Grounding Experts , author =. 2026 , eprint =. doi:10.48550/arXiv.2602.02468 , url =
Show all 131 references
-
[18]
2025 , eprint =
ModServe: Modality- and Stage-Aware Resource Disaggregation for Scalable Multimodal Model Serving , author =. 2025 , eprint =. doi:10.48550/arXiv.2502.00937 , url =
2025 doi
- [19]
-
[20]
2026 , eprint =
Graph2Eval: Automatic Multimodal Task Generation for Agents via Knowledge Graphs , author =. 2026 , eprint =. doi:10.48550/arXiv.2510.00507 , url =
2026 doi
-
[21]
2026 , journal=
MIRAGE: The Illusion of Visual Understanding , author=. 2026 , journal=
2026
-
[22]
2026 , journal =
Diagnosing Visual Ignorance in Vision-Language Models , author =. 2026 , journal =
2026
-
[23]
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , year=
What's in the Image? A Deep-Dive into the Vision of Vision Language Models , author=. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , year=. 2411.17491 , archivePrefix=
-
[24]
arXiv preprint arXiv:2407.21771 , year=
Paying more attention to image: A training-free method for alleviating hallucination in lvlms , author=. arXiv preprint arXiv:2407.21771 , year=
-
[25]
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , year=
XLRS-Bench: Could Your Multimodal LLMs Understand Extremely Large Ultra-High-Resolution Remote Sensing Imagery? , author=. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , year=
-
[26]
2025 , eprint =
AgentRewardBench: Evaluating Automatic Evaluations of Web Agent Trajectories , author =. 2025 , eprint =
2025
-
[27]
International Conference on Machine Learning (ICML) , year=
Plan-and-Act: A Scalable Framework for Enhancing LLM-based Web Agents , author=. International Conference on Machine Learning (ICML) , year=
-
[28]
International Conference on Machine Learning (ICML) , year=
ViLP: A Benchmark for Evaluating Visual Language Priors in VLMs , author=. International Conference on Machine Learning (ICML) , year=
-
[29]
Navigating the Digital World as Humans Do: Universal Visual Grounding for
Gou, Boyu and Wang, Ruohan and Zheng, Boyuan and others , booktitle=. Navigating the Digital World as Humans Do: Universal Visual Grounding for. 2025 , eprint=
2025
-
[30]
International Conference on Learning Representations (ICLR) , year=
WALT: Web Agents that Learn Tools , author=. International Conference on Learning Representations (ICLR) , year=
-
[31]
Transactions on Machine Learning Research , year=
Exposing Limitations of Language Model Agents in Sequential-Task Compositions on the Web , author=. Transactions on Machine Learning Research , year=
-
[32]
arXiv preprint arXiv:2510.15955 , year=
How Good Are LLMs at Processing Tool Outputs? , author=. arXiv preprint arXiv:2510.15955 , year=
-
[33]
The Thirteenth International Conference on Learning Representations , year=
AgentOccam: A Simple Yet Strong Baseline for LLM-Based Web Agents , author=. The Thirteenth International Conference on Learning Representations , year=. 2410.13825 , archivePrefix=
-
[34]
arXiv preprint arXiv:2409.13711 , year=
WebQuest: A Benchmark for Multimodal QA on Web Page Sequences , author=. arXiv preprint arXiv:2409.13711 , year=
-
[35]
arXiv preprint arXiv:2410.17236 , year=
Large Language Models Empowered Personalized Web Agents , author=. arXiv preprint arXiv:2410.17236 , year=
-
[36]
arXiv preprint arXiv:2510.15974 , year=
Limits of Emergent Reasoning of Large Language Models in Agentic Frameworks for Deterministic Games , author=. arXiv preprint arXiv:2510.15974 , year=
-
[37]
arXiv preprint arXiv:2403.07718 , year=
WorkArena: How Capable are Web Agents at Solving Common Knowledge Work Tasks? , author=. arXiv preprint arXiv:2403.07718 , year=
-
[38]
arXiv preprint arXiv:2407.13032 , year=
Agent-E: From Autonomous Web Navigation to Foundational Design Principles in Agentic Systems , author=. arXiv preprint arXiv:2407.13032 , year=
-
[39]
ResearchGate Preprint , year=
VisEscape: A Benchmark for Evaluating Exploration-driven Decision-making in Virtual Escape Rooms , author=. ResearchGate Preprint , year=
-
[40]
2024 , eprint=
Pan, Yichen and Kong, Dehan and Zhou, Sida and Cui, Cheng and Leng, Yifei and others , journal=. 2024 , eprint=
2024
-
[41]
Advances in Neural Information Processing Systems (NeurIPS) , year=
On the Effects of Data Scale on UI Control Agents , author=. Advances in Neural Information Processing Systems (NeurIPS) , year=
-
[42]
arXiv preprint arXiv:2603.07024 , year=
Enhancing Web Agents with a Hierarchical Memory Tree , author=. arXiv preprint arXiv:2603.07024 , year=
-
[43]
arXiv preprint arXiv:2603.28387 , year=
The Scaffold Effect: How Prompt Framing Drives Apparent Multimodal Gains in Clinical VLM Evaluation , author=. arXiv preprint arXiv:2603.28387 , year=
-
[44]
Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal
Tong, Shengbang and Liu, Zhuang and Zhai, Yuexiang and others , booktitle=. Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal. 2024 , eprint=
2024
-
[45]
arXiv preprint arXiv:2604.12424 , year=
Decoding by Perturbation: Mitigating MLLM Hallucinations via Dynamic Textual Perturbation , author=. arXiv preprint arXiv:2604.12424 , year=
-
[46]
arXiv preprint arXiv:2603.27187 , year=
Omni-Modal Dissonance Benchmark: Systematically Breaking Modality Consensus to Probe Robustness and Calibrated Abstention , author=. arXiv preprint arXiv:2603.27187 , year=
-
[47]
arXiv preprint arXiv:2508.07999 , year=
Widesearch: benchmarking agentic broad info-seeking , author=. arXiv preprint arXiv:2508.07999 , year=
-
[48]
arXiv preprint arXiv:2509.09674 , year=
RFTF , author=. arXiv preprint arXiv:2509.09674 , year=
-
[49]
arXiv preprint arXiv:2604.01438 , year=
ClawSafety: "Safe" LLMs, Unsafe Agents , author=. arXiv preprint arXiv:2604.01438 , year=
-
[50]
arXiv preprint arXiv:2504.07951 , year=
Scaling Laws for Native Multimodal Models , author=. arXiv preprint arXiv:2504.07951 , year=
-
[51]
International Conference on Machine Learning (ICML) , year=
Probing Visual Language Priors in VLMs , author=. International Conference on Machine Learning (ICML) , year=
-
[52]
arXiv preprint arXiv:2512.17875 , year=
Visually Prompted Benchmarks Are Surprisingly Fragile , author=. arXiv preprint arXiv:2512.17875 , year=
-
[53]
arXiv preprint arXiv:2601.04073 , year=
Analyzing Reasoning Consistency in Large Multimodal Models under Cross-Modal Conflicts , author=. arXiv preprint arXiv:2601.04073 , year=
-
[55]
2606.10423 , archivePrefix =
Hwang, Jayoo and Zhang, Xiaowen and Padwal, Vedant , year =. 2606.10423 , archivePrefix =
-
[56]
arXiv preprint arXiv:2509.16087 , year=
See&Trek: Training-Free Spatial Prompting for Multimodal Large Language Model , author=. arXiv preprint arXiv:2509.16087 , year=
-
[57]
International Conference on Learning Representations (ICLR) Under Review , annote=
Inference-Optimal Token Compression for Vision Language Models , author=. International Conference on Learning Representations (ICLR) Under Review , annote=
-
[58]
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , year =
HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models , author =. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , year =. 2310.14566 , archivePrefix =
-
[59]
and Chun, Wei-Chiu and Krishna, Ranjay , booktitle =
Fu, Xingyu and Hu, Yushi and Li, Bangzheng and Feng, Yu and Wang, Haoyu and Lin, Xudong and Roth, Dan and Smith, Noah A. and Chun, Wei-Chiu and Krishna, Ranjay , booktitle =. 2024 , pages =. 2404.12390 , archivePrefix =
2024 arXiv
-
[60]
Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP) , pages =
Evaluating Object Hallucination in Large Vision-Language Models , author =. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP) , pages =. 2023 , publisher =. doi:10.18653/v1/2023.emnlp-main.20 , eprint =
2023 doi
-
[61]
Breaking Common Sense:
Bitton-Guetta, Nitzan and Bitton, Yonatan and Hessel, Jack and Schmidt, Ludwig and Elovici, Yuval and Stanovsky, Gabriel and Schwartz, Roy , booktitle =. Breaking Common Sense:. 2023 , pages =. 2303.07274 , archivePrefix =
2023 arXiv
-
[62]
2025 , eprint =
Devils in Middle Layers of Large Vision-Language Models: Interpreting, Detecting, and Mitigating Object Hallucinations via Attention Lens , author =. 2025 , eprint =
2025
-
[63]
2026 , month = may, eprint =
Tool Calling is Linearly Readable and Steerable in Language Models , author =. 2026 , month = may, eprint =
2026
-
[64]
2026 , month = may, eprint =
Flexible Routing via Uncertainty Decomposition , author =. 2026 , month = may, eprint =
2026
-
[65]
Inference Time Causal Probing in
Khorasani, Sadegh and Salehkaleybar, Saber and Kiyavash, Negar and Grossglauser, Matthias , year =. Inference Time Causal Probing in. 2605.07631 , archivePrefix =
-
[66]
2026 , month = may, eprint =
Position: Mechanistic Interpretability Must Disclose Identification Assumptions for Causal Claims , author =. 2026 , month = may, eprint =
2026
-
[67]
2026 , month = may, eprint =
Where's the Plan? Locating Latent Planning in Language Models with Lightweight Mechanistic Interventions , author =. 2026 , month = may, eprint =
2026
-
[68]
Fayyaz, Mohsen and others , year =
-
[69]
Interpretability in the Wild: a Circuit for Indirect Object Identification in
Wang, Kevin and Variengien, Alexandre and Conmy, Arthur and Shlegeris, Buck and Steinhardt, Jacob , booktitle =. Interpretability in the Wild: a Circuit for Indirect Object Identification in. 2023 , eprint =
2023
-
[70]
2024 , eprint =
Best Practices for Activation Patching in Language Models: Metrics and Methods , author =. 2024 , eprint =
2024
-
[71]
Scandinavian Journal of Statistics , volume =
A Simple Sequentially Rejective Multiple Test Procedure , author =. Scandinavian Journal of Statistics , volume =. 1979 , publisher =
1979
-
[72]
and Steinhardt, Jacob , journal =
Lipton, Zachary C. and Steinhardt, Jacob , journal =. Troubling Trends in Machine Learning Scholarship: Some. 2019 , publisher =
2019
-
[73]
2024 , howpublished =
2024
-
[74]
2025 , eprint =
Seeing but Not Believing: Probing the Disconnect Between Visual Attention and Answer Correctness in Vision-Language Models , author =. 2025 , eprint =
2025
-
[75]
Cost-Aware Contrastive Routing for
Shirkavand, Reza and Gao, Shangqian and Yu, Peiran and Huang, Heng , year =. Cost-Aware Contrastive Routing for. 2508.12491 , archivePrefix =
-
[76]
Same Task, Different Circuits: Disentangling Modality-Specific Mechanisms in
Nikankin, Yaniv and Arad, Dana and Gandelsman, Yossi and Belinkov, Yonatan , year =. Same Task, Different Circuits: Disentangling Modality-Specific Mechanisms in. 2506.09047 , archivePrefix =
-
[77]
Understanding
Tao, Xingjian and Wang, Yiwei and Cai, Yujun and Yang, Zhicheng and Tang, Jing , year =. Understanding. 2506.15425 , archivePrefix =
-
[78]
Schiepanski, Thassilo M. and Pi. Beyond Pixels: Exploring. 2025 , eprint =
2025
-
[79]
2025 , eprint =
Runaway is Ashamed, But Helpful: On the Early-Exit Behavior of Large Language Model-based Agents in Embodied Environments , author =. 2025 , eprint =
2025
-
[80]
2026 , eprint =
Judge Circuits , author =. 2026 , eprint =
2026
-
[81]
2605.24785 , archivePrefix =
Li, Yubo and Miao, Yidi and Shen, Yuntian and Liu, Yuxin , year =. 2605.24785 , archivePrefix =
-
[82]
2605.11212 , archivePrefix =
Abaskohi, Amirhossein and He, Yuhang and West, Peter and Carenini, Giuseppe and Chawla, Pranit and Vineet, Vibhav , year =. 2605.11212 , archivePrefix =
-
[83]
Dynamic Mixed-Precision Routing for Efficient Multi-step
Li, Yuanzhe and Deng, Jianing and Hu, Jingtong and Chen, Tianlong and Wang, Song and Yang, Huanrui , year =. Dynamic Mixed-Precision Routing for Efficient Multi-step. 2602.02711 , archivePrefix =
-
[84]
2601.02439 , archivePrefix =
Bai, Hao and Taymanov, Alexey and Zhang, Tong and Kumar, Aviral and others , year =. 2601.02439 , archivePrefix =
-
[85]
2601.04126 , archivePrefix =
Zhang, Ziyun and Wang, Zezhou and Zhang, Xiaoyi and Guo, Zongyu and others , year =. 2601.04126 , archivePrefix =
-
[86]
2602.14296 , archivePrefix =
Wu, Yifan and Peng, Yiran and Chen, Yiyu and Ruan, Jianhao and others , year =. 2602.14296 , archivePrefix =
-
[87]
2511.20766 , archivePrefix =
Ullrich, Karen and Su, Jingtong and Shi, Claudia and Subramonian, Arjun and others , year =. 2511.20766 , archivePrefix =
-
[88]
2605.26114 , archivePrefix =
Wu, Dingbang and Hao, Rui and Wang, Haiyang and Wu, Shuzhe and others , year =. 2605.26114 , archivePrefix =
-
[90]
Proceedings of the 34th International Conference on Machine Learning (ICML) , year =
On Calibration of Modern Neural Networks , author =. Proceedings of the 34th International Conference on Machine Learning (ICML) , year =. 1706.04599 , archivePrefix =
-
[91]
International Conference on Learning Representations (ICLR) , year =
Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generation , author =. International Conference on Learning Representations (ICLR) , year =. 2302.09664 , archivePrefix =
-
[92]
2022 , eprint =
Language Models (Mostly) Know What They Know , author =. 2022 , eprint =
2022
-
[93]
2401.13919 , archivePrefix =
He, Hongliang and Yao, Wenlin and Ma, Kaixin and others , year =. 2401.13919 , archivePrefix =
-
[94]
Can You Trust Your Model's Uncertainty?
Ovadia, Yaniv and Fertig, Emily and Ren, Jie and others , booktitle =. Can You Trust Your Model's Uncertainty?. 2019 , eprint =
2019
-
[95]
2024 , eprint =
Language Model Cascades: Token-Level Uncertainty and Beyond , author =. 2024 , eprint =
2024
-
[96]
Ding, Dujian and Mallick, Ankur and Wang, Chi and others , year =. Hybrid. 2404.14618 , archivePrefix =
-
[97]
2025 , eprint =
WebRouter: Query-specific Router via Variational Information Bottleneck for Cost-sensitive Web Agent , author =. 2025 , eprint =
2025
-
[98]
Energy and Policy Considerations for Deep Learning in
Strubell, Emma and Ganesh, Ananya and McCallum, Andrew , booktitle =. Energy and Policy Considerations for Deep Learning in. 2019 , eprint =
2019
-
[99]
and Borm, George F
IntHout, Joanna and Ioannidis, John P.A. and Borm, George F. , journal =. The Hartung-Knapp-Sidik-Jonkman method for random effects meta-analysis is straightforward and considerably outperforms the standard. 2014 , doi =
2014
-
[100]
Research Synthesis Methods , volume =
Methods to estimate the between-study variance and its uncertainty in meta-analysis , author =. Research Synthesis Methods , volume =. 2016 , doi =
2016
-
[101]
BMC Medical Research Methodology , volume =
Hartung-Knapp-Sidik-Jonkman approach and its modification for random-effects meta-analysis with few studies , author =. BMC Medical Research Methodology , volume =. 2015 , doi =. 1508.01227 , archivePrefix =
2015 arXiv
-
[103]
2026 , eprint =
Read More, Think More: Revisiting Observation Reduction for Web Agents , author =. 2026 , eprint =
2026
-
[104]
2026 , eprint =
Web Agents Should Adopt the Plan-Then-Execute Paradigm , author =. 2026 , eprint =
2026
-
[105]
2026 , eprint =
Revisiting Observation Reduction for Web Agents: Comprehensive Evaluation with a Lightweight Framework , author =. 2026 , eprint =
2026
-
[106]
2026 , eprint =
Focus When Necessary: Adaptive Routing and Collaborative Grounding for Training-Free Visual Grounding , author =. 2026 , eprint =
2026
-
[107]
2026 , eprint =
Learning Agent Routing From Early Experience , author =. 2026 , eprint =
2026
-
[108]
2026 , eprint =
Adaptive Re-Ranking , author =. 2026 , eprint =
2026
-
[109]
2026 , eprint =
Signal-Driven Observation for Long-Horizon Web Agents , author =. 2026 , eprint =
2026
-
[110]
2026 , eprint =
Detect Before You Leap: Mirage Detection in Vision-Language Models , author =. 2026 , eprint =
2026
-
[111]
2026 , eprint =
Sema: Semantic Transport for Real-Time Multimodal Agents , author =. 2026 , eprint =
2026
-
[112]
2025 , eprint =
Building Browser Agents: Architecture, Security, and Practical Solutions , author =. 2025 , eprint =
2025
-
[113]
2026 , eprint =
How Vision Becomes Language: A Layer-wise Information-Theoretic Analysis of Multimodal Reasoning , author =. 2026 , eprint =
2026
-
[114]
Just how sure are you? Improving Verbalized Uncertainty Calibration in Medical
Senoglu, Eren and Toschi, Federico and Brunello, Nicolo and Sassella, Andrea and Carman, Mark James , year =. Just how sure are you? Improving Verbalized Uncertainty Calibration in Medical. 2606.27023 , archivePrefix =
-
[115]
Li, Yanhang and Fan, Zhichao and Zhuang, Zexin , year =. When. 2606.22864 , archivePrefix =
-
[116]
The American Statistician , volume =
Approximate Is Better than ``Exact'' for Interval Estimation of Binomial Proportions , author =. The American Statistician , volume =. 1998 , publisher =. doi:10.1080/00031305.1998.10480550 , annote =
1998 arXiv
-
[117]
Statistics in Medicine , volume =
Quantifying Heterogeneity in a Meta-Analysis , author =. Statistics in Medicine , volume =. 2002 , doi =
2002
-
[118]
Journal of Machine Learning Research , volume =
On Over-Fitting in Model Selection and Subsequent Selection Bias in Performance Evaluation , author =. Journal of Machine Learning Research , volume =. 2010 , annote =
2010
-
[119]
2026 , eprint =
Are Online Skill and Memory Modules Always Worth Their Tokens? A Budget-Constrained Study of Web Agents , author =. 2026 , eprint =
2026
-
[120]
2026 , eprint =
MolmoWeb: Open Visual Web Agent and Open Data for the Open Web , author =. 2026 , eprint =
2026
-
[121]
2026 , eprint =
Benchmarking the Benchmarks: A Validity Audit of Tool-Calling Evaluation , author =. 2026 , eprint =
2026
-
[122]
Cawley and Nicola L
Gavin C. Cawley and Nicola L. C. Talbot. 2010. On over-fitting in model selection and subsequent selection bias in performance evaluation. Journal of Machine Learning Research, 11:2079--2107
2010
- [123]
- [124]
-
[125]
Amine El hattami , Megh Thakkar, Nicolas Chapados, and Christopher Pal. 2025. https://openreview.net/forum?id=94tlGxmqkN WebArena Verified : Reliable evaluation for web agents . Poster, SEA Workshop @ NeurIPS 2025 (non-archival by authors' choice)
2025
-
[126]
Boyu Gou, Ruohan Wang, Boyuan Zheng, et al. 2025. https://arxiv.org/abs/2410.05243 Navigating the digital world as humans do: Universal visual grounding for GUI agents . In International Conference on Learning Representations (ICLR)
2025 arXiv
-
[127]
Neha Gupta, Harikrishna Narasimhan, Wittawat Jitkrittum, et al. 2024. https://arxiv.org/abs/2404.10136 Language model cascades: Token-level uncertainty and beyond . Preprint, arXiv:2404.10136
2024 arXiv
-
[128]
Laradji, Spandana Gella, and Nicolas Gontier
Sina Hajimiri, Masih Aminbeidokhti, Jose Dolz, Ismail Ben Ayed, Issam H. Laradji, Spandana Gella, and Nicolas Gontier. 2026. https://arxiv.org/abs/2606.15017 Are online skill and memory modules always worth their tokens? a budget-constrained study of web agents . Preprint, arX...
2026
-
[129]
Horace He and Thinking Machines Lab . 2025. https://thinkingmachines.ai/blog/defeating-nondeterminism-in-llm-inference/ Defeating nondeterminism in LLM inference . Thinking Machines Lab: Connectionism (blog post, 2025-09-10)
2025
-
[130]
Sture Holm. 1979. A simple sequentially rejective multiple test procedure. Scandinavian Journal of Statistics, 6(2):65--70
1979
-
[131]
Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, et al. 2022. https://arxiv.org/abs/2207.05221 Language models (mostly) know what they know . Preprint, arXiv:2207.05221
2022 arXiv
- [132]
-
[133]
Tao Li, Jinlong Hu, Yang Wang, Junfeng Liu, and Xuejun Liu. 2025. https://arxiv.org/abs/2510.11221 Webrouter: Query-specific router via variational information bottleneck for cost-sensitive web agent . Preprint, arXiv:2510.11221
2025
-
[134]
Xunzhuo Liu, Bowei He, Xue Liu, Andy Luo, Haichen Zhang, and Huamin Chen. 2026. Adaptive vision-language model routing for computer use agents. arXiv preprint arXiv:2603.12823
2026
-
[135]
Pal, and Siva Reddy
Xing Han L \`u , Amirhossein Kazemnejad, Nicholas Meade, Arkil Patel, Dongchan Shin, Alejandra Zambrano, Karolina Sta \'n czak, Peter Shaw, Christopher J. Pal, and Siva Reddy. 2025. https://arxiv.org/abs/2504.08942 Agentrewardbench: Evaluating automatic evaluations of web agen...
2025
-
[136]
Kelleher
Yasmin Moslem and John D. Kelleher. 2026. https://arxiv.org/abs/2603.04445 Dynamic model routing and cascading for efficient LLM inference: A survey . Preprint, arXiv:2603.04445
2026 arXiv
-
[137]
Gonzalez, M
Isaac Ong, Amjad Almahairi, Vincent Wu, Wei-Lin Chiang, Tianhao Wu, Joseph E. Gonzalez, M. Waleed Kadous, and Ion Stoica. 2025. https://doi.org/10.48550/arXiv.2406.18665 Routellm: Learning to route llms with preference data . In International Conference on Learning Representations
- [138]
-
[139]
Tianci Xue, Weijian Qi, Tianneng Shi, Chan Hee Song, et al. 2025. https://arxiv.org/abs/2504.01382 An illusion of progress? A ssessing the current state of web agents . Preprint, arXiv:2504.01382
2025
- [140]
-
[141]
Jiayi Yuan, Hao Li, Xinheng Ding, Wenya Xie, Yu-Jhe Li, Wentian Zhao, Kun Wan, Jing Shi, Xia Hu, and Zirui Liu. 2025. https://arxiv.org/abs/2506.09501 Understanding and mitigating numerical sources of nondeterminism in LLM inference . Preprint, arXiv:2506.09501
2025
- [142]
-
[143]
Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig
Shuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig. 2024. https://doi.org/10.48550/arXiv.2307.13854 Webarena: A realistic web environment for building autonomous agents . ...
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.