Pith. sign in

REVIEW 20 cited by

Autonomous Evaluation and Refinement of Digital Agents

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.06474 v3 pith:UGQICNME submitted 2024-04-09 cs.AI

classification cs.AI
keywords agentsperformanceevaluationimprovecontroldevicedigitalevaluators
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We show that domain-general automatic evaluators can significantly improve the performance of agents for web navigation and device control. We experiment with multiple evaluation models that trade off between inference cost, modularity of design, and accuracy. We validate the performance of these models in several popular benchmarks for digital agents, finding between 74.4 and 92.9% agreement with oracle evaluation metrics. Finally, we use these evaluators to improve the performance of existing agents via fine-tuning and inference-time guidance. Without any additional supervision, we improve state-of-the-art performance by 29% on the popular benchmark WebArena, and achieve around 75% relative improvement in device control settings.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 20 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. StepReflect: Structured UI Transition Reflection for Mobile GUI Agents

    cs.AI 2026-08 conditional novelty 6.0 of 10

    A locally deployed 8B reflector improves task success on three of four mobile GUI agent frameworks and beats GPT-5.2 by 11.83 points on offline AndroidWorld transition accuracy.

  2. WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    WebRetriever is a benchmark of 800 websites and 1,550 tasks with an automated evaluator (NavEval) achieving ~91–97% human agreement, showing current web agents succeed on only 11–37% of realistic tasks across three ev...

  3. Estimating the Empowerment of Language Model Agents

    cs.AI 2025-09 conditional novelty 6.0 of 10

    EELMA estimates the mutual information between an LM agent's actions and future text states, and this 'empowerment' is shown to correlate with task performance across toy games and WebArena.

  4. Reinforcement Learning for Machine Learning Engineering Agents

    cs.LG 2025-09 conditional novelty 6.0 of 10

    RL-trained Qwen2.5-3B outperforms prompted Claude-3.5-Sonnet and GPT-4o on 12 MLEBench tasks by an average of 22% and 24%, using two targeted RL modifications.

  5. MobiAgent: A Systematic Framework for Customizable Mobile Agents

    cs.MA 2025-08 conditional novelty 6.0 of 10

    A full-stack mobile agent framework reports state-of-the-art task completion on its own DAG-based benchmark and 2-3x speedups from replaying recorded trajectories.

  6. CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A decoupled planner-executor GUI agent, trained by per-app reinforcement learning followed by specialist-to-generalist distillation, lifts ScienceBoard success from about 7.6% to 21.0% average and 40% pass@8.

  7. SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience

    cs.AI 2025-08 conditional novelty 6.0 of 10

    A self-evolving computer-use agent trained with full-trajectory state judging and curriculum task generation goes from 11.3% to 34.5% average success on five OSWorld apps, and a specialist-to-generalist variant beats ...

  8. DipLLM: Fine-Tuning LLM for Strategic Decision-making in Diplomacy

    cs.AI 2025-06 conditional novelty 6.0 of 10

    A fine-tuned LLM with a unit-by-unit action decomposition outperforms prior Diplomacy agents while using far less training data.

  9. Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Scaling the number of interaction steps, trained via a curriculum over rollout horizon, improves web-agent task success and outperforms scaling per-step reasoning under fixed token budgets.

  10. Self-Challenging Language Model Agents

    cs.AI 2025-06 conditional novelty 6.0 of 10

    A language model agent can generate its own verifiable training tasks and improve its tool-use success rate by about 2x without human-annotated data.

  11. ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making

    cs.RO 2025-05 conditional novelty 6.0 of 10

    ManiTaskGen automatically generates diverse, feasible mobile manipulation tasks from any input scene, and uses them to benchmark and improve vision-language robot agents.

  12. ProgRM: Build Better GUI Agents with Progress Rewards

    cs.AI 2025-05 conditional novelty 6.0 of 10

    ProgRM, a per-step progress reward model trained with LCS-based self-annotated labels, improves RL-trained GUI agent success rates on WikiHow relative to outcome reward models.

  13. Evaluation Agent: Efficient and Promptable Evaluation Framework for Visual Generative Models

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Evaluation Agent is an LLM-agent framework that evaluates visual generative models with a handful of samples per round, claiming a 10x time reduction while keeping conclusions within one tier of full-benchmark results...

  14. HiLSVA: Design and Evaluation of a Human-in-the-Loop Agentic System for Scientific Visualization

    cs.HC 2026-06 unverdicted novelty 5.0 of 10

    HiLSVA shows that a human-in-the-loop LLM agent system can help novices and experts complete scientific visualization tasks, while human oversight adds measurable execution time.

  15. Branch-and-Browse: Efficient and Controllable Web Exploration with Tree-Structured Reasoning and Action Memory

    cs.AI 2025-10 conditional novelty 5.0 of 10

    Branch-and-Browse, a tree-structured web agent with page-level action memory, reports 35.8% success and up to 40.4% less time than the cited Tree Search baseline on WebArena.

  16. WorkForceAgent-R1: Incentivizing Reasoning Capability in LLM-based Web Agents via Reinforcement Learning

    cs.CL 2025-05 conditional novelty 5.0 of 10

    R1-style reinforcement learning on single-step web actions lifts open-source agents above gpt-4o on WorkArena while avoiding the reward hacking seen with dense rewards.

  17. Digi-Q: Learning Q-Value Functions for Training Device-Control Agents

    cs.LG 2025-02 conditional novelty 5.0 of 10

    An offline RL method learns a Q-function from frozen VLM features and extracts a device-control policy by imitating the best of several actions ranked by that Q-function.

  18. PAFFA: Premeditated Actions For Fast Agents

    cs.AI 2024-12 reject novelty 5.0 of 10

    PAFFA caches LLM-generated web interaction scripts into an action library so that runtime agents only retrieve and parameterize a pre-written API rather than parsing HTML step by step.

  19. Reflection-Based Memory For Web navigation Agents

    cs.AI 2025-06 conditional novelty 4.0 of 10

    Reflection-Augmented Planning (ReAP) retrieves short self-reflections from past web navigation tasks and lifts held-out task success by 11 points on WebArena.

  20. Generalist Virtual Agents: A Survey on Autonomous Agents Across Digital Platforms

    cs.MA 2024-11 conditional novelty 3.0 of 10

    A survey that proposes the Generalist Virtual Agent concept and taxonomies for agent environments, tasks, perceptions, actions, models, and evaluation, concluding that real-world-like environments favor human-like int...

Pith tools