Pith. sign in

REVIEW 9 cited by

Large Language Models Can Self-Improve At Web Agent Tasks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.20309 v2 pith:WOXYTTAQ submitted 2024-05-30 cs.LG cs.AIcs.CL

classification cs.LGcs.AIcs.CL
keywords modelsagentagentsbenchmarkdatalanguagellmsnavigate
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Training models to act as agents that can effectively navigate and perform actions in a complex environment, such as a web browser, has typically been challenging due to lack of training data. Large language models (LLMs) have recently demonstrated some capability to navigate novel environments as agents in a zero-shot or few-shot fashion, purely guided by natural language instructions as prompts. Recent research has also demonstrated LLMs have the capability to exceed their base performance through self-improvement, i.e. fine-tuning on data generated by the model itself. In this work, we explore the extent to which LLMs can self-improve their performance as agents in long-horizon tasks in a complex environment using the WebArena benchmark. In WebArena, an agent must autonomously navigate and perform actions on web pages to achieve a specified objective. We explore fine-tuning on three distinct synthetic training data mixtures and achieve a 31\% improvement in task completion rate over the base model on the WebArena benchmark through a self-improvement procedure. We additionally contribute novel evaluation metrics for assessing the performance, robustness, capabilities, and quality of trajectories of our fine-tuned agent models to a greater degree than simple, aggregate-level benchmark scores currently used to measure self-improvement.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Gated-BEPO: Confidence-Gated Bellman Credit Assignment for Large Language Model Agents

    cs.AI 2026-08 conditional novelty 7.0 of 10

    Gated-BEPO derives step-level Bellman advantages from empirical rollout graphs and uses a confidence gate to mix them with episode-level credit, improving LLM agent success on WebShop, ALFWorld, and Sokoban.

  2. Skill Self-Play: Pushing the Frontier of LLM Capability with Co-Evolving Skills

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Skill Self-Play uses an evolving skill library to guide LLM self-training, improving tool-call and reasoning accuracy beyond unguided self-play on five model backbones.

  3. Morae: Proactively Pausing UI Agents for User Choices

    cs.HC 2025-08 conditional novelty 6.0 of 10

    Morae, a UI agent that proactively pauses at ambiguous decision points, helps blind and low-vision users complete more tasks and express preferences better than fully autonomous agents.

  4. Build the web for agents, not agents for the web

    cs.LG 2025-06 conditional novelty 6.0 of 10

    The paper proposes a paradigm shift: design a standardized Agentic Web Interface for AI agents, rather than adapting agents to human-facing websites.

  5. InSTA: Towards Internet-Scale Training For Agents

    cs.LG 2025-02 conditional novelty 6.0 of 10

    Automated LLM task generation, agent execution, and judge filtering at 150k-site scale lets a 1.7B model match much larger web agents.

  6. OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task Synthesis

    cs.AI 2024-12 conditional novelty 6.0 of 10

    OS-Genesis synthesizes GUI agent training trajectories by reversing task synthesis: it explores interfaces first, derives low- and high-level instructions from observed actions, and filters trajectories with a reward model.

  7. Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training

    cs.AI 2025-01 conditional novelty 5.0 of 10

    Agent-R iteratively fine-tunes language agents on trajectories that splice the agent's own failed prefix at a model-identified error step with a successful continuation, improving scores on WebShop, ScienceWorld, and ...

  8. Compute-Accuracy Pareto Frontiers for Open-Source Reasoning Large Language Models

    cs.CL 2025-12 conditional novelty 4.0 of 10

    Across five reasoning benchmarks, sparse MoE models dominate the accuracy-vs-FLOPs Pareto frontier, inference-compute gains saturate, and wrong answers systematically consume more compute than correct ones.

  9. Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges

    cs.AI 2025-07 reject novelty 1.0 of 10

    A broad survey of LLM alignment that catalogs objectives, benchmarks, SFT/RLHF/DPO methods, and safety challenges, without contributing new experimental or theoretical results.

Pith tools