Pith. sign in

hub

Webshop: Towards scalable real-world web interaction with grounded language agents

24 Pith papers cite this work. Polarity classification is still indexing.

24 Pith papers citing it

hub tools

citation-role summary

background 1 dataset 1

citation-polarity summary

years

2026 23 2025 1

representative citing papers

ScaleWoB: Guiding GUI Agents with Coding Agents via Large-Scale Environmental Synthesis

cs.AI · 2026-05-24 · unverdicted · novelty 7.0

ScaleWoB generates 100+ synthetic interactive GUI environments and 1000+ verifiable tasks as web pages, releasing a 120-task mobile benchmark where state-of-the-art agents achieve 27.92% success (17.82% on long-horizon tasks) versus 92.08% for humans, with synthetic results generalizing to real apps

Diagnosing Task Insensitivity in Language Agents

cs.AI · 2026-06-25 · unverdicted · novelty 6.0

The paper diagnoses task insensitivity in LLM agents as a cause of weak OOD generalization, links it to attention drift, and proposes Task-Perturbed NLL Optimization as a contrastive regularizer to improve task dependence.

citing papers explorer

Showing 24 of 24 citing papers.