WebRetriever is a benchmark of 800 websites and 1,550 tasks with an automated evaluator (NavEval) achieving ~91–97% human agreement, showing current web agents succeed on only 11–37% of realistic tasks across three evaluation protocols.
Advances in neural information processing systems36, 46595–46623 (2023)
7 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
years
2026 7roles
dataset 1polarities
use dataset 1representative citing papers
S^2tory uses narratological theory and a Narrative Expert Agent to identify plot nuclei in movie scripts for high-fidelity summarization at 3.5x compression, with strong zero-shot generalization to books.
SurgCheck benchmark reveals that vision-language models for surgical VQA often depend on linguistic shortcuts rather than visual reasoning, shown by consistent performance drops on less-biased questions.
ANVIL automates analogy-based instructional animations for computer science by chaining LLM analogy generation, screenplay structuring, manim code production with repair, and mixed human-automated evaluations.
GEST-Engine turns game engines into zero-cost dense ground-truth video generators; GTASA reveals frozen video encoders fail inter-entity spatial relation probes.
View-PNDF detects and selectively fine-tunes view-specific neurons for consistent multi-view chest X-ray report generation, followed by LLM consolidation of reports.
SMART expands speculative decoding trees only when a node's marginal benefit-cost ratio exceeds current tree-level speedup, claiming ~15–20% extra wall-clock speedup without quality loss.
citing papers explorer
-
WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation
WebRetriever is a benchmark of 800 websites and 1,550 tasks with an automated evaluator (NavEval) achieving ~91–97% human agreement, showing current web agents succeed on only 11–37% of realistic tasks across three evaluation protocols.
-
S^2tory: Story Spine Distillation for Movie Script Summarization
S^2tory uses narratological theory and a Narrative Expert Agent to identify plot nuclei in movie scripts for high-fidelity summarization at 3.5x compression, with strong zero-shot generalization to books.
-
SurgCheck: Do Vision-Language Models Really Look at Images in Surgical VQA?
SurgCheck benchmark reveals that vision-language models for surgical VQA often depend on linguistic shortcuts rather than visual reasoning, shown by consistent performance drops on less-biased questions.
-
ANVIL: Analogies and Videos for Lecturers
ANVIL automates analogy-based instructional animations for computer science by chaining LLM analogy generation, screenplay structuring, manim code production with repair, and mixed human-automated evaluations.
-
GTASA: Ground Truth Annotations for Spatiotemporal Analysis, Evaluation and Training of Video Models
GEST-Engine turns game engines into zero-cost dense ground-truth video generators; GTASA reveals frozen video encoders fail inter-entity spatial relation probes.
-
Seeing Through Multiple Views: Parameter-Efficient Fine-Tuning via Selective Neurons for Consistent Radiology Report Generation
View-PNDF detects and selectively fine-tunes view-specific neurons for consistent multi-view chest X-ray report generation, followed by LLM consolidation of reports.
-
SMART: When is it Actually Worth Expanding a Speculative Tree?
SMART expands speculative decoding trees only when a node's marginal benefit-cost ratio exceeds current tree-level speedup, claiming ~15–20% extra wall-clock speedup without quality loss.