A new benchmark with cognitive traps shows frontier deep research agents achieve only 13-16% acceptance on expert consulting tasks under combined verifier and rubric criteria.
SciAgent: Tool-augmented language models for scientific reasoning.arXiv preprint arXiv:2402.11451
4 Pith papers cite this work, alongside 2 external citations. Polarity classification is still indexing.
years
2026 4representative citing papers
SCICONVBENCH is a new benchmark evaluating LLMs on multi-turn disambiguation and inconsistency resolution for task formulation in computational science, with frontier models reaching only 52.7% success on fluid mechanics disambiguation cases.
SciHorizon-DataEVA is a multi-agent system that applies Sci-TQA2 principles across four dimensions to assess AI-readiness of heterogeneous scientific data via dynamic profiling and self-correcting evaluation workflows.
citing papers explorer
-
Evaluating Deep Research Agents on Expert Consulting Work: A Benchmark with Verifiers, Rubrics, and Cognitive Traps
A new benchmark with cognitive traps shows frontier deep research agents achieve only 13-16% acceptance on expert consulting tasks under combined verifier and rubric criteria.
-
SCICONVBENCH: Benchmarking LLMs on Multi-Turn Clarification for Task Formulation in Computational Science
SCICONVBENCH is a new benchmark evaluating LLMs on multi-turn disambiguation and inconsistency resolution for task formulation in computational science, with frontier models reaching only 52.7% success on fluid mechanics disambiguation cases.
-
SciHorizon-DataEVA: An Agentic System for AI-Readiness Evaluation of Heterogeneous Scientific Data
SciHorizon-DataEVA is a multi-agent system that applies Sci-TQA2 principles across four dimensions to assess AI-readiness of heterogeneous scientific data via dynamic profiling and self-correcting evaluation workflows.
- PDE-Agents: An LLM-Orchestrated Multi-Agent Framework for Automated Finite Element Simulations with Knowledge Graph-Augmented Reasoning