ARTS improves automated scientific discovery by using reasoning LMs with test-time training to separate hypothesis merit from execution quality in tree search, achieving 15.3% relative gains on 22 MLGym and MLEBench tasks.
ThermoFlex Water Bottle
3 Pith papers cite this work. Polarity classification is still indexing.
fields
cs.AI 3years
2026 3verdicts
UNVERDICTED 3representative citing papers
DPTS shows cold-start bottlenecks at low budgets while SSDP exhibits frontier depletion, indicating fixed ToT strategies are inelastic across compute levels.
SLATE benchmark and Entropy-Guided Branching algorithm improve LLM agent success and efficiency on long-horizon tasks in large tool libraries.
citing papers explorer
-
Learning the ARTS of Search for Automated Discovery
ARTS improves automated scientific discovery by using reasoning LMs with test-time training to separate hypothesis merit from execution quality in tree search, achieving 15.3% relative gains on 22 MLGym and MLEBench tasks.
-
Beyond Fixed Budgets: Characterizing the Inelasticity and Limitations of Tree-of-Thought Reasoning Strategies
DPTS shows cold-start bottlenecks at low budgets while SSDP exhibits frontier depletion, indicating fixed ToT strategies are inelastic across compute levels.
-
Long-Horizon Plan Execution in Large Tool Spaces through Entropy-Guided Branching
SLATE benchmark and Entropy-Guided Branching algorithm improve LLM agent success and efficiency on long-horizon tasks in large tool libraries.