Explorer-Definer and Reflective Orchestrator harnesses raise DeepSeek V3.2 from 15.5% to 67.25% pass@2 on ARC-AGI-1 public eval at $0.25–$0.62 per task without ARC-specific training.
URL https://arxiv.org/abs/2512.11847
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
fields
cs.AI 2years
2026 2representative citing papers
Guided stochastic exploration with three label-free diagnostics lifts Sudoku-Extreme solve accuracy from 85.9% to 98.0% and flags misaligned guides on Maze-Hard.
citing papers explorer
-
Cost-Effective Agent Harnesses for Abstract Reasoning and Generalization on ARC-AGI-1
Explorer-Definer and Reflective Orchestrator harnesses raise DeepSeek V3.2 from 15.5% to 67.25% pass@2 on ARC-AGI-1 public eval at $0.25–$0.62 per task without ARC-specific training.
-
Boosting Inference with Guided Reasoning: Stochastic Exploration for Recursive Models
Guided stochastic exploration with three label-free diagnostics lifts Sudoku-Extreme solve accuracy from 85.9% to 98.0% and flags misaligned guides on Maze-Hard.