A rule-gated 7B agent with deterministic post-execution verification outperforms a direct-prompted 32B baseline on a reliability benchmark where half of correct responses are not SQL answers.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.AI 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Business Truth, not SQL Accuracy: A Rule-Gated 7B Analytics Agent Outperforms a Direct-Prompted 32B Baseline
A rule-gated 7B agent with deterministic post-execution verification outperforms a direct-prompted 32B baseline on a reliability benchmark where half of correct responses are not SQL answers.