Clarity generates benchmarks showing that leading NL2SQL systems degrade significantly under multi-faceted ambiguity and struggle to localize or resolve schema-level issues despite detecting ambiguity.
Title resolution pending
5 Pith papers cite this work. Polarity classification is still indexing.
fields
cs.CL 5years
2026 5verdicts
UNVERDICTED 5representative citing papers
PolySQL enables cross-dialect text-to-SQL evaluation via automated backend isomorphism through result normalization, achieving 100% coverage and revealing that accuracy drops 10.1% from SQLite to other dialects due mostly to logical errors.
APMPO boosts average Pass@1 scores on math reasoning benchmarks by 3 points over GRPO by using an adaptive power-mean policy objective and feedback-driven clipping bounds in RLVR training.
FREIA applies free energy principles and adaptive advantage shaping to unsupervised RL, outperforming baselines by 0.5-3.5 Pass@1 points on math reasoning with a 1.5B model.
A knowledge-aware Text-to-SQL framework constructs domain knowledge bases to generate synthetic data and enhance inference, claiming substantial gains on seven benchmarks especially in low-resource settings.
citing papers explorer
-
CLARITY: A Framework and Benchmark for Conversational Language Ambiguity and Unanswerability in Interactive NL2SQL Systems
Clarity generates benchmarks showing that leading NL2SQL systems degrade significantly under multi-faceted ambiguity and struggle to localize or resolve schema-level issues despite detecting ambiguity.
-
PolySQL: Scaling Text-to-SQL Evaluation Across SQL Dialects via Automated Backend Isomorphism
PolySQL enables cross-dialect text-to-SQL evaluation via automated backend isomorphism through result normalization, achieving 100% coverage and revealing that accuracy drops 10.1% from SQLite to other dialects due mostly to logical errors.
-
Adapt to Thrive! Adaptive Power-Mean Policy Optimization for Improved LLM Reasoning
APMPO boosts average Pass@1 scores on math reasoning benchmarks by 3 points over GRPO by using an adaptive power-mean policy objective and feedback-driven clipping bounds in RLVR training.
-
Free Energy-Driven Reinforcement Learning with Adaptive Advantage Shaping for Unsupervised Reasoning in LLMs
FREIA applies free energy principles and adaptive advantage shaping to unsupervised RL, outperforming baselines by 0.5-3.5 Pass@1 points on math reasoning with a 1.5B model.
-
Knowledge Distillation for Low-Resource Open-source Text-to-SQL Model
A knowledge-aware Text-to-SQL framework constructs domain knowledge bases to generate synthetic data and enhance inference, claiming substantial gains on seven benchmarks especially in low-resource settings.