CRAFTQA uses CodeSTEP to emit executable Python reasoning sequences and CRAFT to synthesize custom functions, yielding claimed gains on complex structured-data QA tasks.
is this text bolded?
7 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
roles
background 1polarities
background 1representative citing papers
Multi-agent AI system formalizes entire 500-page graduate algebraic combinatorics textbook into Lean, creating 130K lines of code in one week at human-expert cost.
Veritas detects out-of-bounds vulnerabilities in stripped binaries at 90% recall by grounding LLM reasoning in static witness-backed flows and runtime validation.
DPC selects correct text-to-SQL outputs by enforcing execution consistency between SQL and Python on an adversarially constructed minimal distinguishing database.
R&B-EnCoRe uses self-supervised importance-weighted variational inference to distill action-predictive reasoning datasets that improve VLA performance on manipulation, navigation, and driving tasks without external verifiers.
Talk-to-Your-Slides uses language-driven structured data manipulation with a hierarchical architecture to edit slides, reporting 34% faster processing, 34% better instruction fidelity, and 87% lower cost than GUI-based baselines on text-centric tasks while releasing the TSBench benchmark.
AI for mathematics is best described as a supervision ladder — final answers, programs, process rewards, proof-assistant kernels — culminating in verified-discovery workflows.
citing papers explorer
-
CRAFTQA: A Code-Driven Adaptive Framework for Complex Structured Data Reasoning
CRAFTQA uses CodeSTEP to emit executable Python reasoning sequences and CRAFT to synthesize custom functions, yielding claimed gains on complex structured-data QA tasks.
-
Automatic Textbook Formalization
Multi-agent AI system formalizes entire 500-page graduate algebraic combinatorics textbook into Lean, creating 130K lines of code in one week at human-expert cost.
-
Veritas: Grounding LLM Agents for Reliable Vulnerability Reasoning over Stripped Binaries
Veritas detects out-of-bounds vulnerabilities in stripped binaries at 90% recall by grounding LLM reasoning in static witness-backed flows and runtime validation.
-
DPC: Training-Free Text-to-SQL Candidate Selection via Dual-Paradigm Consistency
DPC selects correct text-to-SQL outputs by enforcing execution consistency between SQL and Python on an adversarially constructed minimal distinguishing database.
-
Self-Supervised Bootstrapping of Action-Predictive Embodied Reasoning
R&B-EnCoRe uses self-supervised importance-weighted variational inference to distill action-predictive reasoning datasets that improve VLA performance on manipulation, navigation, and driving tasks without external verifiers.
-
Talk to Your Slides: High-Efficiency Slide Editing via Language-Driven Structured Data Manipulation
Talk-to-Your-Slides uses language-driven structured data manipulation with a hierarchical architecture to edit slides, reporting 34% faster processing, 34% better instruction fidelity, and 87% lower cost than GUI-based baselines on text-centric tasks while releasing the TSBench benchmark.
-
Artificial Intelligence for Mathematical Reasoning: An Integrated Survey of Language Models, Neuro-symbolic Systems, and Verified Discovery
AI for mathematics is best described as a supervision ladder — final answers, programs, process rewards, proof-assistant kernels — culminating in verified-discovery workflows.