MathNet compiles 30,676 multilingual Olympiad problems with solutions into a benchmark showing top LLMs score 69–78% while embedding retrievers rarely find mathematically equivalent problems at rank 1.
Mathbert: A pre-trained model for mathematical formula understanding.arXiv preprint arXiv:2105.00377
6 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
roles
background 1polarities
background 1representative citing papers
EDU-CIRCUIT-HW reveals large latent recognition failures in MLLMs on real handwritten university STEM solutions, limiting auto-grading reliability, though hybrid human-AI routing of only 3.3% cases improves outcomes.
Retrieval-of-Thought organizes prior reasoning into a thought graph for retrieval and reward-guided recombination, reducing output tokens by up to 40% and latency by 82% while preserving accuracy on reasoning benchmarks.
ORION reports 77.74 Driving Score and 54.62% Success Rate on Bench2Drive, outperforming prior end-to-end methods by 14.28 DS and 19.61% SR through unified VQA and planning optimization.
Introduces GSM8K dataset and demonstrates that verifier-based selection of solutions from multiple candidates outperforms fine-tuning baselines on math word problems.
AI for math combines task-specific architectures and general foundation models to support research and advance AI reasoning capabilities.
citing papers explorer
-
MathNet: a Global Multimodal Benchmark for Mathematical Reasoning and Retrieval
MathNet compiles 30,676 multilingual Olympiad problems with solutions into a benchmark showing top LLMs score 69–78% while embedding retrievers rarely find mathematically equivalent problems at rank 1.
-
EDU-CIRCUIT-HW: Evaluating Multimodal Large Language Models on Real-World University-Level STEM Student Handwritten Solutions
EDU-CIRCUIT-HW reveals large latent recognition failures in MLLMs on real handwritten university STEM solutions, limiting auto-grading reliability, though hybrid human-AI routing of only 3.3% cases improves outcomes.
-
Retrieval-of-Thought: Efficient Reasoning via Reusing Thoughts
Retrieval-of-Thought organizes prior reasoning into a thought graph for retrieval and reward-guided recombination, reducing output tokens by up to 40% and latency by 82% while preserving accuracy on reasoning benchmarks.
-
ORION: A Holistic End-to-End Autonomous Driving Framework by Vision-Language Instructed Action Generation
ORION reports 77.74 Driving Score and 54.62% Success Rate on Bench2Drive, outperforming prior end-to-end methods by 14.28 DS and 19.61% SR through unified VQA and planning optimization.
-
Training Verifiers to Solve Math Word Problems
Introduces GSM8K dataset and demonstrates that verifier-based selection of solutions from multiple candidates outperforms fine-tuning baselines on math word problems.
-
AI for Mathematics: Progress, Challenges, and Prospects
AI for math combines task-specific architectures and general foundation models to support research and advance AI reasoning capabilities.