OlymMATH is a 350-problem Olympiad math benchmark combining bilingual natural-language evaluation with Lean 4 formal verification to test LLM reasoning.
A survey of large language models
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
verdicts
UNVERDICTED 2representative citing papers
Keystone neurons are a stable, sparse, pretraining-established subset in Transformers; their removal collapses behavior and updating only them yields comparable task gains to full fine-tuning.
citing papers explorer
-
Challenging the Boundaries of Reasoning: An Olympiad-Level Math Benchmark for Large Language Models
OlymMATH is a 350-problem Olympiad math benchmark combining bilingual natural-language evaluation with Lean 4 formal verification to test LLM reasoning.
-
Tiny Brains, Giant Impact: Uncovering the Keystone Neurons of LLM with Just a Few Prompts
Keystone neurons are a stable, sparse, pretraining-established subset in Transformers; their removal collapses behavior and updating only them yields comparable task gains to full fine-tuning.