Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-17T00:44:01.658214Z
Paper Citation Record · LEDGER
As of 23 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 100 inbound Pith citation observations for arXiv:2411.04872.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-17T00:44:01.658214Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T12:39:57.153685Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-10T20:47:34.579097Z
32 of 32 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 39460a96-c5fa-4602-b86c-f4e3b3875d4a · outbound
FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Unresolved cited work
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation cf80a1b0-5933-41af-aafc-d8093549bd21 · outbound
FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Unresolved cited work
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 5309dc1d-68ff-4f05-ade1-2a7ea1675038 · outbound
FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Advances in neural information processing systems , volume=
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation be4f333b-bea8-4592-9273-32be998c95d4 · outbound
FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation c079aed5-3f1b-494a-bae6-bf443fe82477 · outbound
FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 9f7f2aeb-ed36-4447-958b-2a61098fbe23 · outbound
FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Unresolved cited work
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 2c1a2fb4-fd83-4128-81a3-58261ad21fec · outbound
FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Unresolved cited work
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 3c77b4b2-5b20-4d2f-b613-2ebfc97e2dc1 · outbound
FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Unresolved cited work
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 2da17ad2-2d52-4627-b3aa-8bea89bdab7b · outbound
FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Nature , publisher =
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation c868446b-73ee-4711-b2a8-fe6e8e2c5c9b · outbound
FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Unresolved cited work
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ac70278b-34f8-47ab-a9a6-f505cb231d04 · outbound
FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Unresolved cited work
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 183977ef-4d2a-433b-bb34-1c1af7309b15 · outbound
FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation c097ccc9-30c1-4eb5-8923-fbb295536bcd · outbound
FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Unresolved cited work
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 5aa33541-7d41-4731-86e0-483e1659e6c2 · outbound
FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Nature , publisher =
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 20169156-23d9-43ac-a66f-e30a1988562f · outbound
FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Unresolved cited work
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 49f4966f-5311-4f36-906f-a78d8ef8a620 · outbound
FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Unresolved cited work
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation cb113683-04c4-4100-9fa5-7dcd2199b93e · outbound
FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation f5efe89a-38c1-4cda-b0d3-d5e426d94940 · outbound
FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Unresolved cited work
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ecb2761b-ec7c-4459-9f3f-362fbc8cf68b · outbound
FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Unresolved cited work
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 2ccac011-6135-4a9f-bc18-d28ad5d74592 · outbound
FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Unresolved cited work
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 41001959-cce5-4d10-bb2c-844d83df24ba · outbound
FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Unresolved cited work
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation a5280095-894f-4cae-bfc9-9b3fc95a88b1 · outbound
FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Unresolved cited work
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 776e030a-3b12-4862-92ad-422680b006b0 · outbound
FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Unresolved cited work
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 99a5fd29-758b-437d-bd9a-248f998a0b59 · outbound
FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Unresolved cited work
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation f22bb72c-833d-4472-b5fa-c701bd9a23e9 · outbound
FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 07450d21-c93d-4329-b626-a3142cd1aa79 · outbound
FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Measuring Massive Multitask Language Understanding
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 85c98fc7-0d24-46bb-98ab-dbb71bc6ed53 · outbound
FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Unresolved cited work
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 18eade63-0879-4ee7-99d4-1a3908f22e33 · outbound
FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Unresolved cited work
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 57c6425c-6289-4531-b774-b58208d994f1 · outbound
FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Equality of orders of a set of integers modulo a prime
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 473cd1cc-18a3-4f7b-9b8a-2b3dfddfdead · outbound
FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Curves over Finite Fields Attaining the Hasse-Weil Upper Bound
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 3a52ed73-4240-4524-95bb-a66541497403 · outbound
FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Unresolved cited work
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation becd89a7-e3cb-4383-a917-ace9399e0154 · outbound
FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Unresolved cited work
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 1f7f29f8-9d01-4dff-8441-24baffdc9061 · inbound
LMAct: A Benchmark for In-Context Imitation Learning with Long Multimodal Demonstrations FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6864c502-3dbc-4ecf-b948-b44014a6541f · inbound
HARP: A challenging human-annotated math reasoning benchmark FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6f1bc17-a186-43bb-873d-5cb9fde76fe0 · inbound
Formal Mathematical Reasoning: A New Frontier in AI FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea9ee7eb-ea6e-4860-bd8d-baaff5665698 · inbound
Entropy-Guided Attention for Private LLMs FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5d9363e-cc9a-47e2-9b9c-a616155cbe79 · inbound
Reasoning Language Models: A Blueprint FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76e21ff7-a526-4546-8e6c-a79829391353 · inbound
Not Every AI Problem is a Data Problem: We Should Be Intentional About Data Scaling FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72a19ef2-e197-4494-ac82-e0e76871d46d · inbound
Humanity's Last Exam FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 46e64c6b-c79f-4ee2-8b1a-3df5ed868a83 · inbound
SCP-116K: A High-Quality Problem-Solution Dataset and a Generalized Pipeline for Automated Extraction in the Higher Education Science Domain FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de65e791-2b36-486d-ba3d-ba130315b2eb · inbound
HardML: A Benchmark For Evaluating Data Science And Machine Learning knowledge and reasoning in AI FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51fef532-03b3-44af-b139-58e416fb4e51 · inbound
SocratiQ: A Generative AI-Powered Learning Companion for Personalized Education and Broader Accessibility FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02db8153-78d2-4c3d-a026-77c3dbe5cb33 · inbound
Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 910781e2-8a2c-416b-970d-8f6d5443e576 · inbound
Language Models Use Trigonometry to Do Addition FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63181b01-6028-476d-a761-cf84cc0c8618 · inbound
LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c25b27db-ccc0-4271-b56c-714685194501 · inbound
Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 213
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 736a5ebe-d0da-4fca-a4a8-b215d660a2b4 · inbound
FLIP Reasoning Challenge FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc01b0b9-f190-4b6e-9a82-191da7b33cd8 · inbound
In between myth and reality: AI for math -- a case study in category theory FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b38bd7a7-1962-4a49-a194-32d3f739b42e · inbound
Virology Capabilities Test (VCT): A Multimodal Virology Q&A Benchmark FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e5321e4-d706-49ce-91d0-77d03093590d · inbound
Lightweight Latent Verifiers for Efficient Meta-Generation Strategies FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d6f0768-e0fa-438f-a65e-c991e572e963 · inbound
Toward Generalizable Evaluation in the LLM Era: A Survey Beyond Benchmarks FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4c71659-396a-4a72-a822-71c465a159f5 · inbound
Automatic Legal Writing Evaluation of LLMs FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed9e84da-f5f8-4d37-9d15-a0b2c8c5f419 · inbound
Self-Ablating Transformers: More Interpretability, Less Sparsity FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ecb7b5e6-4fda-4e80-806c-fbdf6f736bf5 · inbound
Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6bd53163-92e2-4b33-b976-518b9fbc667c · inbound
R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75cd1dab-74a3-4db0-9f91-c0dee4dfeffc · inbound
Knowledge Augmented Complex Problem Solving with Large Language Models: A Survey FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a259721-15b6-4d1e-aeb6-8a0c23ecdf9f · inbound
R^3-VQA: "Read the Room" by Video Social Reasoning FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ce6c74f-ab92-42ad-9713-a112a38eb66c · inbound
Beyond Theorem Proving: Formulation, Framework and Benchmark for Formal Problem-Solving FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56ff5ae1-4f66-4896-9a39-67551a9bfff4 · inbound
AI Governance to Avoid Extinction: The Strategic Landscape and Actionable Research Questions FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 979e5bd1-3c6a-4f60-9573-ba7d0b2a1a11 · inbound
DeepMath-Creative: A Benchmark for Evaluating Mathematical Creativity of Large Language Models FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aeb6b72e-3edc-414f-8f95-0463b5ecea17 · inbound
HARDMath2: A Benchmark for Applied Mathematics Built by Students as Part of a Graduate Class FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed03e151-8c4a-4482-9086-79b502798798 · inbound
Children's Mental Models of AI Reasoning: Implications for AI Literacy Education FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4b6457d-799c-446b-a68d-314104f6df77 · inbound
Sudoku-Bench: Evaluating creative reasoning with Sudoku variants FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0f4e25b-0026-4598-9290-4d4737d2310e · inbound
Effective Reinforcement Learning for Reasoning in Language Models FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5da3b6b-af00-4747-ba97-d1321d3c4645 · inbound
CapBencher: Give Your LLM Benchmark a Built-in Alarm for Test-Set Overfitting FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4857d67b-0746-4bc9-ad44-c737b17a1ac4 · inbound
DFIR-Metric: A Benchmark Dataset for Evaluating Large Language Models in Digital Forensics and Incident Response FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 634a25d9-e61e-4648-8ad9-d15b3c1048fc · inbound
ASyMOB: Algebraic Symbolic Mathematical Operations Benchmark FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1fa19fda-577b-4940-af64-d789bac285dc · inbound
Using Reasoning Models to Generate Search Heuristics that Solve Open Instances of Combinatorial Design Problems FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd1ea47d-1177-45b1-a559-e261b22f9628 · inbound
Circuit Stability Characterizes Language Model Generalization FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83b0f156-b75d-481a-b848-ea1442e1e07c · inbound
Literature Review Of Multi-Agent Debate For Problem-Solving FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3c96934-e189-4598-8b84-cb32a3653808 · inbound
Breakpoint: Scalable evaluation of system-level reasoning in LLM code agents FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55da0497-55b0-488e-89c7-1624a72df77d · inbound
AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 435a24fe-d609-42ce-8590-fbf85262dc76 · inbound
RE-IMAGINE: Symbolic Benchmark Synthesis for Reasoning Evaluation FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fc04ec4-6484-4b78-9075-bd038d42f644 · inbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d40c2e4-2cad-4008-bc25-793a4dfba619 · inbound
Probing for Arithmetic Errors in Language Models FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a329d34-b70a-4bae-9e30-03f4eb0d9059 · inbound
FormulaOne: Measuring the Depth of Algorithmic Reasoning Beyond Competitive Programming FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 834de928-4ab6-4b41-a962-c2fad6bd8d13 · inbound
Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3960809-b0e2-4714-8998-6068c3e924f0 · inbound
Proof2Hybrid: Automatic Mathematical Benchmark Synthesis for Proof-Centric Problems FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 743df680-9c39-41e3-b4dc-e8ddf532273d · inbound
The Mathematician's Assistant: Integrating AI into Research Practice FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f8136b4-8bc9-4afc-83af-b5331c3858d3 · inbound
AI Reasoning Models for Problem Solving in Physics FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57ac86d0-1350-4f84-bccc-431fd86cc74f · inbound
A Survey of Reinforcement Learning for Large Reasoning Models FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 163
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation cff174b5-0057-4edf-ad18-0884207b7c46 · inbound
CFDLLMBench: A Benchmark Suite for Evaluating Large Language Models in Computational Fluid Dynamics FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 000180cb-00e4-4ce2-ac62-83df23cad931 · inbound
Probing the Critical Point (CritPt) of AI Reasoning: a Frontier Physics Research Benchmark FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 39cbd854-f23d-4167-9e22-43d6ae7d89bb · inbound
Co-Designing Quantum Codes with Transversal Diagonal Gates via Multi-Agent Systems FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 14fa0979-dca1-44e5-bb23-9932c7c8fdba · inbound
FEM-Bench: A Structured Scientific Reasoning Benchmark for Evaluating Code-Generating LLMs FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee92b023-b7d2-482f-bf70-ed8868aa70e3 · inbound
AI for Mathematics: Progress, Challenges, and Prospects FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation bb6e0504-f392-47a6-937c-f5cdb0f72e5b · inbound
Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e6d869f-eb4c-4f57-9385-cd585674287b · inbound
QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32b11f53-8652-454d-89ae-f51e74600799 · inbound
LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 308e4fcb-fae1-4860-b996-7d03eadb80bb · inbound
Can LLMs Prove Robotic Path Planning Optimality? A Benchmark for Research-Level Algorithm Verification FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2922276d-b6af-4b89-bb8e-79ac2994d2b0 · inbound
Can LLMs Prove Robotic Path Planning Optimality? A Benchmark for Research-Level Algorithm Verification FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9e23c67-577a-495e-a4d2-f5ea1c917233 · inbound
Xpertbench: Expert Level Tasks with Rubrics-Based Evaluation FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 5180c4eb-610f-46eb-8faa-32218492a12c · inbound
Automated Conjecture Resolution with Formal Verification FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 1dd9f079-99d9-4385-b390-bdd2948c03bb · inbound
Automated Conjecture Resolution with Formal Verification FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f9bcabb-242f-4a8a-8a44-f1e791c99e28 · inbound
Automatically Generating Hard Math Problems from Hypothesis-Driven Error Analysis FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b616c540-6143-40cb-b0b4-5de9c626106e · inbound
DeonticBench: A Benchmark for Reasoning over Rules FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 76f947a2-81ff-475f-ac0e-f6633f17f381 · inbound
Artificial Intelligence and the Structure of Mathematics FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation a37c7062-38c5-463e-bf5b-5293a35583bf · inbound
Riemann-Bench: A Benchmark for Moonshot Mathematics FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 768a1353-c835-449e-96c1-b29c33efa2ea · inbound
$k$-server-bench: Automating Potential Discovery for the $k$-Server Conjecture FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 895512bc-89d5-4ec5-8bc9-57fb48e1f44f · inbound
Problem Reductions at Scale: Agentic Integration of Computationally Hard Problems FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 363b1d25-5919-476f-994f-1a610104efaf · inbound
Agentic Frameworks for Reasoning Tasks: An Empirical Study FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ce61519e-7cd9-4cef-946f-e8e6ffa34df4 · inbound
Fine-Tuning Small Reasoning Models for Quantum Field Theory FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 7cc1818a-2bf0-4657-b3a6-8b82b28eb397 · inbound
MathDuels: A Self-Play Benchmark That Grows FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 923968c8-a34e-4d42-86a0-63c34adefdb0 · inbound
Forge: Quality-Aware Reinforcement Learning for NP-Hard Optimization in LLMs FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 5bd5622d-e6d2-4695-a04f-e7a1e94b9827 · inbound
Soohak: A Mathematician-Curated Benchmark for Evaluating Research-level Math Capabilities of LLMs FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 5d266f89-6a95-46b4-a4c1-16f6e7516e9e · inbound
Soohak: A Mathematician-Curated Benchmark for Evaluating Research-level Math Capabilities of LLMs FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 430054da-3fda-405e-bbab-74a13dba2b83 · inbound
Formal Conjectures: An Open and Evolving Benchmark for Verified Discovery in Mathematics FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation d2c6d8b1-7f5b-48b0-a474-40e12f42420a · inbound
STAR-P\'olyaMath: Multi-Agent Reasoning under Persistent Meta-Strategic Supervision FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 12f99899-ecf2-47ab-be40-905cbd9e37af · inbound
Scientific reasoning does not reliably translate into scientific forecasting in frontier AI FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 7c4509d2-bd4f-46b1-865c-5527f222b8df · inbound
Scientific reasoning does not reliably translate into scientific forecasting in frontier AI FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3da02bf1-3fb3-4eff-ae75-4c9058268dee · inbound
RMA: an Agentic System for Research-Level Mathematical Problems FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation d8931d30-f951-4dcd-8dfe-c5b8af77d71b · inbound
MadEvolve: Evolutionary Optimization of Trading Systems with Large Language Models FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 962229af-4518-4832-8218-d9276d110c0a · inbound
Self-Improving Language Models with Bidirectional Evolutionary Search FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation d2a75c58-b04b-43a3-be07-bb4cf9de881c · inbound
Why Larger Models Learn More: Effects of Capacity, Interference, and Rare-Task Retention FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 9e5c041c-014f-41ab-b5d9-1825db7f49f9 · inbound
Evaluating Interactive Reasoning in Large Language Models: A Hierarchical Benchmark with Executable Games FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation d07eae8e-f869-471d-81ad-8562e008dc17 · inbound
FVSpec: Real-World Property-Based Tests as Lean Challenges FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 731c206a-b6f2-484b-b76c-575c46a49c03 · inbound
Trust Region On-Policy Distillation FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation dd0dcb8d-b48a-49d9-b464-a09ed08527f9 · inbound
An Enigma of Artificial Reason: Investigating the Production-Evaluation Gap in Large Reasoning Models FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation abeef3e5-3717-44f4-bdf2-0905230e334a · inbound
Lean-GAP: A Dataset of Formalized Graduate Algebra Problems FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation fbfee4f5-4113-4d26-bbf3-a5866889807a · inbound
GTBench: A Curriculum-Grounded Benchmark for Evaluating LLMs as Mathematical Research Assistants in Graph Theory FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation f31297e3-a2f6-43b1-ad07-d2d819cab9ad · inbound
LeanMarathon: Toward Reliable AI Co-Mathematicians through Long-Horizon Lean Autoformalization FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation d8702a46-4bd0-4b1f-a18b-21b44c265c7b · inbound
Automated Proving of Shannon-Type Entropy Inequalities via Fine-Tuned Language Models and Guided Tree Search FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ab30e99c-aef7-4f34-ab18-747f9c148845 · inbound
CrowdMath: A Dataset of Crowdsourced Mathematical Research Discussions FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 2a7de1a8-87b1-4cf5-9585-e2d5fb46f9fa · inbound
How reliable are LLMs when it comes to playing dice? FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 318f2566-3ec4-49ab-b70d-8dc68026cf77 · inbound
Artificial Intelligence for Mathematical Reasoning: An Integrated Survey of Language Models, Neuro-symbolic Systems, and Verified Discovery FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 210
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 37473251-e072-4cb7-8b7d-60e32eb136ed · inbound
Artificial Intelligence for Mathematical Reasoning: An Integrated Survey of Language Models, Neuro-symbolic Systems, and Verified Discovery FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 212
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e91a0f67-a638-4150-85f0-45cfe8d6fe91 · inbound
Sakana Fugu Technical Report FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 5e59d95f-482f-47d8-b827-8ebecf666c0f · inbound
Learning the ARTS of Search for Automated Discovery FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation f8d10b29-0fe4-4ea3-b166-67a56fbff509 · inbound
Curriculum Reinforcement Learning Can Incentivize Reasoning Capacity in LLMs Beyond the Base Model FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 14626137-7b07-4b29-9dd6-d2defe38cf00 · inbound
Theorist Toolbox: Tools for Agent Based LLM-assisted economic theory Research FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation d03d779f-a66c-45e2-8812-73b4a58be69d · inbound
Beyond Shapley: Efficient Computation of Asymmetric Shapley Values FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 02c26e77-4dcd-44f5-80b0-124a5c2740cf · inbound
Data and Evaluation Closed-Loop for Model Capability Enhancement FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.