Pith. sign in

Paper Citation Record · LEDGER

FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

As of 21 July 2026, this Paper Citation Record lists 32 of 32 outbound references and 55 inbound Pith citation observations for arXiv:2411.04872.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.04872 v7

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-17T00:44:01.658214Z

measured 87 of 87 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-07-21T06:31:05.380196+00:00

measured 55 of 55 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-14T17:09:09.167784Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T20:47:34.579097Z

Reference resolution

32 of 32 outbound references displayed

  • verified exact1
  • verified fuzzy27
  • unresolved0
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 39460a96-c5fa-4602-b86c-f4e3b3875d4a · outbound

This paper cites an unresolved cited work.

FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Unresolved cited work

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:44:01.776584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T00:44:01.658214Z digest=sha256:23f20b486137e767f218cd879ad03cbc6585ad0f35a0172578d1e41e36059a08

Observation cf80a1b0-5933-41af-aafc-d8093549bd21 · outbound

This paper cites an unresolved cited work.

FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Unresolved cited work

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:44:01.783257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T00:44:01.658214Z digest=sha256:b357947020508f7ecfe85ce7b152b7d55104865a4afeb48b3e2f894ca0fa2cd4

Observation 5309dc1d-68ff-4f05-ade1-2a7ea1675038 · outbound

This paper cites Advances in neural information processing systems , volume=.

FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Advances in neural information processing systems , volume=

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:44:01.786765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T00:44:01.658214Z digest=sha256:63ee5eee44dd6fb2c90c6804ea3f5792f690a203fe93fe27d0596034b76cd61c

Observation be4f333b-bea8-4592-9273-32be998c95d4 · outbound

This paper cites an unresolved cited work.

FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Unresolved cited work

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:44:01.790311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T00:44:01.658214Z digest=sha256:71dd8a9a6aa27869162af0a4f7042fe49f5cbf91b9742a3d21156119ddb140f1

Observation c079aed5-3f1b-494a-bae6-bf443fe82477 · outbound

This paper cites an unresolved cited work.

FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Unresolved cited work

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:44:01.699208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T00:44:01.658214Z digest=sha256:2c75cbe91608719171fc304f576ab5c02da8ed87423604179d752415617a4534

Observation 9f7f2aeb-ed36-4447-958b-2a61098fbe23 · outbound

This paper cites an unresolved cited work.

FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Unresolved cited work

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:44:01.702938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T00:44:01.658214Z digest=sha256:0e1d34bd43cf6cc60854286fe64bb5fbc4d6727db08032d38be6eb24b5e84c1f

Observation 2c1a2fb4-fd83-4128-81a3-58261ad21fec · outbound

This paper cites an unresolved cited work.

FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Unresolved cited work

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:44:01.706206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T00:44:01.658214Z digest=sha256:031ad5e8135a3510d9430d33a15c98c5bb6629ebe76065b77c5e05edd733f48b

Observation 3c77b4b2-5b20-4d2f-b613-2ebfc97e2dc1 · outbound

This paper cites an unresolved cited work.

FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Unresolved cited work

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:44:01.709516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T00:44:01.658214Z digest=sha256:19661defd519faee24e832e274b37ab0b510d0e507fe51cad3f44fbac683f234

Observation 2da17ad2-2d52-4627-b3aa-8bea89bdab7b · outbound

This paper cites Nature , publisher =.

FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Nature , publisher =

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:44:01.712633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T00:44:01.658214Z digest=sha256:7c99c8b5effa0a0b3ea88e326fffef2e6547050912ce96e6b62f174e42b52666

Observation c868446b-73ee-4711-b2a8-fe6e8e2c5c9b · outbound

This paper cites an unresolved cited work.

FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Unresolved cited work

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:44:01.715737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T00:44:01.658214Z digest=sha256:e153038480092700be484812e6e14d94ed7b8577f243805cdf2dbd7f86769a3a

Observation ac70278b-34f8-47ab-a9a6-f505cb231d04 · outbound

This paper cites an unresolved cited work.

FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Unresolved cited work

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:44:01.718881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T00:44:01.658214Z digest=sha256:507fe3201e57b8abb9137a795a82a813883fb65922f76d25de2fefa57f24a998

Observation 183977ef-4d2a-433b-bb34-1c1af7309b15 · outbound

This paper cites an unresolved cited work.

FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Unresolved cited work

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:44:01.721927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T00:44:01.658214Z digest=sha256:94864d5c2e4511b5204d8cfdf09c33f53a7db95d26e1a61c020e8b73f68a99d3

Observation c097ccc9-30c1-4eb5-8923-fbb295536bcd · outbound

This paper cites an unresolved cited work.

FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Unresolved cited work

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:44:01.724693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T00:44:01.658214Z digest=sha256:e372be83725f34d12d7a0c1198bea775dee9dcdbca7dc0f50fb8571530de8989

Observation 5aa33541-7d41-4731-86e0-483e1659e6c2 · outbound

This paper cites Nature , publisher =.

FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Nature , publisher =

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:44:01.727711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T00:44:01.658214Z digest=sha256:03be6f6208394fec86f65059e28edb524c945e7159e797d288efb1d017a43fd0

Observation 20169156-23d9-43ac-a66f-e30a1988562f · outbound

This paper cites an unresolved cited work.

FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Unresolved cited work

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:44:01.731161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T00:44:01.658214Z digest=sha256:50c7b665c6b35ab5a3726bfd1959659bffe2d2c6bfca7818849c53f41ff065f9

Observation 49f4966f-5311-4f36-906f-a78d8ef8a620 · outbound

This paper cites an unresolved cited work.

FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Unresolved cited work

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:44:01.734398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T00:44:01.658214Z digest=sha256:571a8c319be724380c441a2835d46c39d13c49fd7f0e4fd297c551c8235a9784

Observation cb113683-04c4-4100-9fa5-7dcd2199b93e · outbound

This paper cites Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks.

FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T00:44:01.690041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T00:44:01.658214Z digest=sha256:aa25d4630fe5ebe4a499bc22ca93849a49121ebcd88ab52b04e2335826c6f4aa

Observation f5efe89a-38c1-4cda-b0d3-d5e426d94940 · outbound

This paper cites an unresolved cited work.

FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Unresolved cited work

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:44:01.738569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T00:44:01.658214Z digest=sha256:cb1b8149f9d7a536569c9d16684f695891617a5d0f36b5e371f628e9676b6b5c

Observation ecb2761b-ec7c-4459-9f3f-362fbc8cf68b · outbound

This paper cites an unresolved cited work.

FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Unresolved cited work

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:44:01.743211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T00:44:01.658214Z digest=sha256:b62bebe9e2ad7175b0209808117bfd2ae272489e8d11377ac252c2ca7f026f75

Observation 2ccac011-6135-4a9f-bc18-d28ad5d74592 · outbound

This paper cites an unresolved cited work.

FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Unresolved cited work

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:44:01.746652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T00:44:01.658214Z digest=sha256:c7c649738cf9d79262b5382c97a2cfecca24220357942091195dc20b0053af70

Observation 41001959-cce5-4d10-bb2c-844d83df24ba · outbound

This paper cites an unresolved cited work.

FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Unresolved cited work

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:44:01.749651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T00:44:01.658214Z digest=sha256:65a83f23462aca8ba22c206c63b44a72afa3aeae73a6ec22c541269ef0681562

Observation a5280095-894f-4cae-bfc9-9b3fc95a88b1 · outbound

This paper cites an unresolved cited work.

FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Unresolved cited work

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:44:01.752545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T00:44:01.658214Z digest=sha256:6dc094d9a7a7f86c140e75ee6659a94464e2fd5ae04fab0a37aa7105fa645d22

Observation 776e030a-3b12-4862-92ad-422680b006b0 · outbound

This paper cites an unresolved cited work.

FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Unresolved cited work

Reference 23

Resolution
parse uncertain
raw_fallback, observed 2026-05-17T00:44:01.755642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T00:44:01.658214Z digest=sha256:c16878bbd9f31be32c4c871b452cacd9a51563b662e14828918d0eb644acea0b

Observation 99a5fd29-758b-437d-bd9a-248f998a0b59 · outbound

This paper cites an unresolved cited work.

FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Unresolved cited work

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:44:01.758675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T00:44:01.658214Z digest=sha256:748f11ffb23b254945ce47b280b024a3f9bf1179d71dec36f584588d42983553

Observation f22bb72c-833d-4472-b5fa-c701bd9a23e9 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 25

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T00:44:01.695941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T00:44:01.658214Z digest=sha256:da11298b260c97c8b51fc4877c735c1c8614538120c21c701b98460f3c609b79

Observation 07450d21-c93d-4329-b626-a3142cd1aa79 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Measuring Massive Multitask Language Understanding

Reference 26

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T00:44:01.684458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T00:44:01.658214Z digest=sha256:671e102ce95a0e00785c55ad0cd104665d50ee063dc7c3c9dd70887adfb9976c

Observation 85c98fc7-0d24-46bb-98ab-dbb71bc6ed53 · outbound

This paper cites an unresolved cited work.

FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Unresolved cited work

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:44:01.761645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T00:44:01.658214Z digest=sha256:6417a209cb21f8ffccd182881011be25c51cea5d20c04101f25743cc5ec340c6

Observation 18eade63-0879-4ee7-99d4-1a3908f22e33 · outbound

This paper cites an unresolved cited work.

FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Unresolved cited work

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:44:01.765058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T00:44:01.658214Z digest=sha256:51797f53b2c3dd2cca2b1789a41fc3076a214650874aff9db65467a081325ae6

Observation 57c6425c-6289-4531-b774-b58208d994f1 · outbound

This paper cites Equality of orders of a set of integers modulo a prime.

FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Equality of orders of a set of integers modulo a prime

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:44:01.679683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T00:44:01.658214Z digest=sha256:50004003393aed8e2f7eacee9d0b4feb30b69fe13a88480ac69d577112de4928

Observation 473cd1cc-18a3-4f7b-9b8a-2b3dfddfdead · outbound

This paper cites Curves over Finite Fields Attaining the Hasse-Weil Upper Bound.

FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Curves over Finite Fields Attaining the Hasse-Weil Upper Bound

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:44:01.769426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T00:44:01.658214Z digest=sha256:95f232cc72d86d48ebfb0769eb91263f3e898492e819eca3a68dfef7dff7c61f

Observation 3a52ed73-4240-4524-95bb-a66541497403 · outbound

This paper cites an unresolved cited work.

FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Unresolved cited work

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:44:01.773294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T00:44:01.658214Z digest=sha256:2ac3ea9241f2f415208bc0991ab0e6c560edf8ba56cfb91864a9916548aaeb41

Observation becd89a7-e3cb-4383-a917-ace9399e0154 · outbound

This paper cites an unresolved cited work.

FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Unresolved cited work

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:44:01.779996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T00:44:01.658214Z digest=sha256:c1b1a1e3c35f8ed8c71a07d76c1768e7bb08f8cf529509d14d658d132aa68a77

Pith citing papers

Observation 72a19ef2-e197-4494-ac82-e0e76871d46d · inbound

Humanity's Last Exam cites this paper.

Humanity's Last Exam FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:44:01.791449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:7022844503b8fa078d53f221cf3a4cd8bff34cddf5c7ad087e2882f578deaede

Observation c25b27db-ccc0-4271-b56c-714685194501 · inbound

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models cites this paper.

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 213

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:44:01.791449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-12T08:40:40.910461Z digest=sha256:54d93e995d73d2bcd62dfd78a9786537ab511c64314bebaac2dc9c30c029546f

Observation 57ac86d0-1350-4f84-bccc-431fd86cc74f · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 163

Resolution
verified exact
local_arxiv, observed 2026-05-18T00:02:25.462023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:da9469126d10f133ca0ecb443c91f30b07a7db4946039d58925d7461998ae7fc

Observation cff174b5-0057-4edf-ad18-0884207b7c46 · inbound

CFDLLMBench: A Benchmark Suite for Evaluating Large Language Models in Computational Fluid Dynamics cites this paper.

CFDLLMBench: A Benchmark Suite for Evaluating Large Language Models in Computational Fluid Dynamics FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 19

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T15:06:32.104836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-18T15:05:37.519850Z digest=sha256:8564cd69a2dda5ac96f4d074f26f8c8f8fd359d1871fd102c52ed94cba2b5938

Observation 000180cb-00e4-4ce2-ac62-83df23cad931 · inbound

Probing the Critical Point (CritPt) of AI Reasoning: a Frontier Physics Research Benchmark cites this paper.

Probing the Critical Point (CritPt) of AI Reasoning: a Frontier Physics Research Benchmark FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-05-18T11:52:35.313977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-18T11:52:10.205796Z digest=sha256:5dbbc779061c62fcd6be7974e4e0c09f28211dd42327abbc4a8ab6a05d484b4e

Observation 39cbd854-f23d-4167-9e22-43d6ae7d89bb · inbound

Co-Designing Quantum Codes with Transversal Diagonal Gates via Multi-Agent Systems cites this paper.

Co-Designing Quantum Codes with Transversal Diagonal Gates via Multi-Agent Systems FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-18T04:30:52.976368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-18T04:28:48.400311Z digest=sha256:9a75d5e38c4d8c58739aae49cf0b794341c7d7b4159f427806e888a0c3a5cce1

Observation ee92b023-b7d2-482f-bf70-ed8868aa70e3 · inbound

AI for Mathematics: Progress, Challenges, and Prospects cites this paper.

AI for Mathematics: Progress, Challenges, and Prospects FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:44:01.791449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-16T13:24:57.923863Z digest=sha256:ec5684251b6c1481f5594890a18705a9404f2e41e674947c17486296cd54d347

Observation 308e4fcb-fae1-4860-b996-7d03eadb80bb · inbound

Can LLMs Prove Robotic Path Planning Optimality? A Benchmark for Research-Level Algorithm Verification cites this paper.

Can LLMs Prove Robotic Path Planning Optimality? A Benchmark for Research-Level Algorithm Verification FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-13T22:02:43.946867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:02:43.946867Z digest=sha256:948e7bbbe8da6d35668a0b45ea202375c93284da16e3dd709b6b7b7e155fa62a

Observation c9e23c67-577a-495e-a4d2-f5ea1c917233 · inbound

Xpertbench: Expert Level Tasks with Rubrics-Based Evaluation cites this paper.

Xpertbench: Expert Level Tasks with Rubrics-Based Evaluation FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:44:01.791449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-14T23:22:53.368231Z digest=sha256:54d2369ea0c41bca6d9227220b5f7a290ddfd8c30762e2c4af745781279ecc8e

Observation 5180c4eb-610f-46eb-8faa-32218492a12c · inbound

Automated Conjecture Resolution with Formal Verification cites this paper.

Automated Conjecture Resolution with Formal Verification FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:44:01.791449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-13T18:05:38.340467Z digest=sha256:ef6cfb44504159f616a66fb5bcf6831f469474e540126274c54592f66e9833fa

Observation 1dd9f079-99d9-4385-b390-bdd2948c03bb · inbound

Automated Conjecture Resolution with Formal Verification cites this paper.

Automated Conjecture Resolution with Formal Verification FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-13T12:18:10.165894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T12:18:10.165894Z digest=sha256:d88968cdc1524c932932c7b4c434ffc979d2a78244db7d2b66619630441d6ef1

Observation 0f9bcabb-242f-4a8a-8a44-f1e791c99e28 · inbound

Automatically Generating Hard Math Problems from Hypothesis-Driven Error Analysis cites this paper.

Automatically Generating Hard Math Problems from Hypothesis-Driven Error Analysis FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:44:01.791449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-10T19:39:36.819142Z digest=sha256:2afd62d92b7131da6296da3ae58e3af817e7f488b23ddcca42cb0eb98516cd28

Observation b616c540-6143-40cb-b0b4-5de9c626106e · inbound

DeonticBench: A Benchmark for Reasoning over Rules cites this paper.

DeonticBench: A Benchmark for Reasoning over Rules FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:44:01.791449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-10T20:22:03.956227Z digest=sha256:7d0f8d47037787395c763b147537291eababbe9c9470ab4daa1fb29322b9c5b5

Observation 76f947a2-81ff-475f-ac0e-f6633f17f381 · inbound

Artificial Intelligence and the Structure of Mathematics cites this paper.

Artificial Intelligence and the Structure of Mathematics FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 45

Resolution
malformed identifier
arxiv_id, observed 2026-05-17T00:44:01.791449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-10T19:12:23.402808Z digest=sha256:739658e1f676da9f6a1e7381c792f7387fd1c1112ea9a254c2cabf72bdc5d1cf

Observation a37c7062-38c5-463e-bf5b-5293a35583bf · inbound

Riemann-Bench: A Benchmark for Moonshot Mathematics cites this paper.

Riemann-Bench: A Benchmark for Moonshot Mathematics FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:44:01.791449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-10T18:02:42.607682Z digest=sha256:38e25451f275c447002fd0957eec06b427251c04332ae8ef407b81e3d45f90a3

Observation 768a1353-c835-449e-96c1-b29c33efa2ea · inbound

$k$-server-bench: Automating Potential Discovery for the $k$-Server Conjecture cites this paper.

$k$-server-bench: Automating Potential Discovery for the $k$-Server Conjecture FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:44:01.791449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-10T17:48:17.363568Z digest=sha256:81e33e5d632e24d3cba4aa368a22137ec229908ea1956a2ac530bf543c84182c

Observation 895512bc-89d5-4ec5-8bc9-57fb48e1f44f · inbound

Problem Reductions at Scale: Agentic Integration of Computationally Hard Problems cites this paper.

Problem Reductions at Scale: Agentic Integration of Computationally Hard Problems FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:44:01.791449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-10T16:07:35.395872Z digest=sha256:64d54a91937756b82a7993188aac0f1e2622f89eb4f5d97697861e54ca3bc52f

Observation 363b1d25-5919-476f-994f-1a610104efaf · inbound

Agentic Frameworks for Reasoning Tasks: An Empirical Study cites this paper.

Agentic Frameworks for Reasoning Tasks: An Empirical Study FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 55

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T00:44:01.791449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-10T08:24:51.573913Z digest=sha256:7bf38b546b9ce8872ccff76be0e4751015e543fe647717ac069cfcd922e633c1

Observation ce61519e-7cd9-4cef-946f-e8e6ffa34df4 · inbound

Fine-Tuning Small Reasoning Models for Quantum Field Theory cites this paper.

Fine-Tuning Small Reasoning Models for Quantum Field Theory FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:44:01.791449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-10T03:23:18.770963Z digest=sha256:4557a0806aa05f5c856df6b94dafe46cfdecce545ddee7d49424bc8f85b4101f

Observation 7cc1818a-2bf0-4657-b3a6-8b82b28eb397 · inbound

MathDuels: Evaluating LLMs as Problem Posers and Solvers cites this paper.

MathDuels: Evaluating LLMs as Problem Posers and Solvers FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T00:44:01.791449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-09T21:26:52.011933Z digest=sha256:0b2c7e31410f9c39ee0a7bec351c468fedc65b8898cd36e69a4ebba3e7d8b1a4

Observation 923968c8-a34e-4d42-86a0-63c34adefdb0 · inbound

Forge: Quality-Aware Reinforcement Learning for NP-Hard Optimization in LLMs cites this paper.

Forge: Quality-Aware Reinforcement Learning for NP-Hard Optimization in LLMs FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T00:44:01.791449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-12T02:44:33.143247Z digest=sha256:852a292603e013b07c0d6f67c4554ec63990fbcf86d0c1f3f66febe00c5bf7a6

Observation 5bd5622d-e6d2-4695-a04f-e7a1e94b9827 · inbound

Soohak: A Mathematician-Curated Benchmark for Evaluating Research-level Math Capabilities of LLMs cites this paper.

Soohak: A Mathematician-Curated Benchmark for Evaluating Research-level Math Capabilities of LLMs FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T00:44:01.791449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-12T01:53:12.708706Z digest=sha256:cc953bdeef03dbe19e46355bb1c2cb21f0b5df77e49dd9cfe593d95840653f10

Observation 5d266f89-6a95-46b4-a4c1-16f6e7516e9e · inbound

Soohak: A Mathematician-Curated Benchmark for Evaluating Research-level Math Capabilities of LLMs cites this paper.

Soohak: A Mathematician-Curated Benchmark for Evaluating Research-level Math Capabilities of LLMs FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T22:24:07.822738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-20T22:24:04.625823Z digest=sha256:7fab494b146fc70117c53ee0c49d3d9383dd1066e94a013a0cc7becff5aa5ad6

Observation 430054da-3fda-405e-bbab-74a13dba2b83 · inbound

Formal Conjectures: An Open and Evolving Benchmark for Verified Discovery in Mathematics cites this paper.

Formal Conjectures: An Open and Evolving Benchmark for Verified Discovery in Mathematics FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:44:01.791449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-14T20:20:54.566662Z digest=sha256:3e898876758ffb42981d801b6b72e6bb7fe885a43ea47226f1412dccfd51448d

Observation d2c6d8b1-7f5b-48b0-a474-40e12f42420a · inbound

STAR-P\'olyaMath: Multi-Agent Reasoning under Persistent Meta-Strategic Supervision cites this paper.

STAR-P\'olyaMath: Multi-Agent Reasoning under Persistent Meta-Strategic Supervision FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-20T02:58:00.175045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-20T02:56:19.104905Z digest=sha256:8fa4b1d205786d02bd80df5fc2a10ffb0ac3caa7c4a9723b4d119c83402dd404

Observation 12f99899-ecf2-47ab-be40-905cbd9e37af · inbound

Scientific reasoning does not reliably translate into scientific forecasting in frontier AI cites this paper.

Scientific reasoning does not reliably translate into scientific forecasting in frontier AI FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-05-22T05:21:07.455360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-22T05:18:15.358571Z digest=sha256:4a12672844c91256623c9e8f7b02c30c12695269f5e620a956758211d55fb89d

Observation 3da02bf1-3fb3-4eff-ae75-4c9058268dee · inbound

RMA: an Agentic System for Research-Level Mathematical Problems cites this paper.

RMA: an Agentic System for Research-Level Mathematical Problems FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-25T06:10:24.159655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-25T06:06:47.226726Z digest=sha256:c104256877222d3dfb889cd4bc67e93f94282a490ad4a4388893870f409b1b1e

Observation d8931d30-f951-4dcd-8dfe-c5b8af77d71b · inbound

MadEvolve: Evolutionary Optimization of Trading Systems with Large Language Models cites this paper.

MadEvolve: Evolutionary Optimization of Trading Systems with Large Language Models FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-25T05:25:23.375844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-25T05:24:05.098181Z digest=sha256:74d15ffc1977aa86f122a49e65a6613c228290f8e71832b3e2c0f2aa3d8c447b

Observation 962229af-4518-4832-8218-d9276d110c0a · inbound

Self-Improving Language Models with Bidirectional Evolutionary Search cites this paper.

Self-Improving Language Models with Bidirectional Evolutionary Search FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:03:26.436991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-06-29T12:57:28.931392Z digest=sha256:9b2b0029d095a55adf4e87293ee2e3b7e3e4fe87a68ed4c2553ca8cdc6852f68

Observation d2a75c58-b04b-43a3-be07-bb4cf9de881c · inbound

Why Larger Models Learn More: Effects of Capacity, Interference, and Rare-Task Retention cites this paper.

Why Larger Models Learn More: Effects of Capacity, Interference, and Rare-Task Retention FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-06-29T08:43:15.141444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-06-29T08:39:14.327838Z digest=sha256:268e4375f66ed1345bc42058be6e665cd6914ae52126ba1591ae38a22f2b3e7d

Observation 9e5c041c-014f-41ab-b5d9-1825db7f49f9 · inbound

Evaluating Interactive Reasoning in Large Language Models: A Hierarchical Benchmark with Executable Games cites this paper.

Evaluating Interactive Reasoning in Large Language Models: A Hierarchical Benchmark with Executable Games FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T17:53:47.431723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-06-29T17:46:39.281623Z digest=sha256:6d334eea9471685b24cf868fa39f3bcce5e5de91fcea75c96a50365979efb294

Observation d07eae8e-f869-471d-81ad-8562e008dc17 · inbound

FVSpec: Real-World Property-Based Tests as Lean Challenges cites this paper.

FVSpec: Real-World Property-Based Tests as Lean Challenges FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-06-28T17:12:25.298307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-06-28T17:05:13.012431Z digest=sha256:c647a75e16f2d0c492f9e80072a4d2241ce83406b62b8d3e4e2ed9e010af2938

Observation 731c206a-b6f2-484b-b76c-575c46a49c03 · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 45

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T20:56:13.250290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:04ae90d56a9f2a39c60429f96a4bb755986926d2712499cc96bf0dcc6d615697

Observation dd0dcb8d-b48a-49d9-b464-a09ed08527f9 · inbound

An Enigma of Artificial Reason: Investigating the Production-Evaluation Gap in Large Reasoning Models cites this paper.

An Enigma of Artificial Reason: Investigating the Production-Evaluation Gap in Large Reasoning Models FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-01T21:36:14.993846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-06-28T16:45:33.046568Z digest=sha256:13190f36b48ccdd45c07ba8662defb1da68b2aedb82e117febb4127346477467

Observation abeef3e5-3717-44f4-bdf2-0905230e334a · inbound

Lean-GAP: A Dataset of Formalized Graduate Algebra Problems cites this paper.

Lean-GAP: A Dataset of Formalized Graduate Algebra Problems FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-06-30T17:34:57.542083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-06-30T17:32:00.411535Z digest=sha256:32003021a12cd406660d6e761f3c9ea1daca932d229e2ca181fff25ed3de03cb

Observation fbfee4f5-4113-4d26-bbf3-a5866889807a · inbound

GTBench: A Curriculum-Grounded Benchmark for Evaluating LLMs as Mathematical Research Assistants in Graph Theory cites this paper.

GTBench: A Curriculum-Grounded Benchmark for Evaluating LLMs as Mathematical Research Assistants in Graph Theory FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-02T03:06:30.296889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-06-28T10:19:03.670540Z digest=sha256:ce1828d280516566c5166bb7f990455d6d4da7ef2809d4ccf2c0a321a49af17d

Observation f31297e3-a2f6-43b1-ad07-d2d819cab9ad · inbound

LeanMarathon: Toward Reliable AI Co-Mathematicians through Long-Horizon Lean Autoformalization cites this paper.

LeanMarathon: Toward Reliable AI Co-Mathematicians through Long-Horizon Lean Autoformalization FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T08:36:48.761264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-06-28T05:48:56.691155Z digest=sha256:977d9698a719871756119216e03cae9163391554e8c71a361bba37d7fb360d3c

Observation d8702a46-4bd0-4b1f-a18b-21b44c265c7b · inbound

Automated Proving of Shannon-Type Entropy Inequalities via Fine-Tuned Language Models and Guided Tree Search cites this paper.

Automated Proving of Shannon-Type Entropy Inequalities via Fine-Tuned Language Models and Guided Tree Search FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-07-02T15:27:05.514373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-06-27T23:52:11.730333Z digest=sha256:a25a28532f1b466190aaef211686ce6a373a54749b68ee7d23d4abf6b280f7e1

Observation ab30e99c-aef7-4f34-ab18-747f9c148845 · inbound

CrowdMath: A Dataset of Crowdsourced Mathematical Research Discussions cites this paper.

CrowdMath: A Dataset of Crowdsourced Mathematical Research Discussions FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 20

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T03:46:32.299057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-06-28T09:45:18.385925Z digest=sha256:32d914881b0986617895c04b4e9302dd320f729d748480703a480e0c09d52c87

Observation 2a7de1a8-87b1-4cf5-9585-e2d5fb46f9fa · inbound

How reliable are LLMs when it comes to playing dice? cites this paper.

How reliable are LLMs when it comes to playing dice? FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-02T17:37:14.662833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-06-27T21:56:42.766016Z digest=sha256:8438a1ff2432bc0a3a53a69e83ab0e79acd31ab9d5d6595f2c6b78cc1540f1a6

Observation 318f2566-3ec4-49ab-b70d-8dc68026cf77 · inbound

Artificial Intelligence for Mathematical Reasoning: An Integrated Survey of Language Models, Neuro-symbolic Systems, and Verified Discovery cites this paper.

Artificial Intelligence for Mathematical Reasoning: An Integrated Survey of Language Models, Neuro-symbolic Systems, and Verified Discovery FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 210

Resolution
verified exact
local_arxiv, observed 2026-07-02T22:47:25.916162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-06-27T18:39:44.696961Z digest=sha256:b9b76a2dcbeda67cc65b38e15f01f5d83ee387da3e49cb63cb091f9214164224

Observation e91a0f67-a638-4150-85f0-45cfe8d6fe91 · inbound

Sakana Fugu Technical Report cites this paper.

Sakana Fugu Technical Report FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T06:39:36.939254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-06-26T14:22:37.596720Z digest=sha256:137f6584d26ff3d2093bca6d9506824d9352808112ed214d189654cd463be3f3

Observation 5e59d95f-482f-47d8-b827-8ebecf666c0f · inbound

Learning the ARTS of Search for Automated Discovery cites this paper.

Learning the ARTS of Search for Automated Discovery FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-04T08:09:41.810861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-06-26T12:04:13.117307Z digest=sha256:1910873fad839fac7a1ed1c6b52554a80abe65c5fe6b68ccbf20adb1dfda285e

Observation f8d10b29-0fe4-4ea3-b166-67a56fbff509 · inbound

Curriculum Reinforcement Learning Can Incentivize Reasoning Capacity in LLMs Beyond the Base Model cites this paper.

Curriculum Reinforcement Learning Can Incentivize Reasoning Capacity in LLMs Beyond the Base Model FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-04T08:39:42.300162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-06-26T11:10:56.755558Z digest=sha256:4648bc4a37daf566117d976b48e533aaa25f1664b5ab80ac5dc2367a10e5e015

Observation 14626137-7b07-4b29-9dd6-d2defe38cf00 · inbound

Theorist Toolbox: Tools for Agent Based LLM-assisted economic theory Research cites this paper.

Theorist Toolbox: Tools for Agent Based LLM-assisted economic theory Research FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-04T09:29:43.421885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-06-26T10:01:18.103014Z digest=sha256:504db951c5aa7ac00b1d40c995a10cbab0da14c2fb98d8eed627f866137416a1

Observation d03d779f-a66c-45e2-8812-73b4a58be69d · inbound

Beyond Shapley: Efficient Computation of Asymmetric Shapley Values cites this paper.

Beyond Shapley: Efficient Computation of Asymmetric Shapley Values FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T18:20:04.548897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-06-25T22:56:30.582241Z digest=sha256:770297890f47e304d1194e7475731b8c6f48faa55e2abbda9bf921be666fc02c

Observation 02c26e77-4dcd-44f5-80b0-124a5c2740cf · inbound

Data and Evaluation Closed-Loop for Model Capability Enhancement cites this paper.

Data and Evaluation Closed-Loop for Model Capability Enhancement FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T15:25:48.060133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-06-30T01:33:08.134048Z digest=sha256:96f7f40c0ef4d93bd0cb0eaddfa092cd1663ea0e7a29315cdd5f1128367270f7

Observation bb81e483-d698-401f-8a86-d526400fdf39 · inbound

IsoSci: A Benchmark of Isomorphic Cross-Domain Science Problems for Evaluating Reasoning versus Knowledge Retrieval in LLMs cites this paper.

IsoSci: A Benchmark of Isomorphic Cross-Domain Science Problems for Evaluating Reasoning versus Knowledge Retrieval in LLMs FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-03T21:18:58.261153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-07-03T21:12:57.721334Z digest=sha256:6ba8179810a741cff7610ef4c851b0ac1076cdab634981403622f5d29632c3c3

Observation e26a7cc6-4c92-4e53-9097-065d3beef83d · inbound

MatPhaseBench: A Semantics-Guided Benchmark for Materials Phase Diagrams Understanding cites this paper.

MatPhaseBench: A Semantics-Guided Benchmark for Materials Phase Diagrams Understanding FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-12T06:00:12.267694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:00:12.267694Z digest=sha256:594a69597eed7a199be064a387335de0a011ff127b22ea5c643ee7c10eb2e3f3

Observation bf1d33b7-b1f6-4caa-a905-0a9e94defc9e · inbound

Evaluating SageMath-Augmented LLM Agents for Computational and Experimental Mathematics cites this paper.

Evaluating SageMath-Augmented LLM Agents for Computational and Experimental Mathematics FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 21

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T20:47:34.580650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-07-10T20:39:42.009312Z digest=sha256:19762b0f6bc57505e6fe58fe0cce7d663f2fd8289f017d8cdb530466080c725d

Observation fad7a543-2869-4193-9615-94f21cfc5335 · inbound

Measuring Intelligence Beyond Human Scale cites this paper.

Measuring Intelligence Beyond Human Scale FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-07-09T21:26:34.386221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-07-09T21:21:39.305904Z digest=sha256:1ffb15d0b024e2cc9e6fbf2f461a0eaf6d117a1273d1072000399d31e496a627

Observation a47f6bd2-58ca-4091-8030-fafbf5ae28aa · inbound

From Solvers to Research: Large Language Model-Driven Formal Mathematics at the Research Frontier cites this paper.

From Solvers to Research: Large Language Model-Driven Formal Mathematics at the Research Frontier FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 80

Resolution
verified exact
local_arxiv, observed 2026-07-10T18:17:33.708098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-07-10T18:16:31.176239Z digest=sha256:f0b43234c34e7d5a4788b47c63d637efe296d253d38e788e10ff4de9078792db

Observation f2ed8671-3228-4d34-8058-2387f304e2a2 · inbound

When Does Continual Learning Require Learning cites this paper.

When Does Continual Learning Require Learning FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-10T16:57:24.591984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-07-10T16:49:00.923598Z digest=sha256:9e8892bebc62bb6f736cee0636f01416baacf613a7c3030ad83c5b6560cea708

Observation 1ae42162-0207-430e-9cfb-6f122a4d000a · inbound

The Ramanujan Challenge For AI cites this paper.

The Ramanujan Challenge For AI FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 230

Resolution
unresolved
no resolver link, observed 2026-07-14T17:09:09.167784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T17:09:09.167784Z digest=sha256:9dbf016555c826ece79d7dcd48a4a575ceb00b51698426af40827bd2b53f4f53

Observation 23bbf43f-c0d0-47d9-a1f5-e31a1a0f39ef · inbound

AdvancedMathBench: A Benchmark Suite for Advanced Mathematical Proof Generation and Verification cites this paper.

AdvancedMathBench: A Benchmark Suite for Advanced Mathematical Proof Generation and Verification FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-14T02:43:21.225324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T02:43:21.225324Z digest=sha256:80ab89174cd7eca18f0403caaee9bbb02e5fbdedf7c000db180a2d0364470ba3