Pith. sign in

Paper Citation Record · LEDGER

FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

As of 23 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 100 inbound Pith citation observations for arXiv:2411.04872.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.04872 v7

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-17T00:44:01.658214Z

measured 132 of 132 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 100 of 120 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:39:57.153685Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T20:47:34.579097Z

Reference resolution

32 of 32 outbound references displayed

  • verified exact1
  • verified fuzzy4
  • unresolved23
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 39460a96-c5fa-4602-b86c-f4e3b3875d4a · outbound

This paper cites an unresolved cited work.

FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-05-17T00:44:01.776584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-17T00:44:01.658214Z digest=sha256:872d7b58b564dd017603ce59099abfaab24f41a8b0097c4449b64086e0f0b501

Observation cf80a1b0-5933-41af-aafc-d8093549bd21 · outbound

This paper cites an unresolved cited work.

FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-05-17T00:44:01.783257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-17T00:44:01.658214Z digest=sha256:5f7d0df4a388556d4b70a37403ffde5c4e979e945747c1a50d900d38425281a8

Observation 5309dc1d-68ff-4f05-ade1-2a7ea1675038 · outbound

This paper cites Advances in neural information processing systems , volume=.

FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Advances in neural information processing systems , volume=

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:44:01.786765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-17T00:44:01.658214Z digest=sha256:a108c56f40c177d9db5dd33ff0d3cebfda7d40d0a6418b24236e00aee2150275

Observation be4f333b-bea8-4592-9273-32be998c95d4 · outbound

This paper cites an unresolved cited work.

FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-05-17T00:44:01.790311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-17T00:44:01.658214Z digest=sha256:e4293ce823da93856c15082136bac49ffcafcfbb0309198db0821bfddd3e678f

Observation c079aed5-3f1b-494a-bae6-bf443fe82477 · outbound

This paper cites an unresolved cited work.

FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-05-17T00:44:01.699208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-17T00:44:01.658214Z digest=sha256:001c8a48fe6483c70315ac577d53867c8a4fb6cbd99849b64b2e7cef659dd506

Observation 9f7f2aeb-ed36-4447-958b-2a61098fbe23 · outbound

This paper cites an unresolved cited work.

FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-05-17T00:44:01.702938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-17T00:44:01.658214Z digest=sha256:97e204475249cdcb808f719ae278d4fa61f655d2c88838d121679674246d6905

Observation 2c1a2fb4-fd83-4128-81a3-58261ad21fec · outbound

This paper cites an unresolved cited work.

FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-05-17T00:44:01.706206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-17T00:44:01.658214Z digest=sha256:322a80e5f09443b6601fcd895a259daec49849d5c9ad9889783c8c7cfa726ed9

Observation 3c77b4b2-5b20-4d2f-b613-2ebfc97e2dc1 · outbound

This paper cites an unresolved cited work.

FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-05-17T00:44:01.709516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-17T00:44:01.658214Z digest=sha256:a3dc86164c252c868a592313f9f3616e65e20f22a0f4fc2cadda6fdb6a596d2f

Observation 2da17ad2-2d52-4627-b3aa-8bea89bdab7b · outbound

This paper cites Nature , publisher =.

FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Nature , publisher =

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:44:01.712633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-17T00:44:01.658214Z digest=sha256:d5de13ae6e617876298a573767e5fa926d294e3c13bf8abed858ba01f90884b0

Observation c868446b-73ee-4711-b2a8-fe6e8e2c5c9b · outbound

This paper cites an unresolved cited work.

FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-05-17T00:44:01.715737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-17T00:44:01.658214Z digest=sha256:bae6eb359aa6f1125eef76676831c18cf4e6abf52dbf6ed7b70c6843594f2f6e

Observation ac70278b-34f8-47ab-a9a6-f505cb231d04 · outbound

This paper cites an unresolved cited work.

FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-05-17T00:44:01.718881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-17T00:44:01.658214Z digest=sha256:4c5ce3b02a1e88610d15e32d2fa60b5f1898b720384bee97864e94c86674dab6

Observation 183977ef-4d2a-433b-bb34-1c1af7309b15 · outbound

This paper cites an unresolved cited work.

FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-05-17T00:44:01.721927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-17T00:44:01.658214Z digest=sha256:9a2f793f130c8f08999c905c3d962cd5905885233b21091657547e9bcb11f389

Observation c097ccc9-30c1-4eb5-8923-fbb295536bcd · outbound

This paper cites an unresolved cited work.

FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-05-17T00:44:01.724693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-17T00:44:01.658214Z digest=sha256:876bfcf453115d2b375a85d0ae0c20c97975fe09a91e8f9685196125030b7346

Observation 5aa33541-7d41-4731-86e0-483e1659e6c2 · outbound

This paper cites Nature , publisher =.

FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Nature , publisher =

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:44:01.727711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-17T00:44:01.658214Z digest=sha256:84294b3da51356431affe7bb2788b6c6f76a31249a8edee8c5103e25eeb7dc37

Observation 20169156-23d9-43ac-a66f-e30a1988562f · outbound

This paper cites an unresolved cited work.

FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-05-17T00:44:01.731161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-17T00:44:01.658214Z digest=sha256:f401a21219e56093b038fcce3dca6f218d46f7542c33ec36cbf37c30998d8496

Observation 49f4966f-5311-4f36-906f-a78d8ef8a620 · outbound

This paper cites an unresolved cited work.

FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-05-17T00:44:01.734398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-17T00:44:01.658214Z digest=sha256:f0a3101d82606d0d47229a0438674044cfc0f6f28c036a69fa3dc1789a2515fc

Observation cb113683-04c4-4100-9fa5-7dcd2199b93e · outbound

This paper cites Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks.

FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T00:44:01.690041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-17T00:44:01.658214Z digest=sha256:6e71ae925dcdd4299ae8bab430aa6a28e3cfaa05b74f6e3db16bab2a1acfab77

Observation f5efe89a-38c1-4cda-b0d3-d5e426d94940 · outbound

This paper cites an unresolved cited work.

FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-05-17T00:44:01.738569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-17T00:44:01.658214Z digest=sha256:5d74bb23be03746bfc9fc9190aa21e298de0fb9635782ff0258ec155b9342052

Observation ecb2761b-ec7c-4459-9f3f-362fbc8cf68b · outbound

This paper cites an unresolved cited work.

FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-05-17T00:44:01.743211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-17T00:44:01.658214Z digest=sha256:d94e3dd8940235dcd49e98f9fa5181dc57f00c9bc09a343589dd40e7a5f697b4

Observation 2ccac011-6135-4a9f-bc18-d28ad5d74592 · outbound

This paper cites an unresolved cited work.

FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-05-17T00:44:01.746652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-17T00:44:01.658214Z digest=sha256:1f998b3b48219d235ccc18f28b03b3bb06305e1fa9ece2c739c2edb9ea86d992

Observation 41001959-cce5-4d10-bb2c-844d83df24ba · outbound

This paper cites an unresolved cited work.

FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-05-17T00:44:01.749651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-17T00:44:01.658214Z digest=sha256:ff24e6e768ba1ba342fcd21d340f30546568df5d737dfbad2bbe38f330442135

Observation a5280095-894f-4cae-bfc9-9b3fc95a88b1 · outbound

This paper cites an unresolved cited work.

FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-05-17T00:44:01.752545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-17T00:44:01.658214Z digest=sha256:912f772b1351c94aa0e2bc68b2103c75b2e39e936833ffee1f58e7a0de0292bb

Observation 776e030a-3b12-4862-92ad-422680b006b0 · outbound

This paper cites an unresolved cited work.

FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Unresolved cited work

Reference 23

Resolution
parse uncertain
raw_fallback, observed 2026-05-17T00:44:01.755642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-17T00:44:01.658214Z digest=sha256:7997e609161cab2126ca8bf4206be712c02110ceeef4b88133d75679c751fed0

Observation 99a5fd29-758b-437d-bd9a-248f998a0b59 · outbound

This paper cites an unresolved cited work.

FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-05-17T00:44:01.758675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-17T00:44:01.658214Z digest=sha256:64af7f9b654c3b0cd7b607262c6f001fc056937f50f1f03a02db1db5e4201f98

Observation f22bb72c-833d-4472-b5fa-c701bd9a23e9 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 25

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T00:44:01.695941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-17T00:44:01.658214Z digest=sha256:447f3d99328d5ab933442916b2ec9f3dcf45d317a57cfc82ba86487024409d58

Observation 07450d21-c93d-4329-b626-a3142cd1aa79 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Measuring Massive Multitask Language Understanding

Reference 26

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T00:44:01.684458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-17T00:44:01.658214Z digest=sha256:ba0596f752404104d55c1530f7670512ce001ba0395433bd253043527e076762

Observation 85c98fc7-0d24-46bb-98ab-dbb71bc6ed53 · outbound

This paper cites an unresolved cited work.

FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-05-17T00:44:01.761645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-17T00:44:01.658214Z digest=sha256:d881f7db20737849e750a4789e413fc8c69acd7b7c35c1b12d2286c20c0e90fa

Observation 18eade63-0879-4ee7-99d4-1a3908f22e33 · outbound

This paper cites an unresolved cited work.

FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-05-17T00:44:01.765058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-17T00:44:01.658214Z digest=sha256:0146eaacb686f4337328f00076ecc9b0fbda3291c8cd609b9e67acae12f51235

Observation 57c6425c-6289-4531-b774-b58208d994f1 · outbound

This paper cites Equality of orders of a set of integers modulo a prime.

FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Equality of orders of a set of integers modulo a prime

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:44:01.679683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-17T00:44:01.658214Z digest=sha256:a8738f83bcfc4e8e212a087b869926ef1ef1e6c4753833f3e9fbeaef23a95434

Observation 473cd1cc-18a3-4f7b-9b8a-2b3dfddfdead · outbound

This paper cites Curves over Finite Fields Attaining the Hasse-Weil Upper Bound.

FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Curves over Finite Fields Attaining the Hasse-Weil Upper Bound

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:44:01.769426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-17T00:44:01.658214Z digest=sha256:bb06a5aeaa16169533d1d7af41f6151d5fd41c43602cde10a13b8e756dc2795a

Observation 3a52ed73-4240-4524-95bb-a66541497403 · outbound

This paper cites an unresolved cited work.

FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-05-17T00:44:01.773294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-17T00:44:01.658214Z digest=sha256:92b47624a328e82801dd27d440e57a20dc0180f9c4d746934ae512b04d3acef7

Observation becd89a7-e3cb-4383-a917-ace9399e0154 · outbound

This paper cites an unresolved cited work.

FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-05-17T00:44:01.779996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-17T00:44:01.658214Z digest=sha256:0e50c5919cc6ab144ea5793404c47152b20df620d112af50a27a1e6d74f1dc31

Pith citing papers

Observation 1f7f29f8-9d01-4dff-8441-24baffdc9061 · inbound

LMAct: A Benchmark for In-Context Imitation Learning with Long Multimodal Demonstrations cites this paper.

LMAct: A Benchmark for In-Context Imitation Learning with Long Multimodal Demonstrations FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-12T04:25:49.785818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:25:49.785818Z digest=sha256:9e43ab30c7626ccd5b9b0acfda4c5364499000facfb1e545840a333afc1236a0

Observation 6864c502-3dbc-4ecf-b948-b44014a6541f · inbound

HARP: A challenging human-annotated math reasoning benchmark cites this paper.

HARP: A challenging human-annotated math reasoning benchmark FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T17:34:19.746349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:34:19.746349Z digest=sha256:8979746ec08eb700920f674cfa50e24ecd876c27ffb2e37fc2ab51d673b22564

Observation c6f1bc17-a186-43bb-873d-5cb9fde76fe0 · inbound

Formal Mathematical Reasoning: A New Frontier in AI cites this paper.

Formal Mathematical Reasoning: A New Frontier in AI FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T10:51:29.659200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:51:29.659200Z digest=sha256:5435717d0493b833327f73ab6604e837589260a72fbe8ea13a4aa212037da64e

Observation ea9ee7eb-ea6e-4860-bd8d-baaff5665698 · inbound

Entropy-Guided Attention for Private LLMs cites this paper.

Entropy-Guided Attention for Private LLMs FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T21:57:08.530011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:57:08.530011Z digest=sha256:457545a49d526dfef1ecc4b87203bba83e6f41e7c227102743ec392a31693b33

Observation b5d9363e-cc9a-47e2-9b9c-a616155cbe79 · inbound

Reasoning Language Models: A Blueprint cites this paper.

Reasoning Language Models: A Blueprint FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-10T18:36:54.483513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:36:54.483513Z digest=sha256:0abae56448d4cebefdbfda9f97c92c277c0229fff2daba9cb0660ad83e33bda8

Observation 76e21ff7-a526-4546-8e6c-a79829391353 · inbound

Not Every AI Problem is a Data Problem: We Should Be Intentional About Data Scaling cites this paper.

Not Every AI Problem is a Data Problem: We Should Be Intentional About Data Scaling FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T15:37:27.991160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:37:27.991160Z digest=sha256:13a3ccc5e0a4a3b13c0490d982114b5695a189fff6981a334fb86e1ffc16b82b

Observation 72a19ef2-e197-4494-ac82-e0e76871d46d · inbound

Humanity's Last Exam cites this paper.

Humanity's Last Exam FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:44:01.791449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:71ea1b2af75594e3f1cf71f9370353478ad2f83aa9f08266d4befc75875d1fdb

Observation 46e64c6b-c79f-4ee2-8b1a-3df5ed868a83 · inbound

SCP-116K: A High-Quality Problem-Solution Dataset and a Generalized Pipeline for Automated Extraction in the Higher Education Science Domain cites this paper.

SCP-116K: A High-Quality Problem-Solution Dataset and a Generalized Pipeline for Automated Extraction in the Higher Education Science Domain FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T14:12:06.445859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:12:06.445859Z digest=sha256:492f59593c81815b9dcdd3b4a6984380b93756e1dedb2639a429c25c76bad0cd

Observation de65e791-2b36-486d-ba3d-ba130315b2eb · inbound

HardML: A Benchmark For Evaluating Data Science And Machine Learning knowledge and reasoning in AI cites this paper.

HardML: A Benchmark For Evaluating Data Science And Machine Learning knowledge and reasoning in AI FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T14:11:16.919552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:11:16.919552Z digest=sha256:47c03c559c7fd9d0627b707264eab45a1cfbfa601d439f588c2032bfa9002826

Observation 51fef532-03b3-44af-b139-58e416fb4e51 · inbound

SocratiQ: A Generative AI-Powered Learning Companion for Personalized Education and Broader Accessibility cites this paper.

SocratiQ: A Generative AI-Powered Learning Companion for Personalized Education and Broader Accessibility FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T19:25:37.103792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:25:37.103792Z digest=sha256:937d5a091c39058c59f861ae50f78348d449f58a982b44c67b2da95453fe41c8

Observation 02db8153-78d2-4c3d-a026-77c3dbe5cb33 · inbound

Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge cites this paper.

Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T18:37:34.857548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:37:34.857548Z digest=sha256:479f9df72c13b7f94833348a706df3fe6968c7b7d17366fe40b653cca2ab409b

Observation 910781e2-8a2c-416b-970d-8f6d5443e576 · inbound

Language Models Use Trigonometry to Do Addition cites this paper.

Language Models Use Trigonometry to Do Addition FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T17:29:00.140227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:29:00.140227Z digest=sha256:9815b235fca97d3afa50e4ea6c6c73fd144272ac2a8d6798668e7d91102d4aa6

Observation 63181b01-6028-476d-a761-cf84cc0c8618 · inbound

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient cites this paper.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-09T18:09:01.604922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:09:01.604922Z digest=sha256:9e237054ee8d2bc86f796179e34cee0907a5262804d5ad5a8fa0e48bbd0a6320

Observation c25b27db-ccc0-4271-b56c-714685194501 · inbound

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models cites this paper.

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 213

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:44:01.791449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-12T08:40:40.910461Z digest=sha256:0cd68c5779b5fcd102adfd569c7b8cc027cbd75ff7b1ffcf7819ce1effb8794a

Observation 736a5ebe-d0da-4fca-a4a8-b215d660a2b4 · inbound

FLIP Reasoning Challenge cites this paper.

FLIP Reasoning Challenge FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T12:39:57.153685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:39:57.153685Z digest=sha256:3e0ed0e94a613440c7199f82d626d8c6a019416c029a8bc52d24ec965e1d703e

Observation dc01b0b9-f190-4b6e-9a82-191da7b33cd8 · inbound

In between myth and reality: AI for math -- a case study in category theory cites this paper.

In between myth and reality: AI for math -- a case study in category theory FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T12:15:57.639459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:15:57.639459Z digest=sha256:44e0fe2b84380416b8e2bac620a7e44d92559a05753d3271e31b4a6d4e421e57

Observation b38bd7a7-1962-4a49-a194-32d3f739b42e · inbound

Virology Capabilities Test (VCT): A Multimodal Virology Q&A Benchmark cites this paper.

Virology Capabilities Test (VCT): A Multimodal Virology Q&A Benchmark FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T11:30:11.854368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:30:11.854368Z digest=sha256:10eb705bbed90e056425baa5cc51b08570f6c98aaf328299315adfae9cfc89fd

Observation 1e5321e4-d706-49ce-91d0-77d03093590d · inbound

Lightweight Latent Verifiers for Efficient Meta-Generation Strategies cites this paper.

Lightweight Latent Verifiers for Efficient Meta-Generation Strategies FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T11:00:20.392061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:00:20.392061Z digest=sha256:4f73fa0376269a913e301bebe74b86d1e25f849d87afd7e9d6c64a2bd78c0d68

Observation 0d6f0768-e0fa-438f-a65e-c991e572e963 · inbound

Toward Generalizable Evaluation in the LLM Era: A Survey Beyond Benchmarks cites this paper.

Toward Generalizable Evaluation in the LLM Era: A Survey Beyond Benchmarks FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T10:12:00.056522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:12:00.056522Z digest=sha256:74144d26db492eaad449eff7330d4e0b714b9c68c0b54401134d551febe73f56

Observation a4c71659-396a-4a72-a822-71c465a159f5 · inbound

Automatic Legal Writing Evaluation of LLMs cites this paper.

Automatic Legal Writing Evaluation of LLMs FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T05:13:51.204065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:13:51.204065Z digest=sha256:7a0f2663e2d544023df949348aad65b12607ebd2fd9c34e7d07466ec75d9dae8

Observation ed9e84da-f5f8-4d37-9d15-a0b2c8c5f419 · inbound

Self-Ablating Transformers: More Interpretability, Less Sparsity cites this paper.

Self-Ablating Transformers: More Interpretability, Less Sparsity FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:45.553635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:45.553635Z digest=sha256:b080c52988790fde5d247ee45471d90171cbbf66ba4ef605d9070ef488d9a137

Observation ecb7b5e6-4fda-4e80-806c-fbdf6f736bf5 · inbound

Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation cites this paper.

Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T04:42:08.218905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:42:08.218905Z digest=sha256:a3e517aa8fe0769ced1bcaefb7684dc0af38315cd9da2db79e3aa66f64d0bbd9

Observation 6bd53163-92e2-4b33-b976-518b9fbc667c · inbound

R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation cites this paper.

R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T04:08:36.839708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:08:36.839708Z digest=sha256:6bc80d9e70026fbfba3c439a792ffc152dfbabddad3ef3d4a4e8214f94f2e4a3

Observation 75cd1dab-74a3-4db0-9f91-c0dee4dfeffc · inbound

Knowledge Augmented Complex Problem Solving with Large Language Models: A Survey cites this paper.

Knowledge Augmented Complex Problem Solving with Large Language Models: A Survey FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T23:56:01.914868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:56:01.914868Z digest=sha256:90bf7e29d9a2621ad29026043438f31bfa3892f849cc72c347be17828bbde19d

Observation 0a259721-15b6-4d1e-aeb6-8a0c23ecdf9f · inbound

R^3-VQA: "Read the Room" by Video Social Reasoning cites this paper.

R^3-VQA: "Read the Room" by Video Social Reasoning FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T23:41:34.116200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:41:34.116200Z digest=sha256:86d8659bec7c284c739906e2412160e824ce58fdcadb43dd9a3f8c6d9842c50d

Observation 7ce6c74f-ab92-42ad-9713-a112a38eb66c · inbound

Beyond Theorem Proving: Formulation, Framework and Benchmark for Formal Problem-Solving cites this paper.

Beyond Theorem Proving: Formulation, Framework and Benchmark for Formal Problem-Solving FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T23:31:49.332889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:31:49.332889Z digest=sha256:ef7b3a0ba84760bfebad10f35e1977d1b248cfca126f2745989d83aae82527a4

Observation 56ff5ae1-4f66-4896-9a39-67551a9bfff4 · inbound

AI Governance to Avoid Extinction: The Strategic Landscape and Actionable Research Questions cites this paper.

AI Governance to Avoid Extinction: The Strategic Landscape and Actionable Research Questions FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-15T23:27:30.027273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:27:30.027273Z digest=sha256:46a6a73b64cf8612b91df15ee20408abf29532c38ac304a3dab05fdd75f88813

Observation 979e5bd1-3c6a-4f60-9573-ba7d0b2a1a11 · inbound

DeepMath-Creative: A Benchmark for Evaluating Mathematical Creativity of Large Language Models cites this paper.

DeepMath-Creative: A Benchmark for Evaluating Mathematical Creativity of Large Language Models FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T21:51:21.300560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:51:21.300560Z digest=sha256:b41e20daac7081dd3f35139ff8338cf18e5cf8060fa506e08e27f06516069895

Observation aeb6b72e-3edc-414f-8f95-0463b5ecea17 · inbound

HARDMath2: A Benchmark for Applied Mathematics Built by Students as Part of a Graduate Class cites this paper.

HARDMath2: A Benchmark for Applied Mathematics Built by Students as Part of a Graduate Class FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T20:53:11.212072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:53:11.212072Z digest=sha256:664e30a108c17793559a53273107d81af0e002d649182d4cdbe3d9ec76d6718a

Observation ed03e151-8c4a-4482-9086-79b502798798 · inbound

Children's Mental Models of AI Reasoning: Implications for AI Literacy Education cites this paper.

Children's Mental Models of AI Reasoning: Implications for AI Literacy Education FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T15:11:00.295349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:11:00.295349Z digest=sha256:42e8e0e6340e40f38263cd31954012ab6491bb7aa9fd6a62a16da0ec602f9266

Observation d4b6457d-799c-446b-a68d-314104f6df77 · inbound

Sudoku-Bench: Evaluating creative reasoning with Sudoku variants cites this paper.

Sudoku-Bench: Evaluating creative reasoning with Sudoku variants FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:08:28.907751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:08:28.907751Z digest=sha256:2eec22eed85fa6dec5f7ce2597c09873c92ee4aef603fb79fc75ae36e85e461d

Observation c0f4e25b-0026-4598-9290-4d4737d2310e · inbound

Effective Reinforcement Learning for Reasoning in Language Models cites this paper.

Effective Reinforcement Learning for Reasoning in Language Models FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:16.257566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:56:16.257566Z digest=sha256:638afb5d97765dedcf8363ce3d72f9e6ec2bca3ba086852b190a01ddb4efdd15

Observation b5da3b6b-af00-4747-ba97-d1321d3c4645 · inbound

CapBencher: Give Your LLM Benchmark a Built-in Alarm for Test-Set Overfitting cites this paper.

CapBencher: Give Your LLM Benchmark a Built-in Alarm for Test-Set Overfitting FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:44:33.791384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:44:33.791384Z digest=sha256:2c0e45b2ed2161d617400acab2c2c41386cd0a6bb5ce9201957fd9d632ff8453

Observation 4857d67b-0746-4bc9-ad44-c737b17a1ac4 · inbound

DFIR-Metric: A Benchmark Dataset for Evaluating Large Language Models in Digital Forensics and Incident Response cites this paper.

DFIR-Metric: A Benchmark Dataset for Evaluating Large Language Models in Digital Forensics and Incident Response FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:08:22.589495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:08:22.589495Z digest=sha256:e0511d6efe73b894e9e1d932c93a8e751c22babcae4a5822a429156a1bb8b149

Observation 634a25d9-e61e-4648-8ad9-d15b3c1048fc · inbound

ASyMOB: Algebraic Symbolic Mathematical Operations Benchmark cites this paper.

ASyMOB: Algebraic Symbolic Mathematical Operations Benchmark FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:04.306909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:05:04.306909Z digest=sha256:490bead8086d5467209ae82ccc3191bab3c8345981f52a607417dd5d5b467a24

Observation 1fa19fda-577b-4940-af64-d789bac285dc · inbound

Using Reasoning Models to Generate Search Heuristics that Solve Open Instances of Combinatorial Design Problems cites this paper.

Using Reasoning Models to Generate Search Heuristics that Solve Open Instances of Combinatorial Design Problems FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:43.891256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:43.891256Z digest=sha256:ddb853397d8774802886370d6ec601afe3f37023a38c759349d1f4876bd53c6d

Observation bd1ea47d-1177-45b1-a559-e261b22f9628 · inbound

Circuit Stability Characterizes Language Model Generalization cites this paper.

Circuit Stability Characterizes Language Model Generalization FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:37.237963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:37.237963Z digest=sha256:b606f2c974bf74b925391c74ccba204144837e7525b0c63670329bf92045d794

Observation 83b0f156-b75d-481a-b848-ea1442e1e07c · inbound

Literature Review Of Multi-Agent Debate For Problem-Solving cites this paper.

Literature Review Of Multi-Agent Debate For Problem-Solving FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:47:35.915364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:47:35.915364Z digest=sha256:d8294e42643ecac002cf782e6cdfd7bc1654dcefdc143a9f5251a16c62c293df

Observation e3c96934-e189-4598-8b84-cb32a3653808 · inbound

Breakpoint: Scalable evaluation of system-level reasoning in LLM code agents cites this paper.

Breakpoint: Scalable evaluation of system-level reasoning in LLM code agents FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:17:44.198051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:17:44.198051Z digest=sha256:517be7c72a1de56d070bb5c808f9e9c286663bc4a8a37864123ad73266919c0c

Observation 55da0497-55b0-488e-89c7-1624a72df77d · inbound

AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions cites this paper.

AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:08.185061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:08.185061Z digest=sha256:78f7f4c40d138b0008b3bc8455d53cd4b53649837860d73eb5988bd761ce4e29

Observation 435a24fe-d609-42ce-8590-fbf85262dc76 · inbound

RE-IMAGINE: Symbolic Benchmark Synthesis for Reasoning Evaluation cites this paper.

RE-IMAGINE: Symbolic Benchmark Synthesis for Reasoning Evaluation FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T19:40:08.739793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:40:08.739793Z digest=sha256:2b5a9daba5bc9114c8a177d7494f87627e9f73d20eba4b7d4c38fb580272224a

Observation 2fc04ec4-6484-4b78-9075-bd038d42f644 · inbound

Establishing Best Practices for Building Rigorous Agentic Benchmarks cites this paper.

Establishing Best Practices for Building Rigorous Agentic Benchmarks FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:22.159924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:22.159924Z digest=sha256:b0d05665276f2f25d82e43313d4dbd2117841dcf9696d03bc0a4bae7ab9d257f

Observation 2d40c2e4-2cad-4008-bc25-793a4dfba619 · inbound

Probing for Arithmetic Errors in Language Models cites this paper.

Probing for Arithmetic Errors in Language Models FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T16:58:55.677386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:58:55.677386Z digest=sha256:3ce7b91ca6463d434f3e153b4fa224c368b2efafd80f77ce7c76e2993dcc64cf

Observation 4a329d34-b70a-4bae-9e30-03f4eb0d9059 · inbound

FormulaOne: Measuring the Depth of Algorithmic Reasoning Beyond Competitive Programming cites this paper.

FormulaOne: Measuring the Depth of Algorithmic Reasoning Beyond Competitive Programming FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T16:30:22.229875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:30:22.229875Z digest=sha256:814a31f408ff25d2b0975d3c25639f4ce8d2eddb61dff40eff5711f28a9f6e70

Observation 834de928-4ab6-4b41-a962-c2fad6bd8d13 · inbound

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks cites this paper.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.307985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.307985Z digest=sha256:433250aaf7dae218794e2ea98ed22831d06c71d775aba79f30fd80a25c29a39b

Observation f3960809-b0e2-4714-8998-6068c3e924f0 · inbound

Proof2Hybrid: Automatic Mathematical Benchmark Synthesis for Proof-Centric Problems cites this paper.

Proof2Hybrid: Automatic Mathematical Benchmark Synthesis for Proof-Centric Problems FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T05:11:53.824313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:11:53.824313Z digest=sha256:036de5d83d7dc1c9e0a51bcaf5754bd139d46b65f3f0b8b6e9ba074274c88703

Observation 743df680-9c39-41e3-b4dc-e8ddf532273d · inbound

The Mathematician's Assistant: Integrating AI into Research Practice cites this paper.

The Mathematician's Assistant: Integrating AI into Research Practice FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T15:18:07.646360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:18:07.646360Z digest=sha256:8a023b692a1e1edde78d9b93da6d7c7ad5f34a8a5dbc64d56f140070f185cf41

Observation 3f8136b4-8bc9-4afc-83af-b5331c3858d3 · inbound

AI Reasoning Models for Problem Solving in Physics cites this paper.

AI Reasoning Models for Problem Solving in Physics FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T14:45:39.976190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:45:39.976190Z digest=sha256:67ff9793f7388b9ed25f3269b1af944fa206026bb846c22b2a87e355d9a4ec66

Observation 57ac86d0-1350-4f84-bccc-431fd86cc74f · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 163

Resolution
verified exact
local_arxiv, observed 2026-05-18T00:02:25.462023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:a5a276a194132e1333db824d269c5fd02495ec4921dce31f7a4ae9b3d8b95398

Observation cff174b5-0057-4edf-ad18-0884207b7c46 · inbound

CFDLLMBench: A Benchmark Suite for Evaluating Large Language Models in Computational Fluid Dynamics cites this paper.

CFDLLMBench: A Benchmark Suite for Evaluating Large Language Models in Computational Fluid Dynamics FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 19

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T15:06:32.104836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-18T15:05:37.519850Z digest=sha256:e811d2a176a82f524fdfa80bddb97dfc9846ab4528bf01e15544f8e0a1d85ae2

Observation 000180cb-00e4-4ce2-ac62-83df23cad931 · inbound

Probing the Critical Point (CritPt) of AI Reasoning: a Frontier Physics Research Benchmark cites this paper.

Probing the Critical Point (CritPt) of AI Reasoning: a Frontier Physics Research Benchmark FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-05-18T11:52:35.313977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-18T11:52:10.205796Z digest=sha256:442ab077bc152a68086c9eb920d1c976cac35cd264eae5b774a975b81552dc1e

Observation 39cbd854-f23d-4167-9e22-43d6ae7d89bb · inbound

Co-Designing Quantum Codes with Transversal Diagonal Gates via Multi-Agent Systems cites this paper.

Co-Designing Quantum Codes with Transversal Diagonal Gates via Multi-Agent Systems FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-18T04:30:52.976368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-18T04:28:48.400311Z digest=sha256:070b97f1ead3a0e4b2e4ffcdfe6c7a3c170366dd5742ea7eb26f5efd6e976075

Observation 14fa0979-dca1-44e5-bb23-9932c7c8fdba · inbound

FEM-Bench: A Structured Scientific Reasoning Benchmark for Evaluating Code-Generating LLMs cites this paper.

FEM-Bench: A Structured Scientific Reasoning Benchmark for Evaluating Code-Generating LLMs FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T14:21:27.147042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:21:27.147042Z digest=sha256:1d5aef5d478bbd00689b3d1c2f9e024703a5f5e0ce54d61f70f1bd84d8f7e545

Observation ee92b023-b7d2-482f-bf70-ed8868aa70e3 · inbound

AI for Mathematics: Progress, Challenges, and Prospects cites this paper.

AI for Mathematics: Progress, Challenges, and Prospects FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:44:01.791449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-16T13:24:57.923863Z digest=sha256:6ee09132973cb81e442f0a10796938253adcfe993915fe8b2118ee265ee09d8f

Observation bb6e0504-f392-47a6-937c-f5cdb0f72e5b · inbound

Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation cites this paper.

Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T03:04:43.601158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:04:43.601158Z digest=sha256:91b5def757bf67b64b4353bb41fbcea4b28651f5e01284fa4a312feb0ffcc2fa

Observation 6e6d869f-eb4c-4f57-9385-cd585674287b · inbound

QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs cites this paper.

QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T21:18:01.374648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:18:01.374648Z digest=sha256:9ddda2e9fcd862ad78d9c9ebb782c6ac53009e4cdfa5071ddb534aedaf2e8c8a

Observation 32b11f53-8652-454d-89ae-f51e74600799 · inbound

LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics cites this paper.

LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T20:05:56.402050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:05:56.402050Z digest=sha256:21cf66cbda6c219807c051ecbee00ce84e88a2b124b54a49f2d24e575588c94b

Observation 308e4fcb-fae1-4860-b996-7d03eadb80bb · inbound

Can LLMs Prove Robotic Path Planning Optimality? A Benchmark for Research-Level Algorithm Verification cites this paper.

Can LLMs Prove Robotic Path Planning Optimality? A Benchmark for Research-Level Algorithm Verification FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-13T22:02:43.946867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:02:43.946867Z digest=sha256:bf55653406730804840a0c62bad4987ab6c4121752ecaa91491774ed37ddb7a8

Observation 2922276d-b6af-4b89-bb8e-79ac2994d2b0 · inbound

Can LLMs Prove Robotic Path Planning Optimality? A Benchmark for Research-Level Algorithm Verification cites this paper.

Can LLMs Prove Robotic Path Planning Optimality? A Benchmark for Research-Level Algorithm Verification FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T17:51:37.623160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:51:37.623160Z digest=sha256:422308005b42fe5a78a04769a4b26e17fe62af59906022de5316fbed5d181238

Observation c9e23c67-577a-495e-a4d2-f5ea1c917233 · inbound

Xpertbench: Expert Level Tasks with Rubrics-Based Evaluation cites this paper.

Xpertbench: Expert Level Tasks with Rubrics-Based Evaluation FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:44:01.791449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-14T23:22:53.368231Z digest=sha256:8cb7e55e9690913d57ef83e8842fbe04b5d3e45eb476b5df4ac065332d3b1f8d

Observation 5180c4eb-610f-46eb-8faa-32218492a12c · inbound

Automated Conjecture Resolution with Formal Verification cites this paper.

Automated Conjecture Resolution with Formal Verification FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:44:01.791449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-13T18:05:38.340467Z digest=sha256:f68c9aaf5b1e0765f3063534c50034390e5fb45f3e99c06623bac2c8b846e942

Observation 1dd9f079-99d9-4385-b390-bdd2948c03bb · inbound

Automated Conjecture Resolution with Formal Verification cites this paper.

Automated Conjecture Resolution with Formal Verification FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-13T12:18:10.165894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T12:18:10.165894Z digest=sha256:02fb6a24abf926ecabb7508f7d2dcf269df12a014f1239a858ff18f4ed487303

Observation 0f9bcabb-242f-4a8a-8a44-f1e791c99e28 · inbound

Automatically Generating Hard Math Problems from Hypothesis-Driven Error Analysis cites this paper.

Automatically Generating Hard Math Problems from Hypothesis-Driven Error Analysis FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:44:01.791449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-10T19:39:36.819142Z digest=sha256:af83b699f4e427114dd42c0843c7dbd5a7b24bcb44308fc13ffd03a1cd33da07

Observation b616c540-6143-40cb-b0b4-5de9c626106e · inbound

DeonticBench: A Benchmark for Reasoning over Rules cites this paper.

DeonticBench: A Benchmark for Reasoning over Rules FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:44:01.791449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T20:22:03.956227Z digest=sha256:60b451609d0bed323f34c88613d3f2a83c5005685a2676ea5ec55d7ea169baea

Observation 76f947a2-81ff-475f-ac0e-f6633f17f381 · inbound

Artificial Intelligence and the Structure of Mathematics cites this paper.

Artificial Intelligence and the Structure of Mathematics FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 45

Resolution
malformed identifier
arxiv_id, observed 2026-05-17T00:44:01.791449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T19:12:23.402808Z digest=sha256:936ef89e1751b1f76d686c7f0843de04e87b66ff9a9cc65516804c544ecbe65f

Observation a37c7062-38c5-463e-bf5b-5293a35583bf · inbound

Riemann-Bench: A Benchmark for Moonshot Mathematics cites this paper.

Riemann-Bench: A Benchmark for Moonshot Mathematics FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:44:01.791449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T18:02:42.607682Z digest=sha256:52336f9098ae727fd4285decfa2a1e89de7ff2cff63360af8168c5810d9f2d28

Observation 768a1353-c835-449e-96c1-b29c33efa2ea · inbound

$k$-server-bench: Automating Potential Discovery for the $k$-Server Conjecture cites this paper.

$k$-server-bench: Automating Potential Discovery for the $k$-Server Conjecture FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:44:01.791449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T17:48:17.363568Z digest=sha256:8173e13076adeed3d67f5e0e39701536e419ed351f42bec73a2ca0e60c8b25ff

Observation 895512bc-89d5-4ec5-8bc9-57fb48e1f44f · inbound

Problem Reductions at Scale: Agentic Integration of Computationally Hard Problems cites this paper.

Problem Reductions at Scale: Agentic Integration of Computationally Hard Problems FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:44:01.791449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T16:07:35.395872Z digest=sha256:4691651c4cc25d8a994e9ef93d6f250d37d300cca5856f8913b3971dd5f3afef

Observation 363b1d25-5919-476f-994f-1a610104efaf · inbound

Agentic Frameworks for Reasoning Tasks: An Empirical Study cites this paper.

Agentic Frameworks for Reasoning Tasks: An Empirical Study FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 55

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T00:44:01.791449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T08:24:51.573913Z digest=sha256:477d63a9a824d27eb96a6092d3eebbf8d15501ba5c11dda7e5df85ccef8b19a9

Observation ce61519e-7cd9-4cef-946f-e8e6ffa34df4 · inbound

Fine-Tuning Small Reasoning Models for Quantum Field Theory cites this paper.

Fine-Tuning Small Reasoning Models for Quantum Field Theory FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:44:01.791449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T03:23:18.770963Z digest=sha256:ad0e5e4ad11bad7e661b4f6718f391bc4b1f45229a5970eee08875969d6fe0d3

Observation 7cc1818a-2bf0-4657-b3a6-8b82b28eb397 · inbound

MathDuels: A Self-Play Benchmark That Grows cites this paper.

MathDuels: A Self-Play Benchmark That Grows FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T00:44:01.791449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-09T21:26:52.011933Z digest=sha256:03bcbc85ab013ec831d7bff384fd81adb04108521a2151d58600cb067bca0dc0

Observation 923968c8-a34e-4d42-86a0-63c34adefdb0 · inbound

Forge: Quality-Aware Reinforcement Learning for NP-Hard Optimization in LLMs cites this paper.

Forge: Quality-Aware Reinforcement Learning for NP-Hard Optimization in LLMs FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T00:44:01.791449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-12T02:44:33.143247Z digest=sha256:9aabc509c966a8fefd07a6aeef7ad45b8ffea6ce9edc6d6257c4dfe8e76e3c8a

Observation 5bd5622d-e6d2-4695-a04f-e7a1e94b9827 · inbound

Soohak: A Mathematician-Curated Benchmark for Evaluating Research-level Math Capabilities of LLMs cites this paper.

Soohak: A Mathematician-Curated Benchmark for Evaluating Research-level Math Capabilities of LLMs FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T00:44:01.791449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-12T01:53:12.708706Z digest=sha256:58a71c0e708b60973b989aed4e6b19c5b30072de477d87cca4dc0a67e12d8de7

Observation 5d266f89-6a95-46b4-a4c1-16f6e7516e9e · inbound

Soohak: A Mathematician-Curated Benchmark for Evaluating Research-level Math Capabilities of LLMs cites this paper.

Soohak: A Mathematician-Curated Benchmark for Evaluating Research-level Math Capabilities of LLMs FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T22:24:07.822738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T22:24:04.625823Z digest=sha256:722d30ea17406c2c32cfc39cda70d59eb20f68b61686480ca6be9f814d2963d5

Observation 430054da-3fda-405e-bbab-74a13dba2b83 · inbound

Formal Conjectures: An Open and Evolving Benchmark for Verified Discovery in Mathematics cites this paper.

Formal Conjectures: An Open and Evolving Benchmark for Verified Discovery in Mathematics FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:44:01.791449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-14T20:20:54.566662Z digest=sha256:8412419942fe3a245671579478a633c762a02fa7596221838c37a4f447ab41a7

Observation d2c6d8b1-7f5b-48b0-a474-40e12f42420a · inbound

STAR-P\'olyaMath: Multi-Agent Reasoning under Persistent Meta-Strategic Supervision cites this paper.

STAR-P\'olyaMath: Multi-Agent Reasoning under Persistent Meta-Strategic Supervision FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-20T02:58:00.175045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T02:56:19.104905Z digest=sha256:db4b6fe85cfc2a858f81fbd2634f7cbfa85367094dc7b572d224d3f89b049fc0

Observation 12f99899-ecf2-47ab-be40-905cbd9e37af · inbound

Scientific reasoning does not reliably translate into scientific forecasting in frontier AI cites this paper.

Scientific reasoning does not reliably translate into scientific forecasting in frontier AI FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-05-22T05:21:07.455360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-22T05:18:15.358571Z digest=sha256:0041b2a2897725827418656f17f273b8e18e784283f266c02f276385a4241702

Observation 7c4509d2-bd4f-46b1-865c-5527f222b8df · inbound

Scientific reasoning does not reliably translate into scientific forecasting in frontier AI cites this paper.

Scientific reasoning does not reliably translate into scientific forecasting in frontier AI FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T13:31:46.340326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:31:46.340326Z digest=sha256:e28d90b14b983e4bf495ad390e77578e9266c4bdd54ee984b335e1b414a5d919

Observation 3da02bf1-3fb3-4eff-ae75-4c9058268dee · inbound

RMA: an Agentic System for Research-Level Mathematical Problems cites this paper.

RMA: an Agentic System for Research-Level Mathematical Problems FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-25T06:10:24.159655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-25T06:06:47.226726Z digest=sha256:b694fdb33429e4f0067f35fbe65f97d077c4d6db942a237c62347a999b560528

Observation d8931d30-f951-4dcd-8dfe-c5b8af77d71b · inbound

MadEvolve: Evolutionary Optimization of Trading Systems with Large Language Models cites this paper.

MadEvolve: Evolutionary Optimization of Trading Systems with Large Language Models FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-25T05:25:23.375844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-25T05:24:05.098181Z digest=sha256:29867dec80442c47bfd9e70cb000d79a0b63f3cf2f14cca4c43dc93ad48427c7

Observation 962229af-4518-4832-8218-d9276d110c0a · inbound

Self-Improving Language Models with Bidirectional Evolutionary Search cites this paper.

Self-Improving Language Models with Bidirectional Evolutionary Search FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:03:26.436991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T12:57:28.931392Z digest=sha256:b5f27ce0ca10bc877621edfa1d8fdd7b79ef58f3511f66c623353015939db419

Observation d2a75c58-b04b-43a3-be07-bb4cf9de881c · inbound

Why Larger Models Learn More: Effects of Capacity, Interference, and Rare-Task Retention cites this paper.

Why Larger Models Learn More: Effects of Capacity, Interference, and Rare-Task Retention FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-06-29T08:43:15.141444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T08:39:14.327838Z digest=sha256:7024b187ccc8d433327690e9dd99005a8c3053452334459ccaa6bbb1d49f01cf

Observation 9e5c041c-014f-41ab-b5d9-1825db7f49f9 · inbound

Evaluating Interactive Reasoning in Large Language Models: A Hierarchical Benchmark with Executable Games cites this paper.

Evaluating Interactive Reasoning in Large Language Models: A Hierarchical Benchmark with Executable Games FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T17:53:47.431723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-29T17:46:39.281623Z digest=sha256:7ee8f0b04e7843d2d670816ad24f4c5707b00156d15d9487ba0ba1c43458ecb8

Observation d07eae8e-f869-471d-81ad-8562e008dc17 · inbound

FVSpec: Real-World Property-Based Tests as Lean Challenges cites this paper.

FVSpec: Real-World Property-Based Tests as Lean Challenges FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-06-28T17:12:25.298307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-28T17:05:13.012431Z digest=sha256:aebc35459b5b3b8b96acbacd88b3b4504b3659d6b1686b8e08b418d80f2e1343

Observation 731c206a-b6f2-484b-b76c-575c46a49c03 · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 45

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T20:56:13.250290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:6cef8ad39e7cf14b7494249394a5ea720be6dad8807ce540496ba51bca53733f

Observation dd0dcb8d-b48a-49d9-b464-a09ed08527f9 · inbound

An Enigma of Artificial Reason: Investigating the Production-Evaluation Gap in Large Reasoning Models cites this paper.

An Enigma of Artificial Reason: Investigating the Production-Evaluation Gap in Large Reasoning Models FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-01T21:36:14.993846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-28T16:45:33.046568Z digest=sha256:73183d066f61f5f1e9cce4f492ebc46d85b57b99fcb63a3608c211465cc073a1

Observation abeef3e5-3717-44f4-bdf2-0905230e334a · inbound

Lean-GAP: A Dataset of Formalized Graduate Algebra Problems cites this paper.

Lean-GAP: A Dataset of Formalized Graduate Algebra Problems FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-06-30T17:34:57.542083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T17:32:00.411535Z digest=sha256:991a8d3b94869925508c2dfdbb40e2087ef9c45050efab390337015e4b0ee5f7

Observation fbfee4f5-4113-4d26-bbf3-a5866889807a · inbound

GTBench: A Curriculum-Grounded Benchmark for Evaluating LLMs as Mathematical Research Assistants in Graph Theory cites this paper.

GTBench: A Curriculum-Grounded Benchmark for Evaluating LLMs as Mathematical Research Assistants in Graph Theory FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-02T03:06:30.296889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-28T10:19:03.670540Z digest=sha256:0d5ccf372e427a3656b0a050ddfe1142243dd1e3eb01b4177cf5f434f6bc980e

Observation f31297e3-a2f6-43b1-ad07-d2d819cab9ad · inbound

LeanMarathon: Toward Reliable AI Co-Mathematicians through Long-Horizon Lean Autoformalization cites this paper.

LeanMarathon: Toward Reliable AI Co-Mathematicians through Long-Horizon Lean Autoformalization FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T08:36:48.761264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-28T05:48:56.691155Z digest=sha256:212675f579cd4ba358c4c91f5bc5ed0488f5cf0f1f8a98bfabcc64b82948de7b

Observation d8702a46-4bd0-4b1f-a18b-21b44c265c7b · inbound

Automated Proving of Shannon-Type Entropy Inequalities via Fine-Tuned Language Models and Guided Tree Search cites this paper.

Automated Proving of Shannon-Type Entropy Inequalities via Fine-Tuned Language Models and Guided Tree Search FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-07-02T15:27:05.514373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-27T23:52:11.730333Z digest=sha256:86a4d92b8a83319883aea7988857870cbc0cb18195445c2dedf36af5dee56143

Observation ab30e99c-aef7-4f34-ab18-747f9c148845 · inbound

CrowdMath: A Dataset of Crowdsourced Mathematical Research Discussions cites this paper.

CrowdMath: A Dataset of Crowdsourced Mathematical Research Discussions FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 20

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T03:46:32.299057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-28T09:45:18.385925Z digest=sha256:d77019da84e829114369eb67cd58e33b0b129e0bfcebde6b3363de19f1ce0ef5

Observation 2a7de1a8-87b1-4cf5-9585-e2d5fb46f9fa · inbound

How reliable are LLMs when it comes to playing dice? cites this paper.

How reliable are LLMs when it comes to playing dice? FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-02T17:37:14.662833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-27T21:56:42.766016Z digest=sha256:3ffaac5d4b24ccf9063e039bd5f883b77d0dce728c9638eee5bb7fb2678c5d80

Observation 318f2566-3ec4-49ab-b70d-8dc68026cf77 · inbound

Artificial Intelligence for Mathematical Reasoning: An Integrated Survey of Language Models, Neuro-symbolic Systems, and Verified Discovery cites this paper.

Artificial Intelligence for Mathematical Reasoning: An Integrated Survey of Language Models, Neuro-symbolic Systems, and Verified Discovery FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 210

Resolution
verified exact
local_arxiv, observed 2026-07-02T22:47:25.916162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-27T18:39:44.696961Z digest=sha256:50bcccda83e564d7bb95296c216e57a4dc5d4096f102be0637efda21d1d01e2c

Observation 37473251-e072-4cb7-8b7d-60e32eb136ed · inbound

Artificial Intelligence for Mathematical Reasoning: An Integrated Survey of Language Models, Neuro-symbolic Systems, and Verified Discovery cites this paper.

Artificial Intelligence for Mathematical Reasoning: An Integrated Survey of Language Models, Neuro-symbolic Systems, and Verified Discovery FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 212

Resolution
unresolved
no resolver link, observed 2026-08-02T12:05:18.237071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:05:18.237071Z digest=sha256:f49e5d7095933e286bd8f8965a277024ba781b7767d87529f4bde96c1e0e773c

Observation e91a0f67-a638-4150-85f0-45cfe8d6fe91 · inbound

Sakana Fugu Technical Report cites this paper.

Sakana Fugu Technical Report FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T06:39:36.939254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-26T14:22:37.596720Z digest=sha256:a78e27c72885c285385573977cdd3a36e6c7b7a723e2c6273b0a9ce81c678a4c

Observation 5e59d95f-482f-47d8-b827-8ebecf666c0f · inbound

Learning the ARTS of Search for Automated Discovery cites this paper.

Learning the ARTS of Search for Automated Discovery FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-04T08:09:41.810861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-26T12:04:13.117307Z digest=sha256:f8a3d31201e5c0840ff200e4ccf64f019007bc50eeafcfe536dd1f2987ee6728

Observation f8d10b29-0fe4-4ea3-b166-67a56fbff509 · inbound

Curriculum Reinforcement Learning Can Incentivize Reasoning Capacity in LLMs Beyond the Base Model cites this paper.

Curriculum Reinforcement Learning Can Incentivize Reasoning Capacity in LLMs Beyond the Base Model FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-04T08:39:42.300162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-26T11:10:56.755558Z digest=sha256:605d9b905132848dd1572b0d339dafca2b13d37bfb7aaf29995edca71998e012

Observation 14626137-7b07-4b29-9dd6-d2defe38cf00 · inbound

Theorist Toolbox: Tools for Agent Based LLM-assisted economic theory Research cites this paper.

Theorist Toolbox: Tools for Agent Based LLM-assisted economic theory Research FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-04T09:29:43.421885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-26T10:01:18.103014Z digest=sha256:3d3b74f3359eef18b5cf13b7f2dc9c1a8cd3a7163e8cf14c49e05310b5e5586e

Observation d03d779f-a66c-45e2-8812-73b4a58be69d · inbound

Beyond Shapley: Efficient Computation of Asymmetric Shapley Values cites this paper.

Beyond Shapley: Efficient Computation of Asymmetric Shapley Values FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T18:20:04.548897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-25T22:56:30.582241Z digest=sha256:70ca2ae87e599aaaa9d038b9ba6d775c11889582783bc442a8de3a4b3ad0f301

Observation 02c26e77-4dcd-44f5-80b0-124a5c2740cf · inbound

Data and Evaluation Closed-Loop for Model Capability Enhancement cites this paper.

Data and Evaluation Closed-Loop for Model Capability Enhancement FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T15:25:48.060133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-30T01:33:08.134048Z digest=sha256:b0929b9097b283289e8c4409f39d2bf7d385d51f3a68749c4f64c28d06fad4c4