Pith. sign in

Paper Citation Record · LEDGER

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

As of 11 August 2026, this Paper Citation Record lists 77 of 77 outbound references and 79 inbound Pith citation observations for arXiv:2410.07985.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.07985 v3

Coverage vector

measured 77 of 77 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-15T09:09:14.884516Z

measured 156 of 156 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 79 of 79 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T00:22:29.920179Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

77 of 77 outbound references displayed

  • verified exact19
  • verified fuzzy48
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch8

External citation measurements

3
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 111cc7e6-2451-4361-aa0d-ca88ac5a2b0e · outbound

This paper cites 2021 , eprint=.

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models 2021 , eprint=

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T09:09:15.117310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-15T09:09:14.884516Z digest=sha256:05c9b11b23ad8317441d9040fac25b37826f62338320313665f9cf43e9f399ee

Observation e9aeaf62-f410-429f-9999-35287294270d · outbound

This paper cites 2021 , eprint=.

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models 2021 , eprint=

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T09:09:15.150144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-15T09:09:14.884516Z digest=sha256:dd31b9f8faa833d819f40a8c6a6340b66267e51a8a9861c5af366beed5ec1ebd

Observation 3af6d83b-34fc-4a7a-85f1-b2e1da473f04 · outbound

This paper cites 2023 , eprint=.

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models 2023 , eprint=

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T09:09:15.152511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-15T09:09:14.884516Z digest=sha256:fa6d0022eea96af8fa1fe436e8812958f34bfb0b79c5e348d72657026c30f439

Observation 505469b3-8300-4c43-842f-44e6e7ce5d39 · outbound

This paper cites 2022 , eprint=.

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models 2022 , eprint=

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T09:09:15.154937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-15T09:09:14.884516Z digest=sha256:81f79bbc6915ddd7e6afba7846f4115fc77efc73cc7dda707c10b1a935ae112c

Observation 589dd251-4d28-460a-a33e-6d168d24bb01 · outbound

This paper cites 2023 , eprint=.

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models 2023 , eprint=

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T09:09:15.157422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-15T09:09:14.884516Z digest=sha256:95cffda4be9d21771f28ad1a04b4c09dea6d90e577b41f67ec4bcd63fbbddea9

Observation c6d9f78e-e229-40f4-896a-0c92ffbe1de5 · outbound

This paper cites 2024 , eprint=.

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models 2024 , eprint=

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T09:09:15.159911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-15T09:09:14.884516Z digest=sha256:b16a3fe7608c82f6ff63531b76d395cfdf878ee3009952d53ec9c7277a14ce97

Observation 7b13af36-1990-4106-9565-642bbbdbed9d · outbound

This paper cites 2024 , eprint=.

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models 2024 , eprint=

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T09:09:15.162210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-15T09:09:14.884516Z digest=sha256:49496962694a45af7fe003fe56b7614e06a79ca275357e5c1b17cef9c26fda78

Observation 9e155c53-2963-40f3-94ab-6a628b2580f0 · outbound

This paper cites 2024 , eprint=.

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models 2024 , eprint=

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T09:09:15.164666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-15T09:09:14.884516Z digest=sha256:b6f258041f3fe6f13241c32217163734947ddf8258d30865e62d492c57ea94ff

Observation 3c47b969-6e11-404f-b692-08b0e7f5e12c · outbound

This paper cites 2024 , eprint=.

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models 2024 , eprint=

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T09:09:15.167404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-15T09:09:14.884516Z digest=sha256:93b7299dd92791175b9bd768be4d90e559af8ea7a85b68c4daa6e741e89f51e2

Observation 12844f27-7b36-4e0f-a029-e27d43e9de05 · outbound

This paper cites 2022 , eprint=.

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models 2022 , eprint=

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T09:09:15.169670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-15T09:09:14.884516Z digest=sha256:5ea0e1f208ed1174e3c4affb3c8962da3a419da07861366bd24ac4de4710772c

Observation 642ec3fe-198c-440f-b011-7be415e9772f · outbound

This paper cites Nature , volume=.

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models Nature , volume=

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T09:09:15.173492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-15T09:09:14.884516Z digest=sha256:2553828cbf7d8865745af8b12c27d6c07e68c3066a48fd3e58899ba3b268aee2

Observation 50dc1e93-050b-4592-bd36-586791f598f7 · outbound

This paper cites Hugging Face repository , howpublished =.

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models Hugging Face repository , howpublished =

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T09:09:15.176750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-15T09:09:14.884516Z digest=sha256:1d645e97dd27b177492a951450cd4980873ad174e26396b75237c93956aa3ba6

Observation d485c82f-a1ce-4ae3-b53c-af0ba529e8ff · outbound

This paper cites 2024 , journal=.

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models 2024 , journal=

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T09:09:15.179645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-15T09:09:14.884516Z digest=sha256:09bf6b8954243e85db2edefb8d27dfcd6b96eb16c16af6fb13c3367cbe3b934c

Observation e3d42753-0873-4f37-bad3-b09685b7b503 · outbound

This paper cites 2023 , eprint=.

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models 2023 , eprint=

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T09:09:15.182379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-15T09:09:14.884516Z digest=sha256:e98654dfba098963e74fffcf05b2765f2f501e7d92003fe9725ed5511e1777bb

Observation f52c174f-8e6a-4658-9266-560fa96a752f · outbound

This paper cites 2023 , eprint=.

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models 2023 , eprint=

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T09:09:15.185389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-15T09:09:14.884516Z digest=sha256:0520b79c5b955197dfccf4e302d58855e8f979f9665c08fb4d8fc743a2decb30

Observation 754cb1f7-23a0-479e-93f7-77da931dd116 · outbound

This paper cites 2024 , eprint=.

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models 2024 , eprint=

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T09:09:15.188878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-15T09:09:14.884516Z digest=sha256:ba64438f139de80a3a00a40bab4f75a22de3f03e3e4c382ab72d4802193c11b4

Observation 366d2d5f-a2f8-4c43-a123-f5d006708522 · outbound

This paper cites 2024 , eprint=.

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models 2024 , eprint=

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T09:09:15.191825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-15T09:09:14.884516Z digest=sha256:9a7e4727d500f2f3cc6e95c035afb26df9d40e66de1f74792efe2598a0c0c2cd

Observation 6ea17305-02fe-442e-b22d-4cd4e769fd1a · outbound

This paper cites Journal of the American statistical Association , volume=.

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models Journal of the American statistical Association , volume=

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T09:09:15.194968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-15T09:09:14.884516Z digest=sha256:524bea41ac0e59eb6ab5d83f4f0880add2a9f192e69623f8fcf6992233f01816

Observation f5c263ef-df71-4895-b674-4f1e1af72780 · outbound

This paper cites 2023 , url=.

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models 2023 , url=

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T09:09:15.197398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-15T09:09:14.884516Z digest=sha256:85db4e399c4e960eb8f289cb1869f77d613ed86586eadf5f2b770defa27f3638

Observation adddf800-f79b-437f-94f4-5494c36538ea · outbound

This paper cites 2024 , eprint=.

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models 2024 , eprint=

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T09:09:15.199781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-15T09:09:14.884516Z digest=sha256:7f9c42767ed49eb907367b03b39071d6aa4dd780ea015234c7fad6e7d399d450

Observation a3a10510-fd94-4695-a29b-e60e04daffcc · outbound

This paper cites 2024 , eprint=.

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models 2024 , eprint=

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T09:09:15.202197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-15T09:09:14.884516Z digest=sha256:97d56a687c755e01a8d4af337e67585b4d69e893f067e1b2437af6cdc1f0a574

Observation 706739d1-378b-48fe-9511-b442ee8857c0 · outbound

This paper cites 2024 , eprint=.

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models 2024 , eprint=

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T09:09:15.204467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-15T09:09:14.884516Z digest=sha256:60205d13ad233b4d3009e97e9e3895719124efb73959ff98f0c2eff288b923da

Observation bf6928b8-9e69-447e-9d7a-06c97cde6188 · outbound

This paper cites 2024 , eprint=.

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models 2024 , eprint=

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T09:09:15.206638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-15T09:09:14.884516Z digest=sha256:d80f85f644b88c8f4d9e5e5584c09ed6beaa959d8ce6f5b1ea29e0f187ca5732

Observation 57804fda-d753-418c-8e5b-7738767212dc · outbound

This paper cites 2024 , eprint=.

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models 2024 , eprint=

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T09:09:15.208772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-15T09:09:14.884516Z digest=sha256:7769f656054c8dedbebaad4620a7d46975a2be548ed0bf567bb52dc97269f34b

Observation 0a304a85-e77f-4e5f-b010-f4b0c4fe4440 · outbound

This paper cites 2024 , eprint=.

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models 2024 , eprint=

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T09:09:15.211792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-15T09:09:14.884516Z digest=sha256:f7fccae5485a4c06fe37d967a548ea1014f261743f6c5f468c9b62c1a4997b4f

Observation 5af3e1bf-e4e6-44db-a025-42db2db54a46 · outbound

This paper cites 2024 , eprint=.

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models 2024 , eprint=

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T09:09:15.213928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-15T09:09:14.884516Z digest=sha256:51d096795db2814992450e89e724b1f54b6ba582b6444c1bae1c4fccfc6d22a9

Observation 43125d1c-8e83-489a-95e5-c2b4f1dd63b5 · outbound

This paper cites 2024 , eprint=.

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models 2024 , eprint=

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T09:09:15.215972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-15T09:09:14.884516Z digest=sha256:0367adad76364d5666d1004afe7d22c7ab75ac91ebb9112227e30669dc587265

Observation 6d6e929d-b0dd-47b2-98d3-d2aad9758169 · outbound

This paper cites 2023 , eprint=.

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models 2023 , eprint=

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T09:09:15.094323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-15T09:09:14.884516Z digest=sha256:dd0ff09f7a4c148e4ad43830775b687a22adaac5c1aea175efd555ac567df0f8

Observation 249eeee5-f771-4c02-9941-a345786ecd45 · outbound

This paper cites 2024 , eprint=.

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models 2024 , eprint=

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T09:09:15.096947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-15T09:09:14.884516Z digest=sha256:83cd0a0674751505fc7d05fe6be1b8416767d1e4e2b6dbcac369b920919d8077

Observation 93c1207e-74cc-44d3-a4ae-0b2981a7c2a5 · outbound

This paper cites 2024 , eprint=.

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models 2024 , eprint=

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T09:09:15.099259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-15T09:09:14.884516Z digest=sha256:3d5a83354546937eecdddd19b9a48f7b5365adb8dc157e61ac50c78c4b9df808

Observation f3cf1fc6-0ed7-4a76-877b-80afac2210a7 · outbound

This paper cites 2023 , eprint=.

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models 2023 , eprint=

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T09:09:15.101389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-15T09:09:14.884516Z digest=sha256:67f832749b6a7af773fd4a00c6d3929bc61645ac647ac40a5b7f4d35d6b1776a

Observation 8d83c05e-14f3-4525-ae54-d39510b7ac8d · outbound

This paper cites An Iterative Optimizing Framework for Radiology Report Summarization With ChatGPT , volume=.

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models An Iterative Optimizing Framework for Radiology Report Summarization With ChatGPT , volume=

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-15T09:09:14.981330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-15T09:09:14.884516Z digest=sha256:9ad1109b9c40f321937b9762bae8e5a665e2401f537930e787157d19a79e4479

Observation 9fa867a5-85c0-4392-8a3f-bd22735b11e9 · outbound

This paper cites 2024 , eprint=.

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models 2024 , eprint=

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T09:09:15.103601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-15T09:09:14.884516Z digest=sha256:50c2d8392d7b950396e8823a361de7ec06253002ddb627268ce71af61352e32d

Observation 09c0cc3f-75bd-41ab-aa35-54d484f05674 · outbound

This paper cites 2024 , eprint=.

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models 2024 , eprint=

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T09:09:15.105832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-15T09:09:14.884516Z digest=sha256:bd8ea3ae42c0f4e768b69fda2140e03133c0272dc0656c7411569881faaf8823

Observation 7a97d095-6e60-46b8-8736-271fc0467445 · outbound

This paper cites an unresolved cited work.

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-05-15T09:09:15.108004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-15T09:09:14.884516Z digest=sha256:669b1b26175af8ad23952a6aa657c0b1689dea5306708dcc5d61adccd708cbbb

Observation ad6e1c53-3095-41f6-965e-f75d0ed9d5e1 · outbound

This paper cites 2024 , eprint=.

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models 2024 , eprint=

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T09:09:15.110403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-15T09:09:14.884516Z digest=sha256:d85b10a71ed9d31f905251f65b9e99ebb7cad41b7f944a951a11ee70ddb08fab

Observation 756c2935-d152-4b32-9643-c94dc7c26072 · outbound

This paper cites an unresolved cited work.

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-05-15T09:09:15.112558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-15T09:09:14.884516Z digest=sha256:2b1e31bbf86e9ba66457288a50aecfa7230e993a2b1c1dc5d49f355b1e0f73fe

Observation 02aa4479-120f-4bd4-9a46-c054b88a38a9 · outbound

This paper cites 2017 , publisher=.

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models 2017 , publisher=

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T09:09:15.115053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-15T09:09:14.884516Z digest=sha256:8a78ec7a352599e58e61bec59dedbab162b2808b82ed1e4c807ac047d3d9456c

Observation 4e1d2122-6d4f-44de-b677-8d883edf776c · outbound

This paper cites 2024 , eprint=.

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models 2024 , eprint=

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T09:09:15.147956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-15T09:09:14.884516Z digest=sha256:7caa7ba03d7155881b2b33becf483a317dbeb9b38dbe192deee133b2f0ad456d

Observation 2c09868b-e5cb-401f-8cff-d0ebf2cb2936 · outbound

This paper cites 2024 , eprint=.

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models 2024 , eprint=

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T09:09:15.119504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-15T09:09:14.884516Z digest=sha256:3a61611e707f7849170ed5e11bdc87689d1487f0a9ea116d9c019f0ac5cd0726

Observation 0af20043-a3a2-47ad-833c-293a8b067571 · outbound

This paper cites 2024 , howpublished =.

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models 2024 , howpublished =

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T09:09:15.121912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-15T09:09:14.884516Z digest=sha256:adfe02877f432088360f0e4f04f538600d2758e3323003d4f5221334bfb02cc7

Observation a776d289-69a4-4f03-ba6a-fe9c4860b581 · outbound

This paper cites 2024 , howpublished =.

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models 2024 , howpublished =

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T09:09:15.124416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-15T09:09:14.884516Z digest=sha256:31dad90913959d62b32df4970ad0682d0312924141d20fe87a93e057af7491d7

Observation c9de084d-101d-42eb-bbe8-1a268c138f4f · outbound

This paper cites The Llama 3 Herd of Models.

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models The Llama 3 Herd of Models

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-15T09:09:15.020097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-15T09:09:14.884516Z digest=sha256:17cd68c4c2a0502d541aaccf1b625eeb6489ccd5bf062da5c6d9bd0eb687f668

Observation c6123048-fa25-469c-8661-199bd963979f · outbound

This paper cites Mathstral.

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models Mathstral

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T09:09:15.126660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-15T09:09:14.884516Z digest=sha256:aa62a258ee3c666dd6b2c9bc5acca55a43dad1f34715e76cdc39df95fa8e2bd8

Observation 6c182db5-ea1b-4417-be5b-c356e81147b9 · outbound

This paper cites Claude 3.5.

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models Claude 3.5

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T09:09:15.129028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-15T09:09:14.884516Z digest=sha256:5d95338c8187c9e6398a5fea8b7e507a05b392802096ef8796abf69b1ea4b136

Observation dcdd4251-351b-4034-a861-a2f72e6272b3 · outbound

This paper cites Have LLMs Advanced Enough? A Challenging Problem Solving Benchmark For Large Language Models.

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models Have LLMs Advanced Enough? A Challenging Problem Solving Benchmark For Large Language Models

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T09:09:15.030511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-15T09:09:14.884516Z digest=sha256:545a4bbe9ea1bfad89f11fa84c37c82af02f406dc6a8ad82ec64039062c71d57

Observation 718c13df-d833-4de6-92f7-37674a21ecbe · outbound

This paper cites ProofNet: Autoformalizing and Formally Proving Undergraduate-Level Mathematics.

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models ProofNet: Autoformalizing and Formally Proving Undergraduate-Level Mathematics

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T09:09:15.036018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-15T09:09:14.884516Z digest=sha256:3890417212f149b452c646b2cf7456ba350135f4b398981fac9a414e0cd15c3c

Observation 38fd9800-16e8-4cc8-bc16-d25f807c602d · outbound

This paper cites Llemma: An Open Language Model For Mathematics.

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models Llemma: An Open Language Model For Mathematics

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T08:17:47.048289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-15T09:09:14.884516Z digest=sha256:14505da3834791bd6c2d29b3263218155891ffc385c8ea88550dbe541d173566

Observation 345a8872-941e-40c6-a086-813961ca63e8 · outbound

This paper cites Distribution of residual autocorrelations in autoregressive-integrated moving average time series models.

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models Distribution of residual autocorrelations in autoregressive-integrated moving average time series models

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T09:09:15.131469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-15T09:09:14.884516Z digest=sha256:b571f3bd44bd35b3999e349b360bc85448e149620630ef05d37ebd438976c9d3

Observation 200f367c-3777-429e-bee8-0ac5ee91faeb · outbound

This paper cites Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference.

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 52

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T09:09:15.078017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-15T09:09:14.884516Z digest=sha256:b201e9ed9803584c24a6a3889860ea6d0a47ee8747bf4eb651a7ba15124d0667

Observation 41b902db-e93b-49dc-95fa-b5941b896d78 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models Training Verifiers to Solve Math Word Problems

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-05-15T09:09:15.084609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-15T09:09:14.884516Z digest=sha256:1e7c968d11be7c640b3347513111d104d7428ca7654c732054322cbaa1cead88

Observation 55910282-9409-4c43-9999-233ff335fcb5 · outbound

This paper cites DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence.

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-16T01:06:07.964742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-15T09:09:14.884516Z digest=sha256:e420815f15c7aec4cae9faa3538587c378d636c18565c3577bc9b5fe66e5974b

Observation 9ca48e62-59a7-42c7-a127-2ce6a7806efe · outbound

This paper cites Problem-Solving Strategies.

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models Problem-Solving Strategies

Reference 55

Resolution
verified exact
doi, observed 2026-05-15T09:09:14.975632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-15T09:09:14.884516Z digest=sha256:818d50585aa9a865fc7279ce2d54b24fde1f5085558fe4f5f8c8ac3aa98bd2aa

Observation ed9b2a01-ae68-458c-8954-79d7f91a31ce · outbound

This paper cites MathOdyssey: Benchmarking Mathematical Problem-Solving Skills in Large Language Models Using Odyssey Math Data.

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models MathOdyssey: Benchmarking Mathematical Problem-Solving Skills in Large Language Models Using Odyssey Math Data

Reference 56

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T09:09:14.987082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-15T09:09:14.884516Z digest=sha256:79112b12cc8796e98bf0ec9a1d265cb199db3b070a1527d305a85364d6977d81

Observation 174bd1a2-1c24-4758-98fe-7cd4aa632bd7 · outbound

This paper cites LLM Critics Help Catch Bugs in Mathematics: Towards a Better Mathematical Verifier with Natural Language Feedback.

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models LLM Critics Help Catch Bugs in Mathematics: Towards a Better Mathematical Verifier with Natural Language Feedback

Reference 57

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T09:09:14.991946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-15T09:09:14.884516Z digest=sha256:83d26930c6a7a679931a887a1fa4ed10ce23b36972f88492af57d46884997cf4

Observation 36ffdc63-201c-4d1f-88d0-2d8df8affe94 · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-05-15T09:09:14.996702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-15T09:09:14.884516Z digest=sha256:75e95311fa90ef384f2f0b53fe1c8e52a39b0da4345ce7f992970eb59390547e

Observation 089a28c4-3dbf-4bf2-ba7f-979b43201aa9 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models Measuring Mathematical Problem Solving With the MATH Dataset

Reference 59

Resolution
verified exact
local_arxiv, observed 2026-05-15T09:09:15.001647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-15T09:09:14.884516Z digest=sha256:138297d16147518446824baa5f3bb69e2e6c1c4bcefb670d0765df483f68cccd

Observation e1fe4b74-821d-4190-b2de-d826281acc9a · outbound

This paper cites OlympicArena: Benchmarking Multi-discipline Cognitive Reasoning for Superintelligent AI.

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models OlympicArena: Benchmarking Multi-discipline Cognitive Reasoning for Superintelligent AI

Reference 60

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T09:09:15.006579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-15T09:09:14.884516Z digest=sha256:725e42c099fdb8da6fb2e1e977dfb2513e5dfd0df29f124e97cb8dcc9b98d542

Observation 02d4b553-6a98-4c4a-9895-575d3465edef · outbound

This paper cites Qwen2.5-Coder Technical Report.

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models Qwen2.5-Coder Technical Report

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-05-15T09:09:15.011138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-15T09:09:14.884516Z digest=sha256:e3e6495f8f86c002505b45a029ee93372aaa9111d661fce516b684e9cbfa4c6f

Observation b936ff09-3887-48b4-9ff5-c3e429fdd334 · outbound

This paper cites Solving Quantitative Reasoning Problems with Language Models.

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models Solving Quantitative Reasoning Problems with Language Models

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-05-15T09:09:15.015382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-15T09:09:14.884516Z digest=sha256:c0130b08c272f8444aeca48bc8ffa17e6c23ef31fc9970b6cb18616fb0329081

Observation f606ce41-2851-4a0d-a558-cd6fd85cf1f8 · outbound

This paper cites Numinamath.

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models Numinamath

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T09:09:15.133576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-15T09:09:14.884516Z digest=sha256:de9fc64162122796592609089dd35dead7f42e6c35979c7ca9def00e4e65dd6f

Observation dffef66e-f120-4612-9347-99c579b7279a · outbound

This paper cites CHAMP: A Competition-level Dataset for Fine-Grained Analyses of LLMs' Mathematical Reasoning Capabilities.

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models CHAMP: A Competition-level Dataset for Fine-Grained Analyses of LLMs' Mathematical Reasoning Capabilities

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-15T09:09:15.024469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-15T09:09:14.884516Z digest=sha256:aa2c16eda713504334da32c1128a0db8c42eed60863b817c2bd06c1a142b6297

Observation e9627305-6f7f-45f6-9fe3-6b5438bff278 · outbound

This paper cites Gpt-4 technical report.

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models Gpt-4 technical report

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T09:09:15.136344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-15T09:09:14.884516Z digest=sha256:5d261056efbcccb08340e46cbe85a6663372d531c46f4074f89baacff7bb4c06

Observation 67f2e29f-5fbb-41e7-a92b-897233ffd297 · outbound

This paper cites Learning to reason with llms.

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models Learning to reason with llms

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T09:09:15.138543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-15T09:09:14.884516Z digest=sha256:f139de89e753e054b15140db2cf8cd29b7e4fe3e8e6ce0490b3c04e4dd21b2a1

Observation cc2dce73-a10a-48bd-ade3-79fe29458d89 · outbound

This paper cites Code Llama: Open Foundation Models for Code.

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models Code Llama: Open Foundation Models for Code

Reference 67

Resolution
verified exact
local_arxiv, observed 2026-05-15T09:09:15.045633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-15T09:09:14.884516Z digest=sha256:522ee8f96f48d43cd4d53c8fb0943f64f7668964478ac17402924eb0b0642130

Observation 372f2faa-7aed-4b5f-8ab0-6b23354cbcab · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 68

Resolution
verified exact
local_arxiv, observed 2026-05-15T09:09:15.049967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-15T09:09:14.884516Z digest=sha256:88be68552fd8ef6d2f276f2b56bb2320a6d4c504f397db031fea7d651eddb3e8

Observation 8a75b2ab-cca7-4914-ac35-87d091de9a1e · outbound

This paper cites Solving olympiad geometry without human demonstrations.

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models Solving olympiad geometry without human demonstrations

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T09:09:15.141280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-15T09:09:14.884516Z digest=sha256:31a3d3b7adf1fe8b8a472893712b9ebcb9c795e1a8412334a6d172355d31abed

Observation 817fd360-ae84-4c74-a0cf-b7d6cf23ac86 · outbound

This paper cites Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations.

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations

Reference 70

Resolution
verified exact
local_arxiv, observed 2026-05-15T09:09:15.058502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-15T09:09:14.884516Z digest=sha256:789fc49bdd7f9c0c498d4575032635845dcf32f48edaa7955264a837933101ea

Observation 5500fd44-6682-4a2d-8452-5aca098e8a37 · outbound

This paper cites Preserving in-context learning ability in large language model fine-tuning.

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models Preserving in-context learning ability in large language model fine-tuning

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T09:09:15.143700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-15T09:09:14.884516Z digest=sha256:49357706452606b076f69805434a2553b363dec5aeee0bca17bf4ee03ee86be2

Observation 35845f36-317b-421a-94c2-f2807d182c79 · outbound

This paper cites Benchmarking Benchmark Leakage in Large Language Models.

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models Benchmarking Benchmark Leakage in Large Language Models

Reference 72

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T09:09:15.062242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-15T09:09:14.884516Z digest=sha256:6d2bfade3f50a5c12e0f38ec8e10911ed5be6a5c4537d4a2e77c4a3c296781a4

Observation 136e4829-a1bb-4adb-b945-fd2378d31e71 · outbound

This paper cites Qwen2 Technical Report.

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models Qwen2 Technical Report

Reference 73

Resolution
verified exact
local_arxiv, observed 2026-05-15T09:09:15.065363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-15T09:09:14.884516Z digest=sha256:f54d715bc07e17a284e02eae722ab27843e18a3771cf778e63c20159f66b8e21

Observation 4c88fc30-871c-45d3-a12d-950f9e21c535 · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 74

Resolution
verified exact
local_arxiv, observed 2026-05-15T09:09:15.068275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-15T09:09:14.884516Z digest=sha256:b84e397919a146b5374083cb6a69208ca37f12edf814540030592715352b8f0a

Observation f8baea53-1534-4f46-8b4d-0b2b7d06328e · outbound

This paper cites Can Large Language Models Always Solve Easy Problems if They Can Solve Harder Ones?.

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models Can Large Language Models Always Solve Easy Problems if They Can Solve Harder Ones?

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-15T09:09:15.071639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-15T09:09:14.884516Z digest=sha256:9567215fab2e555e83db7f0bebf4a3adb7a8a992a51beb8fbd7f9db2595d2b99

Observation ee60e291-cb29-4f52-9319-9e737ce448b5 · outbound

This paper cites InternLM-Math: Open Math Large Language Models Toward Verifiable Reasoning.

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models InternLM-Math: Open Math Large Language Models Toward Verifiable Reasoning

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-05-15T09:09:15.075148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-15T09:09:14.884516Z digest=sha256:438a01ec0c0a2445cceb4ecfc2cd55d63c96f574f837916dac890a89ed1c8dad

Observation 3e027ceb-a0f1-4b9e-8ad4-8990c7b47dd4 · outbound

This paper cites The art and craft of problem solving.

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models The art and craft of problem solving

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T09:09:15.145821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-15T09:09:14.884516Z digest=sha256:d9c83dbeeb8f81d5ae465f534b2b88a3596468f80b6e94a7db1397c06f11a907

Observation 25abb680-ff6a-40a1-8d0d-78b693ab81ef · outbound

This paper cites MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?.

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?

Reference 78

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:29:30.364997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-15T09:09:14.884516Z digest=sha256:2b0b009f5f1340736f3a1c7d82ffbfe827f22a0317e34129ca7c441ed1e0823c

Observation 218aff1b-80eb-48ba-8f6f-138e2144ab7d · outbound

This paper cites MiniF2F: a cross-system benchmark for formal Olympiad-level mathematics.

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models MiniF2F: a cross-system benchmark for formal Olympiad-level mathematics

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:03:04.481585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-15T09:09:14.884516Z digest=sha256:85cc5d377037f425064e140e0c5e00ff536408d0d267ec38016e61f53320ea43

Pith citing papers

Observation 15afcf64-1beb-4bcd-9a12-3cb633a8a7d3 · inbound

Confidence v.s. Critique: A Decomposition of Self-Correction Capability for LLMs cites this paper.

Confidence v.s. Critique: A Decomposition of Self-Correction Capability for LLMs Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T00:22:29.920179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:22:29.920179Z digest=sha256:754aaa4b42e373650aa848b2207725f99965e87d08611228519543a9de509e44

Observation 30800675-fe33-4030-8767-e9e2aca13a71 · inbound

CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings cites this paper.

CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 1978

Resolution
unresolved
no resolver link, observed 2026-08-10T22:37:58.412771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:37:58.412771Z digest=sha256:47071e1ed32d3afd7a02de57713d51ae91cb0e958f3980e5647eb5eb849f9955

Observation 2bb0e081-ff39-4c5d-9db7-e4aa955c24ef · inbound

End-to-End Bangla AI for Solving Math Olympiad Problem Benchmark: Leveraging Large Language Model Using Integrated Approach cites this paper.

End-to-End Bangla AI for Solving Math Olympiad Problem Benchmark: Leveraging Large Language Model Using Integrated Approach Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T21:36:40.255920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:36:40.255920Z digest=sha256:50856dd6f9bc446f7e17ab45d24a3be2050eca602fea4b885e349458af1ca499

Observation e6f7cb30-a856-48c3-b2c4-36aabad2993d · inbound

T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling cites this paper.

T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T18:05:42.995015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:05:42.995015Z digest=sha256:b98bd671ea2b91ff8c9e550c0992736f90b4a22acd58d41d49da31183c8562b8

Observation 6e914fe3-e237-4b2a-9699-84239dbf0294 · inbound

Humanity's Last Exam cites this paper.

Humanity's Last Exam Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-15T09:09:15.216863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:390456ed5241a5e1f15763563887c70f0c48f9143737b21a2f78248253ffc70a

Observation 12a7c8d0-1ed8-4fc8-82b8-03ae4d6ffb86 · inbound

UGPhysics: A Comprehensive Benchmark for Undergraduate Physics Reasoning with Large Language Models cites this paper.

UGPhysics: A Comprehensive Benchmark for Undergraduate Physics Reasoning with Large Language Models Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-09T19:27:46.645948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T19:27:46.645948Z digest=sha256:5e5989de7dcc10c6ff604268955feb6dc920c30a8e298ad3118fd565461b669f

Observation 906baf4e-c8d2-41ed-9100-66a0d22d47f7 · inbound

L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning cites this paper.

L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-18T00:19:22.206556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T00:19:22.140009Z digest=sha256:0099cc444494e0ce847c600e003a850a9fcde755ffad55cb5bd28f1b4c475d7c

Observation 1a6b304e-0c60-4719-85e4-d978156008d5 · inbound

LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL cites this paper.

LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-16T15:15:46.382972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T15:15:46.255296Z digest=sha256:f230e6cee97b4c413c167a7a0c8ba5d51686027fa52aa09d211beb63aa5e529e

Observation b94038d3-60fe-489f-aee2-ede6b1573753 · inbound

MathFlow: Enhancing the Perceptual Flow of MLLMs for Visual Mathematical Problems cites this paper.

MathFlow: Enhancing the Perceptual Flow of MLLMs for Visual Mathematical Problems Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-22T22:57:13.232377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-22T22:55:34.238427Z digest=sha256:79e27e65807d3e52f4e7ca183cb03ceb9be1fbca9a79976332dfd14eaa90b305

Observation 91278f42-2e3b-46c9-a8e9-461d97f33d26 · inbound

Challenging the Boundaries of Reasoning: An Olympiad-Level Math Benchmark for Large Language Models cites this paper.

Challenging the Boundaries of Reasoning: An Olympiad-Level Math Benchmark for Large Language Models Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-22T22:15:11.196939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-22T22:14:53.522206Z digest=sha256:ab01206ac76741a192c20d6b5e10ca7c111d702cd894b4467828dd4a4c13d882

Observation 85cff00c-7d90-46d5-a7ff-42656da1cbae · inbound

DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning cites this paper.

DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T10:31:04.794480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T10:31:04.728005Z digest=sha256:340aa8af163dc87b18b238a10d4b3c976c4e9faa540099c05afa893453af4053

Observation 2e341750-2138-46c8-9ec1-da57a5347267 · inbound

Reinforcement Learning for Reasoning in Large Language Models with One Training Example cites this paper.

Reinforcement Learning for Reasoning in Large Language Models with One Training Example Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-05-15T19:51:05.052431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T19:51:04.779597Z digest=sha256:878c0803f336d6ca39021e7eeb899ce1a39ba56b9d1d96585ac67b994a984a9c

Observation 3a56fac1-4f09-493f-ad1f-d88b1ad02e51 · inbound

The Hallucination Tax of Reinforcement Finetuning cites this paper.

The Hallucination Tax of Reinforcement Finetuning Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:30.864841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:30.864841Z digest=sha256:c097fd5112d6b27d2feece29ba07f4ce7520bc60ab7b198b9a4b6254845b433f

Observation f67a7076-9d68-40a0-920b-580c6667ec10 · inbound

SCOPE: Compress Mathematical Reasoning Steps for Efficient Automated Process Annotation cites this paper.

SCOPE: Compress Mathematical Reasoning Steps for Efficient Automated Process Annotation Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:39:16.630641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:39:16.630641Z digest=sha256:3d89df815fd38f5dd05fd9d7c8efab097bbdc9c97e4cb06ea07234f34ad4bfda

Observation ae6bea58-246b-4a2f-afdb-9148b5a5af40 · inbound

Incentivizing Dual Process Thinking for Efficient Large Language Model Reasoning cites this paper.

Incentivizing Dual Process Thinking for Efficient Large Language Model Reasoning Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:58.842174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:58.842174Z digest=sha256:0b924fb48ee29d84ffbfdbd5bed3be621aca5f4b7b0381fced9f8358cf7d9510

Observation 887e1314-3873-42ed-9526-052b06796094 · inbound

AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning cites this paper.

AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:35.322184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:35.322184Z digest=sha256:28f8fb8930a2976507e51d2b9c9e86fd26c7caae193039b9145b75ddfcf58cac

Observation 31ca3bd3-192a-4723-a0f4-dd4082dc77b7 · inbound

RBench-V: A Primary Assessment for Visual Reasoning Models with Multi-modal Outputs cites this paper.

RBench-V: A Primary Assessment for Visual Reasoning Models with Multi-modal Outputs Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:58:42.014658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:58:42.014658Z digest=sha256:18fe0297764a1aa1d242443c2ca4434484b6c799ea75f4266202a2bcf8385470

Observation 5ee06ac6-e9c8-4669-a56a-d1e83aa8f32b · inbound

Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards cites this paper.

Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:01.324509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:01.324509Z digest=sha256:754cda2b01a6b5fd7fdf53db54535935d822a7e18d1bc71ede41aca237698eed

Observation 2c3b15ba-1857-43f8-b4b8-5ab6c984caad · inbound

Formally Solving Answer-Construction Problems in Lean cites this paper.

Formally Solving Answer-Construction Problems in Lean Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:07.262018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:07.262018Z digest=sha256:f25da1202ec823069541cf8f6655e283b866743bb9419111608297c46cddaba6

Observation cb9f2095-3250-48b0-8e20-4e991cea372b · inbound

CoTGuard: Using Chain-of-Thought Triggering for Copyright Protection in Multi-Agent LLM Systems cites this paper.

CoTGuard: Using Chain-of-Thought Triggering for Copyright Protection in Multi-Agent LLM Systems Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:17.780680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:17.780680Z digest=sha256:7cb2688ccafc1f5aa0126d98f79f9d31db91eda98890c112bd5a0bae801819a9

Observation a8570983-096a-48d2-96c9-cb11b5ddd572 · inbound

Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning cites this paper.

Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:41:29.732447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:41:29.732447Z digest=sha256:26be2274ffb3a1139c90723d814b38e242566c77aef0321be86153fe7c4e6c57

Observation 69415067-5e5a-4773-98af-3df7ece59e28 · inbound

Decomposing Elements of Problem Solving: What "Math" Does RL Teach? cites this paper.

Decomposing Elements of Problem Solving: What "Math" Does RL Teach? Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:25.427253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:05:25.427253Z digest=sha256:933ac8585b811319d40ca71863d837347bcaf6f2b4c07e569b9bf655a72dad76

Observation 07659698-598a-4887-90eb-6701309475c2 · inbound

MathArena: Evaluating LLMs on Uncontaminated Math Competitions cites this paper.

MathArena: Evaluating LLMs on Uncontaminated Math Competitions Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-15T09:09:15.216863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T00:10:14.812539Z digest=sha256:44c0df3b2903c5b7804eb3ab234ba022170cbf637bfcb1da1254c61e540c6dba

Observation 35e7b365-250a-4634-b96e-a7e534e7b8db · inbound

Revisiting Test-Time Scaling: A Survey and a Diversity-Aware Method for Efficient Reasoning cites this paper.

Revisiting Test-Time Scaling: A Survey and a Diversity-Aware Method for Efficient Reasoning Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T10:42:38.944246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:42:38.944246Z digest=sha256:ff463550b100c212722e2f87f5d5f4dfc4c30b01bf53188a6a27bba7e2ae1e93

Observation e711829c-0ac3-47de-a042-751ba8441949 · inbound

Reinforcement Pre-Training cites this paper.

Reinforcement Pre-Training Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:22.086147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:26:22.086147Z digest=sha256:0b5b81b99624aea1ebe5131698d53c3752100fa774362bc97ba29cbdf36505de

Observation d7d0511f-bc7d-4b3f-9910-c310f4ad58b5 · inbound

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search cites this paper.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:20.862537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:20.862537Z digest=sha256:9dbbb82895376fd68d8a29682bdeb6cd0769d1f9364afbee1672b92b1a5ca188

Observation 7c5f84c5-589f-4ffe-9226-2f2ca8fa44f7 · inbound

SciDA: Scientific Dynamic Assessor of LLMs cites this paper.

SciDA: Scientific Dynamic Assessor of LLMs Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T00:40:17.637897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:40:17.637897Z digest=sha256:af5f773b159be38c6ea6dfdf76202dfa535b847b35001bdcf67318f657a2b36e

Observation 62e2253e-bac5-43a5-9fd5-69693d221042 · inbound

CriticLean: Critic-Guided Reinforcement Learning for Mathematical Formalization cites this paper.

CriticLean: Critic-Guided Reinforcement Learning for Mathematical Formalization Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T19:14:14.957669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:14:14.957669Z digest=sha256:2999fe5641da9a7e1435fa963f5d6bed8c6dc65eec3022de2362145fdaaae57b

Observation fa384924-2e54-4492-a8df-de4af4faa6ab · inbound

One Token to Fool LLM-as-a-Judge cites this paper.

One Token to Fool LLM-as-a-Judge Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T18:15:37.080727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:15:37.080727Z digest=sha256:349663283e6c9cd40ec108f3b34fba9f692f49a894739dd855d7a3a660bc2c8c

Observation 32029885-2627-44ca-a02e-fa8a0510044d · inbound

Towards Concise and Adaptive Thinking in Large Reasoning Models: A Survey cites this paper.

Towards Concise and Adaptive Thinking in Large Reasoning Models: A Survey Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T17:53:47.832078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:53:47.832078Z digest=sha256:106569b244aeeb97c46336fb8c1c7fb5a173c7cbf5a46cd7dddc7d5a12104940

Observation 06093aef-6291-4c7e-aa5b-3a6027b33ad1 · inbound

REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once cites this paper.

REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:31.417001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:35:31.417001Z digest=sha256:ea02687778a1f4020e18a3914fde6bafb3b0217270301406c3f24e0ae4106696

Observation 0c4b594e-f0a2-43cc-8651-0f6f4b6a0a1d · inbound

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) cites this paper.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T16:44:49.647176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:44:49.647176Z digest=sha256:f4e8c1d111effb8f50c44457e5f4a1d5ad8683d50ff4f4163e2aee956fcbb110

Observation 400cd132-617b-436b-a6b7-43a1cc4c6d71 · inbound

Proof2Hybrid: Automatic Mathematical Benchmark Synthesis for Proof-Centric Problems cites this paper.

Proof2Hybrid: Automatic Mathematical Benchmark Synthesis for Proof-Centric Problems Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T05:11:53.696502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:11:53.696502Z digest=sha256:c504c0deb26a89993e6f8c1e071258d0c7d355e41f1b8f66fd02f03a6cdfcf06

Observation 46efc729-4bcc-469c-9444-c125fa14dbf2 · inbound

Grove MoE: Towards Efficient and Superior MoE LLMs with Adjugate Experts cites this paper.

Grove MoE: Towards Efficient and Superior MoE LLMs with Adjugate Experts Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T21:55:07.854608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:55:07.854608Z digest=sha256:58e3918f9235738e5457aa02a4868c72acb02b2e6554feec2d579d4d0061266f

Observation b3ec06ee-2212-451f-a16d-175cb0800d5e · inbound

Throttling Web Agents Using Reasoning Gates cites this paper.

Throttling Web Agents Using Reasoning Gates Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T12:28:03.943989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:28:03.943989Z digest=sha256:0f0bc8b735b7165fdd2a0431667b9d76ec92fd92b9c8eb6948b600a0abc3f7bb

Observation 98543ed1-8fd0-45c8-9acf-285ac669280c · inbound

Another Turn, Better Output? A Turn-Wise Analysis of Iterative LLM Prompting cites this paper.

Another Turn, Better Output? A Turn-Wise Analysis of Iterative LLM Prompting Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T23:12:10.137275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:12:10.137275Z digest=sha256:c4689d7b97b2b05745e8a7de800d395065cb417783b2025af03a27b4c8bcdab1

Observation 4aca38e9-be89-4ee7-a789-2e9140c49eef · inbound

RIMO: An Easy-to-Evaluate, Hard-to-Solve Olympiad Benchmark for Advanced Mathematical Reasoning cites this paper.

RIMO: An Easy-to-Evaluate, Hard-to-Solve Olympiad Benchmark for Advanced Mathematical Reasoning Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T21:54:48.183408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:54:48.183408Z digest=sha256:a60e495c274d3c93a21081a7d919445ded0c8a8b9cc3994729f365ab508b2187

Observation 4c49559d-194a-47fc-bfaf-fecdf23671ac · inbound

Student-Centered Distillation Narrows the Agentic Gap Between Small and Large LLMs cites this paper.

Student-Centered Distillation Narrows the Agentic Gap Between Small and Large LLMs Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T17:57:52.348384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:57:52.348384Z digest=sha256:975189dadcef679e112b40e6919fafea1ffef18440c27a0e5fda10a4b69598f9

Observation dfd12c5a-cb70-4bee-8fb7-e4975cf227cf · inbound

EngiBench: A Benchmark for Evaluating Large Language Models on Engineering Problem Solving cites this paper.

EngiBench: A Benchmark for Evaluating Large Language Models on Engineering Problem Solving Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-18T15:11:32.661126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T15:08:03.793304Z digest=sha256:076738a62c3fac49a2324ba2ad5f2bf7e2b0ddc67126ff2aebb8c6c51520c8f7

Observation c4988a27-d134-48c3-8ddb-236d8af9ec9e · inbound

StatEval: A Comprehensive Benchmark for Large Language Models in Statistics cites this paper.

StatEval: A Comprehensive Benchmark for Large Language Models in Statistics Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T10:36:23.280848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:36:23.280848Z digest=sha256:6483b6915e5ed657932ec845f40dbbbdc9a837986e15ac8f41c01d87630275cc

Observation 06dfca44-1c00-4d8c-8ece-9b7bb9ab5c00 · inbound

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation cites this paper.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.581084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.581084Z digest=sha256:533d16444f6b24ba271669ccc945d7cbd0ce3535de476f99376bdd5f5bcb1c75

Observation 17d1b9ca-423f-438e-a586-6552ca628b97 · inbound

LLaDA2.0: Scaling Up Diffusion Language Models to 100B cites this paper.

LLaDA2.0: Scaling Up Diffusion Language Models to 100B Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-15T09:09:15.216863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-14T18:53:20.911374Z digest=sha256:8f316d5640e808d0dbace9a1b9529fc0f931ed096f6eeaea1d3dd9b2aa729ef5

Observation 7b1cd48c-cbdb-468f-9303-5e6e6d6268d3 · inbound

The Geometric Reasoner: Manifold-Informed Latent Foresight Search for Long-Context Reasoning cites this paper.

The Geometric Reasoner: Manifold-Informed Latent Foresight Search for Long-Context Reasoning Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-16T10:42:45.214512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T10:42:33.579501Z digest=sha256:7b2ac244bf519a85896c213b4d442494f1a657985177bfedca76663c9f20abe0

Observation e0d579c9-2615-42c2-8439-a081d552c22b · inbound

ECHO-2: A Large-Scale Distributed Rollout Framework for Cost-Efficient Reinforcement Learning cites this paper.

ECHO-2: A Large-Scale Distributed Rollout Framework for Cost-Efficient Reinforcement Learning Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T05:32:42.552709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:32:42.552709Z digest=sha256:d0fe209415a47a98fc3ecf1ccff306a5abfb834ddf5ff8bb9eae59569ae880f6

Observation 8dd39d68-61f9-42c5-a67e-6b7a1480586d · inbound

LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics cites this paper.

LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T20:05:56.217637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:05:56.217637Z digest=sha256:d32d3673866cbbe3b3aef276c2218a11bda164870f6db7aa4e606d2c88249c37

Observation ab5b8ce9-1ab1-4587-93b8-4af6cc0e48a2 · inbound

OpenHospital: A Thing-in-itself Arena for Evolving and Benchmarking LLM-based Collective Intelligence cites this paper.

OpenHospital: A Thing-in-itself Arena for Evolving and Benchmarking LLM-based Collective Intelligence Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-14T21:00:05.911327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T21:00:05.911327Z digest=sha256:cad77a7e3745a536082c733fc2156f306fe6090a8c35c79fb693e983904a7243

Observation db0b71ae-c859-4823-9861-72ff92aa634a · inbound

Open, Reliable, and Collective: A Community-Driven Framework for Tool-Using AI Agents cites this paper.

Open, Reliable, and Collective: A Community-Driven Framework for Tool-Using AI Agents Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-14T19:59:59.030585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T19:59:59.030585Z digest=sha256:c632b4982fa55fb9f6c8a1c648588661526c851caada7d4a435723734a5bef79

Observation a524c0e0-bb4c-4317-bf1a-13eaff005a87 · inbound

Riemann-Bench: A Benchmark for Moonshot Mathematics cites this paper.

Riemann-Bench: A Benchmark for Moonshot Mathematics Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-15T09:09:15.216863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T18:02:42.607682Z digest=sha256:ad2eb32b2f715b7ecfefe5ef410651b9725e902f95de59f09a17412962067122

Observation 282951e6-a92a-4ea0-b2db-4fbf7f691e22 · inbound

HEALing Entropy Collapse: Enhancing Exploration in Few-Shot RLVR via Hybrid-Domain Entropy Dynamics Alignment cites this paper.

HEALing Entropy Collapse: Enhancing Exploration in Few-Shot RLVR via Hybrid-Domain Entropy Dynamics Alignment Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 98

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T09:09:15.216863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-10T05:23:08.478393Z digest=sha256:640281e4c243657991b91aa7dbcf40625fff5477a9de3dad1ccf495a267d9755

Observation 93bd2711-d30d-48f5-afb6-351a5c074111 · inbound

MathNet: a Global Multimodal Benchmark for Mathematical Reasoning and Retrieval cites this paper.

MathNet: a Global Multimodal Benchmark for Mathematical Reasoning and Retrieval Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T09:09:15.216863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-10T04:09:46.616019Z digest=sha256:d4b28d176c8536e114483099365ccd8520b9618015cfa7b6096aeeb23321a697

Observation 838ade7f-4b5b-4def-8cd6-70bb4f0a10e2 · inbound

OLLM: Options-based Large Language Models cites this paper.

OLLM: Options-based Large Language Models Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-15T09:09:15.216863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T02:04:12.174834Z digest=sha256:963cf632f5449c5c613a00f378e0c44fc4aa72df81ccf742884a0d4f5db9c3b0

Observation f1c004be-7c19-4585-9da1-f9a2e6d475f6 · inbound

Dual-Cluster Memory Agent: Resolving Multi-Paradigm Ambiguity in Optimization Problem Solving cites this paper.

Dual-Cluster Memory Agent: Resolving Multi-Paradigm Ambiguity in Optimization Problem Solving Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 54

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T09:09:15.216863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-10T00:37:55.147350Z digest=sha256:2e3a78e9380161765752752a030aadc14e7920e4f54e08814bb580cdf89f5ba7

Observation 4cd45cc3-8a0d-497b-a96f-578296a69bb5 · inbound

OptiVerse: A Comprehensive Benchmark towards Optimization Problem Solving cites this paper.

OptiVerse: A Comprehensive Benchmark towards Optimization Problem Solving Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 71

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T09:09:15.216863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-09T22:04:19.654714Z digest=sha256:2e47bfadce9e23552c303859f036fc1c991a1cbbcaecdb80202f02e972d4a5fa

Observation ca08261e-a418-411e-ba16-d09ca7c1cfa7 · inbound

Rethinking Math Reasoning Evaluation: A Robust LLM-as-a-Judge Framework Beyond Symbolic Rigidity cites this paper.

Rethinking Math Reasoning Evaluation: A Robust LLM-as-a-Judge Framework Beyond Symbolic Rigidity Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-15T09:09:15.216863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-08T11:45:54.181019Z digest=sha256:9f4f18c25139292fa6f6b177759a947d65a47baeafb529f38eeab28e651e680b

Observation 282031cd-6c11-4d82-b87c-0e9cc8b09057 · inbound

Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity cites this paper.

Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T09:09:15.216863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-09T19:56:39.465133Z digest=sha256:015d0cd25fb4d05a4aacaf3cbdfd130add50a4ed0dd74a4c55a5bcd117cd2e1e

Observation df0b4dbe-0944-4a03-9b9b-33ac3c1ab829 · inbound

Verifiable Counterfactual Supervision for Process Reward Models cites this paper.

Verifiable Counterfactual Supervision for Process Reward Models Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-15T09:09:15.216863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T19:24:18.022593Z digest=sha256:bb00e0b915621ad0d744dec9ce456610278190aa308f5bb456cb176e30d5d02c

Observation 9bc3e4fd-85be-4e4a-a740-1c965bd41702 · inbound

You Snooze, You Lose: Automatic Safety Alignment Restoration through Neural Weight Translation cites this paper.

You Snooze, You Lose: Automatic Safety Alignment Restoration through Neural Weight Translation Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-15T09:09:15.216863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T17:02:20.836208Z digest=sha256:6ccd8e6dcb4bb972528df09c31163d285df9ebd4d9c75d6ef29286d2ed4f1b6f

Observation cb507dec-0fa6-4004-94fb-048b23aa4fc6 · inbound

Policy-Guided Stepwise Model Routing for Cost-Effective Reasoning cites this paper.

Policy-Guided Stepwise Model Routing for Cost-Effective Reasoning Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-15T09:09:15.216863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T10:29:23.538806Z digest=sha256:57daea26982731a3b21f537a23a1249126bb5ddc83e0a6f6c69c5b7015b7e376

Observation 53ef6107-7bd0-4e64-a5f9-2b8924308ecf · inbound

Iterative Critique-and-Routing Controller for Multi-Agent Systems with Heterogeneous LLMs cites this paper.

Iterative Critique-and-Routing Controller for Multi-Agent Systems with Heterogeneous LLMs Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-15T09:09:15.216863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:10:03.459912Z digest=sha256:f2b2c37b118560cf26d5b8e0d5ed01ee5529366f1fb9d2398602411e199c773e

Observation fdf3d0ef-00a8-4fef-aa75-ada9cc889f2b · inbound

Beyond Accuracy: Evaluating Strategy Diversity in LLM Mathematical Reasoning cites this paper.

Beyond Accuracy: Evaluating Strategy Diversity in LLM Mathematical Reasoning Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-15T09:09:15.216863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:30:29.268117Z digest=sha256:863e3efc3d0c9421a83f3466c90c6c9a621326ca02bc77a0618384bdf66a174b

Observation af13ff09-c550-4d0e-98ec-439e0bb4be80 · inbound

TIDE-Bench: Task-Aware and Diagnostic Evaluation of Tool-Integrated Reasoning cites this paper.

TIDE-Bench: Task-Aware and Diagnostic Evaluation of Tool-Integrated Reasoning Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T09:09:15.216863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-12T02:47:31.830410Z digest=sha256:df4a7a221777ffb087bdd8dd4698b0c84a758a913cb72420ba913a5bc5106a59

Observation cc48f11c-e718-463d-919b-7b742f5c0ecd · inbound

Nice Fold or Hero Call: Learning Budget-Efficient Thinking for Adaptive Reasoning cites this paper.

Nice Fold or Hero Call: Learning Budget-Efficient Thinking for Adaptive Reasoning Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-15T09:09:15.216863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T01:31:21.377700Z digest=sha256:7f96c170e148093b23d6593eeefb7a7247a6accf3a7d9de60fd0bd41790872a8

Observation e12b708f-6244-40cb-8a63-c6f760dcef65 · inbound

Stress-Testing the Reasoning Competence of LLMs With Proofs Under Minimal Formalism cites this paper.

Stress-Testing the Reasoning Competence of LLMs With Proofs Under Minimal Formalism Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 103

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T09:09:15.216863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-14T21:03:13.407192Z digest=sha256:992683b154ec3bacfdf50b53c5c914c5044753b0046ea60b25ef699faf8dfe38

Observation 362cf041-0fe3-45e8-93cd-7629f955998b · inbound

RMA: an Agentic System for Research-Level Mathematical Problems cites this paper.

RMA: an Agentic System for Research-Level Mathematical Problems Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-25T06:10:24.176711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-25T06:06:47.226726Z digest=sha256:ff1f646b4b59892fd4fae61379946bc0ce14ffcd2d73449c251d679a20fbb518

Observation 4b75cfb3-a1ca-458a-bf81-03c02e72d9f7 · inbound

SkillOpt: Executive Strategy for Self-Evolving Agent Skills cites this paper.

SkillOpt: Executive Strategy for Self-Evolving Agent Skills Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-25T03:55:21.082658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-25T03:51:49.812710Z digest=sha256:bc741e3cd5329adde0674ed78b63bf0dd426beaa0868b6c4e583e8268dcb67e8

Observation f66cc439-9b5f-4175-9773-67ea13d92506 · inbound

SkillOpt: Executive Strategy for Self-Evolving Agent Skills cites this paper.

SkillOpt: Executive Strategy for Self-Evolving Agent Skills Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-06-30T16:35:12.863669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-30T16:26:23.130418Z digest=sha256:abe7efee750e334f18f35559cb3089d5de8592874da1e390c2d4f039987c5b9e

Observation 90451569-2903-46be-946b-89165f4f4bdc · inbound

RLVR Datasets and Where to Find Them: Tracing Data Lineage for Better Training Data cites this paper.

RLVR Datasets and Where to Find Them: Tracing Data Lineage for Better Training Data Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-06-29T19:43:55.098336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-29T19:34:12.081362Z digest=sha256:486ade7ce7ac9ccd25906b9178db235a06a419c1f517e8eac64b031be7ab7c96

Observation 17200321-6e33-417a-8953-6feca06a231d · inbound

Bridging the Detection-to-Abstention Gap in Reasoning Models under Insufficient Information cites this paper.

Bridging the Detection-to-Abstention Gap in Reasoning Models under Insufficient Information Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-06-29T12:43:25.795466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-29T12:37:16.138816Z digest=sha256:89251650015cbf75cf69adcc8be7e4c8a89a2fdee399fc3ecbfae411c3282145

Observation 377fb678-6ebb-41f5-a322-7c1f5ba8d447 · inbound

Artificial Intelligence for Mathematical Reasoning: An Integrated Survey of Language Models, Neuro-symbolic Systems, and Verified Discovery cites this paper.

Artificial Intelligence for Mathematical Reasoning: An Integrated Survey of Language Models, Neuro-symbolic Systems, and Verified Discovery Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 219

Resolution
verified exact
local_arxiv, observed 2026-07-02T22:47:26.006683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-27T18:39:44.696961Z digest=sha256:e7fc0947f8a04102b3893bbf0a5b51ac339bda235cf4f37725e3937b9a607666

Observation 38de89c2-d281-47f0-ae08-e4f395a9ff85 · inbound

Artificial Intelligence for Mathematical Reasoning: An Integrated Survey of Language Models, Neuro-symbolic Systems, and Verified Discovery cites this paper.

Artificial Intelligence for Mathematical Reasoning: An Integrated Survey of Language Models, Neuro-symbolic Systems, and Verified Discovery Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 221

Resolution
unresolved
no resolver link, observed 2026-08-02T12:05:18.284698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:05:18.284698Z digest=sha256:ac0a8b07e01472dcc5b1f5acab066120c71ca968afa00d1f5a3dcb6eb3b47c99

Observation 79c3b2ca-0396-4c14-ae59-cdf5706853f7 · inbound

KCSAT-ML: Probing Reasoning Models with Nationwide-Cohort Human Difficulty cites this paper.

KCSAT-ML: Probing Reasoning Models with Nationwide-Cohort Human Difficulty Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-06-27T13:10:55.958616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T13:07:59.509660Z digest=sha256:825fbd19f3c20e8da9ed784d0bde0bc836106adc978a2d1ea01179eb049c23a4

Observation 2aecc018-1a2a-4a5e-8c03-b6b612e973a6 · inbound

Cognitive Episodes in LLM Reasoning Traces Enable Interpretable Human Item Difficulty Prediction cites this paper.

Cognitive Episodes in LLM Reasoning Traces Enable Interpretable Human Item Difficulty Prediction Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-01T17:15:51.636611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-29T03:58:32.372896Z digest=sha256:a9911b379761f6d0a62176bb85943bbe1adadfe68a98f6e9627c8a9aec3ff657

Observation b4d53ea3-6e0a-436d-82fe-6f0fc22e85eb · inbound

Cognitive Episodes in LLM Reasoning Traces Enable Interpretable Human Item Difficulty Prediction cites this paper.

Cognitive Episodes in LLM Reasoning Traces Enable Interpretable Human Item Difficulty Prediction Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-14T17:12:24.565155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T17:12:24.565155Z digest=sha256:7bdf30bedfa9a0737df71a6650ce25665633a94bdfa3a521a7bf9e7149f352db

Observation 604b69a1-63f9-41af-9cbd-e914229b0396 · inbound

Beyond Compilation: Evaluating Faithful Natural-Language-to-Lean Statement Formalization cites this paper.

Beyond Compilation: Evaluating Faithful Natural-Language-to-Lean Statement Formalization Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T13:15:45.420152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-07-01T00:58:00.053021Z digest=sha256:a4741e921a6299020cc286dd4a8dbd8cf448b154f896c6368e337b9f93e3b2fc

Observation 54ae2a75-170f-4185-8607-c99e59f0dca6 · inbound

OS-Pruner: Pruning Chains-of-Thought of Reasoning Models via Optimal Stopping cites this paper.

OS-Pruner: Pruning Chains-of-Thought of Reasoning Models via Optimal Stopping Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-14T07:09:11.486851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T07:09:11.486851Z digest=sha256:44509219e1a298a8ab59214fd2ec5e9cb760c74c4c53efe544ed5d70a0b44328

Observation 96536195-28b5-464e-870e-60c9d5909ec8 · inbound

AdvancedMathBench: A Benchmark Suite for Advanced Mathematical Proof Generation and Verification cites this paper.

AdvancedMathBench: A Benchmark Suite for Advanced Mathematical Proof Generation and Verification Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-14T02:43:21.225324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T02:43:21.225324Z digest=sha256:69f37112c49520ef34cc2c4d1cd67dcc766a3c87208d78e87bdf64cb6d95b8df

Observation 5f3ef200-f229-4067-a7f3-498e6f885606 · inbound

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning cites this paper.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:06.382219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:06.382219Z digest=sha256:b191049ab53d5d44ffc8adf11643944249eef890fcc6fb5a9a9b2b50671e1d13

Observation 95d6c613-43c4-4846-8694-4ff07d2545a9 · inbound

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information cites this paper.

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-01T12:54:32.701608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:54:32.701608Z digest=sha256:13d780e17bebfed93881214ce78f5645d08916d019d7964243d94c38cdb1f60f

Observation c4263efc-e780-4a88-91e7-e55119155782 · inbound

When RLVR Shrinks the Reasoning Boundary: Diagnosing Pass@k Inversion cites this paper.

When RLVR Shrinks the Reasoning Boundary: Diagnosing Pass@k Inversion Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T07:14:52.979212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:14:52.979212Z digest=sha256:9967406af77baf3d068b0fcce0a324ab53d5eb731c830c79b16f2f9f3536faaa