Pith. sign in

Paper Citation Record · LEDGER

DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

As of 4 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 70 inbound Pith citation observations for arXiv:2504.11456.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.11456 v2

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-16T10:31:04.728005Z

measured 92 of 92 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 70 of 70 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T22:56:02.509530Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-09T03:25:57.840663Z

Reference resolution

22 of 22 outbound references displayed

  • verified exact10
  • verified fuzzy1
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch10

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 26485edc-3072-4b24-925c-5d34a4ec5208 · outbound

This paper cites Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models.

DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:31:04.818649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T10:31:04.728005Z digest=sha256:b7ea0977635031db251a9219a7cdff290569c344ea9c175edd82112d26d930a1

Observation b931d842-000d-4550-aadb-f6fe85a48a11 · outbound

This paper cites SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model.

DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T10:31:04.753989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T10:31:04.728005Z digest=sha256:f5ce3d1a046a04029ca5265372215ea3aaf7ab69415bc04c26132a6b40e2663d

Observation 1689582a-6927-475c-8f48-ed77d2dfcab2 · outbound

This paper cites doi: 10.18653/v1/2023.emnlp-main.468.

DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning doi: 10.18653/v1/2023.emnlp-main.468

Reference 3

Resolution
verified exact
doi, observed 2026-05-16T10:31:04.746785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T10:31:04.728005Z digest=sha256:dd0c71e6f9bad28791ad3c26c6edb2e1d3b09492ef450567b3d5628ae8b86af8

Observation 2d5c5678-3006-4d48-a05d-14ba7ab3a382 · outbound

This paper cites Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs.

DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T10:31:04.759798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T10:31:04.728005Z digest=sha256:675f7a044cd076b866eebc8f4c4cc2e1054085fc9fba48e5fe1b8dc0b61ff4db

Observation d4bb8b0b-0c52-479f-810f-09d643e3c0e6 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning Training Verifiers to Solve Math Word Problems

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-16T10:31:04.765546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T10:31:04.728005Z digest=sha256:f79875a8cbeb4e72732eb0547c1282b4d43dd5fe2e50131dd4255d384cf9eaac

Observation 46dd5051-9a31-473e-b041-bdf3603445f9 · outbound

This paper cites Process Reinforcement through Implicit Rewards.

DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning Process Reinforcement through Implicit Rewards

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-16T10:31:04.770958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T10:31:04.728005Z digest=sha256:9b132b346dff9acaefad65a473c7188426b1731e5174fb61296639ecc52514dc

Observation adfc2564-27b3-4004-9acd-8578b2f4171b · outbound

This paper cites Reinforcement learning for reasoning in small llms: What works and what doesn’t.

DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning Reinforcement learning for reasoning in small llms: What works and what doesn’t

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:31:04.776854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T10:31:04.728005Z digest=sha256:e23c41bdcf8256d7000b430dc79092d9e33ea7ab9af0238f2ba755617464b3dd

Observation d7fd0904-dc6b-4b76-8db5-42c463690aca · outbound

This paper cites MathOdyssey: Benchmarking Mathematical Problem-Solving Skills in Large Language Models Using Odyssey Math Data.

DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning MathOdyssey: Benchmarking Mathematical Problem-Solving Skills in Large Language Models Using Odyssey Math Data

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T10:31:04.783711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T10:31:04.728005Z digest=sha256:d9b1ee5e075990f41bdbb6582670d288f141ac053ad7267b6bacf260ec254764

Observation 2033577b-b2b8-4cda-a571-46374563bd24 · outbound

This paper cites Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs.

DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-17T11:40:33.380427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T10:31:04.728005Z digest=sha256:fb0b9a5e28fce41ba51ef52bae0846efeb817024ae771fcec7c48db41270f311

Observation 85cff00c-7d90-46d5-a7ff-42656da1cbae · outbound

This paper cites Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models.

DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T10:31:04.794480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T10:31:04.728005Z digest=sha256:5f52a5351a36efba2d508fc4e737c88e72856bdbb86ad71e0ad524e67b28b2ab

Observation 995d756a-08af-4938-8ce5-992406dc104c · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-16T10:31:04.799231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T10:31:04.728005Z digest=sha256:d856e80b6c5614e523451e87615b25b01654ef8b64069a4c3382c272a67346aa

Observation 630c41d4-ba8c-462e-826c-c081c0fff7ab · outbound

This paper cites OlympicArena: Benchmarking Multi-discipline Cognitive Reasoning for Superintelligent AI.

DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning OlympicArena: Benchmarking Multi-discipline Cognitive Reasoning for Superintelligent AI

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:31:04.804127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T10:31:04.728005Z digest=sha256:6b378f202efab1e8037e3a738aefc8d0db9cbe5ab7ce68ec4e1ba17c23fe97fa

Observation 447177ca-6f1b-4e07-a775-d9c9af612eb1 · outbound

This paper cites an unresolved cited work.

DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-05-16T10:31:04.847365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T10:31:04.728005Z digest=sha256:43984204d935968f64edd3bb99743895817b1cba922b10a8eb12ecd2fa2a6413

Observation d5894fd5-b3aa-497a-bb74-22d163c92241 · outbound

This paper cites Augmenting Math Word Problems via Iterative Question Composing.

DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning Augmenting Math Word Problems via Iterative Question Composing

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T10:31:04.823272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T10:31:04.728005Z digest=sha256:1503e7755b16d47b510c8269081d568d5e8e898e815992ac79a5435512fa4dea

Observation da5f633b-de98-4e75-bf85-fc411aeb19d8 · outbound

This paper cites AIMO-2 Winning Solution: Building State-of-the-Art Mathematical Reasoning Models with OpenMathReasoning dataset.

DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning AIMO-2 Winning Solution: Building State-of-the-Art Mathematical Reasoning Models with OpenMathReasoning dataset

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T10:31:04.827741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T10:31:04.728005Z digest=sha256:3312a5e54d819e4c4a899ced2b5d73e3be32c483963bc6f0b6245f0fd2a573bc

Observation e3eba730-9185-4187-8c01-ad0fe37424e1 · outbound

This paper cites Sentence-bert: Sentence embeddings using siamese bert-networks.

DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning Sentence-bert: Sentence embeddings using siamese bert-networks

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:31:04.850220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T10:31:04.728005Z digest=sha256:3232b812b4d6bdb6c0cd7085388205618bba2d1a50ae075f68f5b0bdd315b32e

Observation 8a257143-90ed-4c71-aeae-4cc1c5ca0954 · outbound

This paper cites Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks.

DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T10:31:04.831503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T10:31:04.728005Z digest=sha256:92ee67ddbd994008d63adcfdbfcd69c84465276699d908107770ef0e6d49bdad

Observation 8564cbcf-f582-49d1-a4be-81653f58aff0 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 19

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T10:31:04.835825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T10:31:04.728005Z digest=sha256:bcbb2e7658b34c8a285230388a2c6646f18a6c96fe680f5dd3f074fe00b76810

Observation 5e1f518d-fd3c-4fe2-a780-f461cb2b8e55 · outbound

This paper cites OpenMathInstruct-2: Accelerating AI for Math with Massive Open-Source Instruction Data.

DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning OpenMathInstruct-2: Accelerating AI for Math with Massive Open-Source Instruction Data

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T10:31:04.839943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T10:31:04.728005Z digest=sha256:c7fa73e9e063bf24cae6245d1bced79843dd235cf5b80341557c113bf7a1399e

Observation 9cc0c66f-1531-4f89-8ac8-883264aa5fc7 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 21

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T10:31:04.843987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T10:31:04.728005Z digest=sha256:5d7d598e7144da2298a665780f025dc238546fd3bebee2ee8dcd5f92b40053fb

Observation 629dd6a8-b5e8-4559-b592-026a21ba435b · outbound

This paper cites SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild.

DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-16T10:31:04.813214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T10:31:04.728005Z digest=sha256:aba33e2636faa40512f4d334a317a8b0c60d533e25e51de93a24f71dca2a5b56

Observation ae8240d8-0d89-450b-876c-8366c45afb75 · outbound

This paper cites MegaMath: Pushing the Limits of Open Math Corpora.

DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning MegaMath: Pushing the Limits of Open Math Corpora

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:31:04.809020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T10:31:04.728005Z digest=sha256:9709b0f182c5413d2722d675e4bc745de8874c571e7ee105e80fd72f92c9310c

Pith citing papers

Observation 47bbb0d6-6ebb-4c6e-a885-d781984ddf9a · inbound

OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles cites this paper.

OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-19T06:59:03.404797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T06:59:03.112252Z digest=sha256:f77ed846fdc1476163e2ed99034c006a4dac40325fca5fe6297ec9886da39efb

Observation 058bade5-4ca3-4d40-aae4-2ec52444afe0 · inbound

Beyond Two-Stage Training: Cooperative SFT and RL for LLM Reasoning cites this paper.

Beyond Two-Stage Training: Cooperative SFT and RL for LLM Reasoning DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T22:56:02.509530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:56:02.509530Z digest=sha256:428e75d3de3bb9f913977877878fbaa94824c7d0ed85a06119f937a1d985e4c0

Observation a60124cb-8aa4-426d-9994-f1153aa919bd · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 193

Resolution
verified exact
local_arxiv, observed 2026-05-18T00:05:31.546896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:830c673a671627105628547b1284d12f97a8851e732fe20b4ee802f3028c4dd5

Observation 723679c4-7e2d-4f29-aa74-cb04aaf5aadf · inbound

Plan Then Action:High-Level Planning Guidance Reinforcement Learning for LLM Reasoning cites this paper.

Plan Then Action:High-Level Planning Guidance Reinforcement Learning for LLM Reasoning DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T12:52:26.866176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:52:26.866176Z digest=sha256:8ba107eaab406df2d8fd3094eba9fda35f847a9070a244ef860ed8b168387ef3

Observation efedbebc-b4da-49ff-8023-bf8b809b4272 · inbound

Probing the Difficulty Perception Mechanism of Large Language Models cites this paper.

Probing the Difficulty Perception Mechanism of Large Language Models DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T11:17:47.562551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:17:47.562551Z digest=sha256:9a87332198025dd88f0efb0018c229481018a4e1c3101461d38a8a1bed51b548

Observation 967a59f4-c332-409d-a416-48d6fa20ece1 · inbound

Rewarding Structural Conformance of Reasoning using Process Mining cites this paper.

Rewarding Structural Conformance of Reasoning using Process Mining DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T07:37:16.657252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:37:16.657252Z digest=sha256:969b87705b5ccb19a69acde2a22da89ceb8e017f9aab25c8f3dbb85794ebf1c8

Observation 7eff8bab-4ce0-4f8b-b813-c430af938f21 · inbound

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments cites this paper.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:05.560042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:05.560042Z digest=sha256:7e72d5da842ad069f9bad2673c86cb3f1ff77c6ab3c4e909c22824f3976bf38c

Observation 5e0f3afa-9e22-4072-a047-c2e47d2f7927 · inbound

SPHINX: A Synthetic Environment for Visual Perception and Reasoning cites this paper.

SPHINX: A Synthetic Environment for Visual Perception and Reasoning DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-17T04:21:30.768641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T04:19:26.808804Z digest=sha256:50b08ee844ecc6c68ae93f1e189b437cff8a336678b3e021c8cf3b6b23e10fde

Observation c73c02c1-8d22-4ee2-bebf-4d4972f38c20 · inbound

SCALER:Synthetic Scalable Adaptive Learning Environment for Reasoning cites this paper.

SCALER:Synthetic Scalable Adaptive Learning Environment for Reasoning DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-16T16:28:05.679417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T16:24:04.132572Z digest=sha256:9ba757dc96221a680b92cfb775ec087376892596d89380ad6e311f258fbc4eac

Observation df99729a-bc3a-4ae8-9baa-f3fa3296fdb8 · inbound

Conversation for Non-verifiable Learning: Self-Evolving LLMs through Meta-Evaluation cites this paper.

Conversation for Non-verifiable Learning: Self-Evolving LLMs through Meta-Evaluation DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:31:04.851073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T10:04:44.379128Z digest=sha256:0cf395dc2900ce54482273f5aee2137bf3d2e767d52e1652d000aa4181e43e7b

Observation cf38665a-e33a-400c-b7a5-84870c2e7f42 · inbound

On the Overscaling Curse of Parallel Thinking: System Efficacy Contradicts Sample Efficiency cites this paper.

On the Overscaling Curse of Parallel Thinking: System Efficacy Contradicts Sample Efficiency DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-16T10:40:51.353235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T10:38:09.875786Z digest=sha256:643716e741cf3b3952ff3f9fe91507571fc33723bfb1af8aa795dda6d70351ef

Observation 860ebc19-c46d-4c0c-8fb0-6d1bd6a6f7f8 · inbound

GradAlign: Gradient-Aligned Data Selection for LLM Reinforcement Learning cites this paper.

GradAlign: Gradient-Aligned Data Selection for LLM Reinforcement Learning DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T21:05:25.136603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:05:25.136603Z digest=sha256:32ba3ac86ed645c92392d3e8ced0a9065d8e3bc095ffdebdf7e2f38025a84c62

Observation f9d9e73e-2c42-43b6-88a5-c55fc99f6445 · inbound

Compress the Easy, Explore the Hard: Difficulty-Aware Entropy Regularization for Efficient LLM Reasoning cites this paper.

Compress the Easy, Explore the Hard: Difficulty-Aware Entropy Regularization for Efficient LLM Reasoning DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T20:44:44.881935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:44:44.881935Z digest=sha256:e7b526fa1fdf8c8ddb5c992af907e1368ba2a2f8235c8d04c90a02145320e785

Observation 690dc376-9abc-4168-8dd7-8c8e5a7cf793 · inbound

Your Model Diversity, Not Method, Determines Reasoning Strategy cites this paper.

Your Model Diversity, Not Method, Determines Reasoning Strategy DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:31:04.851073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T15:16:30.448400Z digest=sha256:24f8316b11f479a16506332ab8cef34db705f338e038832110f1fa2aa9ed840d

Observation 07c5dbba-d228-4d05-aeae-fa4ae0032d28 · inbound

PubSwap: Public-Data Off-Policy Coordination for Federated RLVR cites this paper.

PubSwap: Public-Data Off-Policy Coordination for Federated RLVR DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:31:04.851073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:05:27.319466Z digest=sha256:6cc2e7d052b6cf0d1e84d5f6118811d88b30d5228c2a47ffc3e521253fd54852

Observation c4d001c6-05c1-4863-a9cb-572442a58a36 · inbound

Rethinking the Comparison Unit in Sequence-Level Reinforcement Learning: An Equal-Length Paired Training Framework from Loss Correction to Sample Construction cites this paper.

Rethinking the Comparison Unit in Sequence-Level Reinforcement Learning: An Equal-Length Paired Training Framework from Loss Correction to Sample Construction DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:31:04.851073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T06:23:14.905706Z digest=sha256:268f4d2200831b7a9ae649c34d3d5b726b8e116b7b5624bb60becbc122fc9a52

Observation 94fced4b-444c-4351-a118-0b4e0b12381f · inbound

Learning from Less: Measuring the Effectiveness of RLVR in Low Data and Compute Regimes cites this paper.

Learning from Less: Measuring the Effectiveness of RLVR in Low Data and Compute Regimes DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:31:04.851073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:55:18.468593Z digest=sha256:3751f88851dca658d40d771e9893b74d75d438807f35e9055ec7e5ac77992384

Observation 0d3b114b-7072-4f2c-83fe-8eb9bab0386a · inbound

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data cites this paper.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 73

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T10:31:04.851073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:0b4fd1ffe94502a461e1d6ea112e4823a42c3ba6dd85b5395833b73cb3b91997

Observation 8997fca0-e118-4973-a178-c56d88e2a3c4 · inbound

Length Value Model: Scalable Value Pretraining for Token-Level Length Modeling cites this paper.

Length Value Model: Scalable Value Pretraining for Token-Level Length Modeling DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:31:04.851073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-07T10:42:27.644514Z digest=sha256:67c7bf4323e7a5a4d6cd3913be8701bf6bf9cf20af90ae78187bcda44a68b65f

Observation 3e7b28f8-4b77-43c5-b8f0-cb93f6526045 · inbound

Length Value Model: Scalable Value Pretraining for Token-Level Length Modeling cites this paper.

Length Value Model: Scalable Value Pretraining for Token-Level Length Modeling DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T15:28:11.698960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T15:28:11.698960Z digest=sha256:2016a23f1afd077aabfa047588afee585c8c64cf4cb612c6b16b8f54f4a15590

Observation 84e34f6c-4458-468d-9d27-fca1959a941f · inbound

Internalizing Outcome Supervision into Process Supervision: A New Paradigm for Reinforcement Learning for Reasoning cites this paper.

Internalizing Outcome Supervision into Process Supervision: A New Paradigm for Reinforcement Learning for Reasoning DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:31:04.851073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T06:13:09.898530Z digest=sha256:90a902e22b86383a9417ab2b595f0a35d65e791fcf3c5346f211f0bd814f4524

Observation bfebc584-9595-40c8-a585-a60125dd2e88 · inbound

Asymmetric On-Policy Distillation: Bridging Exploitation and Imitation at the Token Level cites this paper.

Asymmetric On-Policy Distillation: Bridging Exploitation and Imitation at the Token Level DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T10:31:04.851073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-08T12:57:24.822525Z digest=sha256:e672f8827148489076081fccb20fe56bef1285ef92714e98a2533415c2dca692

Observation 32a3b2f2-5629-4d01-ab47-dc53fee74265 · inbound

Asymmetric On-Policy Distillation: Bridging Exploitation and Imitation at the Token Level cites this paper.

Asymmetric On-Policy Distillation: Bridging Exploitation and Imitation at the Token Level DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T10:31:04.851073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T01:51:29.021423Z digest=sha256:6e26947d7f1df038a03ac85cea8feb32409ed98f09d826e6b40281cb34511c75

Observation 0515b3fd-630d-4969-bf5d-2338303bdbb1 · inbound

Asymmetric On-Policy Distillation: Bridging Exploitation and Imitation at the Token Level cites this paper.

Asymmetric On-Policy Distillation: Bridging Exploitation and Imitation at the Token Level DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T10:31:04.851073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-14T21:14:22.718769Z digest=sha256:52656cc5b5d6370441a5f9a530eb69d662ce6cf50521c10160bc9cecfda80333

Observation efa33de1-b97f-4433-a238-25112f42335b · inbound

On the Implicit Reward Overfitting and the Low-rank Dynamics in RLVR cites this paper.

On the Implicit Reward Overfitting and the Low-rank Dynamics in RLVR DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:31:04.851073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T12:40:53.063991Z digest=sha256:6dfb498c043d06d9002478377618b2b2178611ae194189c35d0249808109ddc0

Observation 685b4bac-0184-4bf8-99b8-1ceab8937d93 · inbound

AIPO: Learning to Reason from Active Interaction cites this paper.

AIPO: Learning to Reason from Active Interaction DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:31:04.851073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T01:17:28.124867Z digest=sha256:4ae82ecb67efe1aeedd1fbe99aa0790f103726e6b3228a5cdab82bd88106dc20

Observation 47286ca3-2199-4373-9bf6-b33efca7cdb3 · inbound

AIPO: Learning to Reason from Active Interaction cites this paper.

AIPO: Learning to Reason from Active Interaction DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-19T18:07:42.397935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T18:07:27.492419Z digest=sha256:5a7af5c2bcd1eafdf60a8525f5fdbe2a0b6ab0cb14791ea1824ca68d2448d6df

Observation b35cf74e-1a9d-4ee5-ada3-20189188e3c4 · inbound

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models cites this paper.

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:31:04.851073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T01:11:50.343466Z digest=sha256:4378ca3012a73d8e980947caaea9bc7401ba31c936f64fb5ab2b3678206bf986

Observation 894a6985-5629-4100-99a9-42a827d32dfa · inbound

Training with Harnesses: On-Policy Harness Self-Distillation for Complex Reasoning cites this paper.

Training with Harnesses: On-Policy Harness Self-Distillation for Complex Reasoning DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:31:04.851073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T03:51:22.454987Z digest=sha256:3eae5099d6244e8a93460abca44c2191fbe37820c2e2afcac4bf1f26b51b6825

Observation 7472f34e-0af5-472c-885f-591472bf5d2a · inbound

Forge: Quality-Aware Reinforcement Learning for NP-Hard Optimization in LLMs cites this paper.

Forge: Quality-Aware Reinforcement Learning for NP-Hard Optimization in LLMs DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T10:31:04.851073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T02:44:33.143247Z digest=sha256:bd26db0223c2914bd51fa2258b0099a009d2aa02bb968ff3e760e78976bc8d41

Observation 22d33486-cbbf-4c6a-95c9-20ae25de4097 · inbound

DARE: Difficulty-Adaptive Reinforcement Learning with Co-Evolved Difficulty Estimation cites this paper.

DARE: Difficulty-Adaptive Reinforcement Learning with Co-Evolved Difficulty Estimation DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:31:04.851073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T03:12:49.428954Z digest=sha256:4598857afd50c0a8508459322087f0a0837dc6d9934c1702a76fca2379d90364

Observation e0eb2663-a354-47ec-91a2-097eb8ad1ec1 · inbound

Nice Fold or Hero Call: Learning Budget-Efficient Thinking for Adaptive Reasoning cites this paper.

Nice Fold or Hero Call: Learning Budget-Efficient Thinking for Adaptive Reasoning DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:31:04.851073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T01:31:21.377700Z digest=sha256:cde28956ea8b3269a97642c6622656eda882e151d438903046aeecc6ea374b62

Observation 63e4046a-1b0f-43a9-99ce-44fb76fbc868 · inbound

Learning to Foresee: Unveiling the Unlocking Efficiency of On-Policy Distillation cites this paper.

Learning to Foresee: Unveiling the Unlocking Efficiency of On-Policy Distillation DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-22T10:11:22.765572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T10:10:14.565995Z digest=sha256:fd6d9a8d1cf597af873357cfd97f2ffa6bf589cbf4cff775254d2ea3e390b7e4

Observation 23eeed74-71e7-412f-b5eb-958a2afb8867 · inbound

Pion: A Spectrum-Preserving Optimizer via Orthogonal Equivalence Transformation cites this paper.

Pion: A Spectrum-Preserving Optimizer via Orthogonal Equivalence Transformation DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:31:04.851073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T04:53:52.898843Z digest=sha256:a1de303b3f2539c2e95124470da9a9ff3ca8401b654276aa2f9d7e2d8f728256

Observation ad9b5f8d-ce88-4956-b600-102f09a6f772 · inbound

Multi-Rollout On-Policy Distillation via Peer Successes and Failures cites this paper.

Multi-Rollout On-Policy Distillation via Peer Successes and Failures DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 73

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T10:31:04.851073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-14T21:29:36.803832Z digest=sha256:556ba166396b4d55ee79557ba2e8aca579f3190ae6b0a309972a4e221875ebd1

Observation 550971ca-a461-4aae-96f8-1821c3e2c580 · inbound

Taming the Thinker: Conditional Entropy Shaping for Adaptive LLM Reasoning cites this paper.

Taming the Thinker: Conditional Entropy Shaping for Adaptive LLM Reasoning DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 23

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T06:43:05.906903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-20T06:40:06.103206Z digest=sha256:b362207e8e1776433c1787182b97624f9eeab62d49f393c22bebae405d42861c

Observation e4b2e5c6-750a-41d0-b97e-86ef77a03cc4 · inbound

Beyond Mode Collapse: Distribution Matching for Diverse Reasoning cites this paper.

Beyond Mode Collapse: Distribution Matching for Diverse Reasoning DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-20T05:33:04.086260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T05:30:37.685873Z digest=sha256:ab12ac6c119ec80c76c9d89fdde72a46bd95a104dacaef9b2b4556a702416a79

Observation 971dcc5f-8264-42b0-8b6c-464b61374a60 · inbound

DEL: Digit Entropy Loss for Numerical Learning of Large Language Models cites this paper.

DEL: Digit Entropy Loss for Numerical Learning of Large Language Models DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:39:49.460219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T07:35:22.992256Z digest=sha256:b90627f005c7e568b2d52731dfdb85416e00d7a333508024151d155a10817974

Observation bd7ea9ff-da6f-420d-bdd9-6ee31bad6352 · inbound

DelTA: Discriminative Token Credit Assignment for Reinforcement Learning from Verifiable Rewards cites this paper.

DelTA: Discriminative Token Credit Assignment for Reinforcement Learning from Verifiable Rewards DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-05-21T05:29:40.132496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-21T05:24:46.545570Z digest=sha256:599b5a191dfb955c1df4c4e63cf5c05ddbd54bc369a6740646524a1938572b5b

Observation 8bb3ca51-c57b-4648-895e-466727ea8ea2 · inbound

LANG: Reinforcement Learning for Multilingual Reasoning with Language-Adaptive Hint Guidance cites this paper.

LANG: Reinforcement Learning for Multilingual Reasoning with Language-Adaptive Hint Guidance DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 31

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T06:21:10.412848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-22T06:19:44.377733Z digest=sha256:812bda1fdd9a70f6b3d89428944e4dced70196ef83b1c5f5dc69d6027fb58432

Observation e880cc08-4db1-45fc-9a62-cd7c46fb811b · inbound

Not only where, But when: Temporal Scheduling for RLVR cites this paper.

Not only where, But when: Temporal Scheduling for RLVR DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:44:01.863471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T22:34:44.791173Z digest=sha256:3933a89aae3a23c61db2d593b489043aab05b9b77eebd60959400f10d0e548d9

Observation ea838eab-8e97-476f-8d80-a2736669e7f2 · inbound

Spend Your Rollouts Where It Counts: Rollout Allocation for Group-Based RL Post-Training cites this paper.

Spend Your Rollouts Where It Counts: Rollout Allocation for Group-Based RL Post-Training DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-06-29T19:53:55.807818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T19:47:09.817243Z digest=sha256:63c9f1cfbb3ffbd0ee3021003616859df6053b77258687e1c9ea8520d6ddb439

Observation bda730de-94b9-4c7c-8f93-9450c7624124 · inbound

RLVR Datasets and Where to Find Them: Tracing Data Lineage for Better Training Data cites this paper.

RLVR Datasets and Where to Find Them: Tracing Data Lineage for Better Training Data DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-06-29T19:43:55.115046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-29T19:34:12.081362Z digest=sha256:a51bf71a52a6a8d638a3203b44f18a11e0bd9e49f444e2e501c1f12622a3c9eb

Observation 9bee5b38-0f96-4678-a3e3-9faac621eb5c · inbound

DenoiseRL: Bootstrapping Reasoning Models to Recover from Noisy Prefixes cites this paper.

DenoiseRL: Bootstrapping Reasoning Models to Recover from Noisy Prefixes DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-06-29T11:53:23.707364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T11:51:01.769882Z digest=sha256:830ae2d74c681774876f6a0adfa1f70a8c7f9ee53ea7cc4735d41e28c2ac3c76

Observation d7258fbd-1837-4aa7-8859-85c5e3b47204 · inbound

DenoiseRL: Bootstrapping Reasoning Models to Recover from Noisy Prefixes cites this paper.

DenoiseRL: Bootstrapping Reasoning Models to Recover from Noisy Prefixes DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T12:58:10.897147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:58:10.897147Z digest=sha256:1d4108661217ff16b02f73d5d9edbe43ec5c5364017bcfa809bca0e15627264a

Observation a769edab-07f2-4b0d-9d8a-e2a919c54835 · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 287

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T20:56:13.352648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:cd46816c941f46b66d3c65b9eeda51091f62897261cc639b30200127eda2805e

Observation a757bd81-2922-43b1-bea3-55363c01d36a · inbound

Filter, Then Reweight: Rethinking Optimization Granularity in On-Policy Distillation cites this paper.

Filter, Then Reweight: Rethinking Optimization Granularity in On-Policy Distillation DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 22

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T22:36:17.403843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-28T15:14:42.489647Z digest=sha256:65019c89ec7cb2090d1eec173d0720e455e1f8ff768d2c0d21696eed5335e4e4

Observation 7fb7b005-ff50-4cc2-a597-31b394148e56 · inbound

ThoughtFold: Folding Reasoning Chains via Introspective Preference Learning cites this paper.

ThoughtFold: Folding Reasoning Chains via Introspective Preference Learning DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T03:26:28.773294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T10:07:16.700499Z digest=sha256:6694437e2d1ccd02da355c3a12563d7fec06594af03bb3a7c0887bec78a66391

Observation c980f253-16a5-419b-9818-56453ac27ce8 · inbound

GeoMin: Data-Efficient Semi-Supervised RLVR via Geometric Distribution Modeling cites this paper.

GeoMin: Data-Efficient Semi-Supervised RLVR via Geometric Distribution Modeling DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 26

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T06:06:40.776163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-28T07:45:43.320339Z digest=sha256:241d43fe245e29a6a82895f9be5899439baec0e87db116179eb02c430b4b7689

Observation 9a80c423-7e11-43b9-8251-bb8b81bd5859 · inbound

GRAIL: Gradient-Reweighted Advantages for Reinforcement Learning with Verifiable Rewards cites this paper.

GRAIL: Gradient-Reweighted Advantages for Reinforcement Learning with Verifiable Rewards DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-02T08:36:47.619392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T05:59:00.005336Z digest=sha256:df8ab13e72a0bf60c7249a38d3d54bc39c74386321b8e548da2910879bc679d7

Observation b425e865-08fb-4d09-aafc-324d5546c6da · inbound

State commitment learning: training language models to distinguish computation from memory cites this paper.

State commitment learning: training language models to distinguish computation from memory DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-06-30T15:24:49.920750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-30T15:21:49.337222Z digest=sha256:435ef9e6736c14e38ede939a983a7c6e9151a8a35dcb5d1778bc6ef4a884380b

Observation 457f3a84-162b-4381-8e5e-f7a65fda66f9 · inbound

ConSteer-RL: Steering Reasoning Capabilities in Large Language Models via Confidence-Aware Reinforcement Learning cites this paper.

ConSteer-RL: Steering Reasoning Capabilities in Large Language Models via Confidence-Aware Reinforcement Learning DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T21:07:23.905688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-27T19:56:09.820812Z digest=sha256:d34eb62127af5836e1db721f6899013585fbfc9f901faf45bac8a850e3214775

Observation 8249fa7b-1492-4cc3-a41a-584ac4c82c17 · inbound

SG-OPD: Sign-Gated On-Policy Distillation via Sign-Consistency Gating and Phased Teacher Sampling cites this paper.

SG-OPD: Sign-Gated On-Policy Distillation via Sign-Consistency Gating and Phased Teacher Sampling DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T01:17:31.317061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-27T16:37:23.465811Z digest=sha256:85066551050bc49608949625fec65c8df3f616f85f4e6273887979b4c41cd120

Observation 8c83603f-2553-43f7-9c8a-8d0ba9286330 · inbound

RLCSD: Reinforcement Learning with Contrastive On-Policy Self-Distillation cites this paper.

RLCSD: Reinforcement Learning with Contrastive On-Policy Self-Distillation DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-03T09:07:48.370451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-27T10:28:18.490452Z digest=sha256:d2a731cca094856bf8dabeb7ee1499e0f90aff3b42c9a5880dbfa6e75e986375

Observation 365dd4e4-44fb-4a6a-980e-69e8631d0689 · inbound

Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs cites this paper.

Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-07-03T21:18:58.389882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-27T00:48:56.892634Z digest=sha256:74cd1fee2280adafb5453f4a8e247b38b671583587f148ab0c7a8f583b8b389a

Observation 74c5ef6d-7dda-442c-b49a-f823e2885d23 · inbound

Finding the Evidence: Discovering Decision-Supporting Tokens for On-Policy Reasoning Distillation cites this paper.

Finding the Evidence: Discovering Decision-Supporting Tokens for On-Policy Reasoning Distillation DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T10:29:44.952266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T08:49:58.735297Z digest=sha256:28d926c13d595f0dd42612a75f4bdc652638f7a734a9f8a4e553d98374f6a055

Observation 1cc6cde3-6c36-4222-87cd-4a0aa0896f7b · inbound

AsyncOPD: How Stale Can On-Policy Distillation Be? cites this paper.

AsyncOPD: How Stale Can On-Policy Distillation Be? DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-04T16:29:56.821408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T00:40:57.112748Z digest=sha256:0a1335c55648ab077234fc177db20b56875f372ace077322eb37ec6295326b2f

Observation 2832645c-535e-4fe9-8480-08a3aa464086 · inbound

EntroRouter: Learning Efficient Model Routing via Entropy Regulation cites this paper.

EntroRouter: Learning Efficient Model Routing via Entropy Regulation DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 45

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T07:44:21.623493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-30T07:42:35.962864Z digest=sha256:9f83a9f09c0f04d0c0cf742954c50a9eec833e0bd277b107e221dcad8e5bd77a

Observation ea29cd36-3041-4a6c-aa58-ad151fe1e294 · inbound

The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning cites this paper.

The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T07:24:20.891805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T07:23:19.403334Z digest=sha256:2dcd1043dd3e8f9f0a47d6edc174e5ac1b91f9dd22eb32aa40ddcb672cf0c6fa

Observation 0cd46127-b12a-4cf3-9c6a-f76ce529268f · inbound

Beyond Compilation: Evaluating Faithful Natural-Language-to-Lean Statement Formalization cites this paper.

Beyond Compilation: Evaluating Faithful Natural-Language-to-Lean Statement Formalization DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 19

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T13:15:45.431035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-07-01T00:58:00.053021Z digest=sha256:28aa9fc33196343dd3e86ea453c9c8acaa01e04572b1348a5a012dae7eae8159

Observation f640058f-4066-40c0-b8b7-8541b8dd9730 · inbound

On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models cites this paper.

On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-11T20:01:06.623978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:01:06.623978Z digest=sha256:664f94d5e4a9a5388c45fe318039cf17c322c47380715084f8acd06f8689fc56

Observation 78c2130d-9121-4dd8-9f88-afced1b0401f · inbound

Weak-to-Strong Generalization via Direct On-Policy Distillation cites this paper.

Weak-to-Strong Generalization via Direct On-Policy Distillation DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-07T12:33:45.202425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-07T12:31:42.224094Z digest=sha256:bb5c5b23799050e2629d8f05b70e5ddd8efc1a4020c3bca9a3313b48ffcdeaa1

Observation 6cf49e12-3bfb-4301-9504-f6cef34cf187 · inbound

Weak-to-Strong Generalization via Direct On-Policy Distillation cites this paper.

Weak-to-Strong Generalization via Direct On-Policy Distillation DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:4bd2ba1133cbeeef1fae690a21c29e8dee71de3c3fbdc67788ef692f198076ed

Observation 2dea5867-8b5f-49dc-9414-306428d24929 · inbound

Max Out GRPO Signal: Adaptive Trace Prefix Control for Hard Reasoning Problems cites this paper.

Max Out GRPO Signal: Adaptive Trace Prefix Control for Hard Reasoning Problems DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:25:57.841864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-09T03:20:40.229376Z digest=sha256:6a4df3e8479ceaeedb4ecff566bacb67e49265308570ad672b182da9d063cf7f

Observation 8b390d5b-bf7a-42a2-b48e-ce75570111dc · inbound

Agon: Competitive Cross-Model RL with Implicit Rival Grading of Reasoning cites this paper.

Agon: Competitive Cross-Model RL with Implicit Rival Grading of Reasoning DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-09T02:35:53.868114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-09T02:29:12.366018Z digest=sha256:808568d3c0fede8098b342bba41e9378e409f887a51f28798ba034b9650d5e86

Observation 8c12d294-4409-4f4e-a72f-2f576ca10288 · inbound

Proxy Exploration and Reusable Guidance: A Modular LLM Post-Training Paradigm via Proxy-Guided Update Signals cites this paper.

Proxy Exploration and Reusable Guidance: A Modular LLM Post-Training Paradigm via Proxy-Guided Update Signals DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-14T05:09:00.865375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T05:09:00.865375Z digest=sha256:8e78c71624f13f4713fae32eea328520f4c7ef10d9cc065ca947c30b24097dd1

Observation 3b5f5768-87a2-432a-88c7-92682d12d1ef · inbound

Cross-Tokenizer On-Policy Distillation via Byte-Prefix Marginalization cites this paper.

Cross-Tokenizer On-Policy Distillation via Byte-Prefix Marginalization DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-01T05:09:58.972025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T05:09:58.972025Z digest=sha256:d9ec9ba7538b2bd2c6006ff33188c31c09ce1d40a1264e59b3e6495f7cd6c398

Observation 3c36225b-6de3-4a63-8363-09d4bb8883ad · inbound

Weak-to-Strong On-Policy Distillation cites this paper.

Weak-to-Strong On-Policy Distillation DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:24.869605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:24.869605Z digest=sha256:3ab876574070221ceb76bf8f4da4932c9e1fb1bbdbe489dc4e5e2aaa1212926c

Observation 72fd49f8-e1e1-40ae-961e-20159b36095b · inbound

LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts cites this paper.

LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T01:28:56.250910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:28:56.250910Z digest=sha256:14e8505cb2dbfa75f91d864b7920a4637ed38a83b61f70f308a22ad63d7f3c91

Observation b63273ff-1c3c-404e-80a4-b52983f59eaf · inbound

ThinkReset: Learnable Intermediate Interface Construction for Bounded-Context Long-Horizon Reasoning cites this paper.

ThinkReset: Learnable Intermediate Interface Construction for Bounded-Context Long-Horizon Reasoning DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T00:54:34.691461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T00:54:34.691461Z digest=sha256:4faede988b6031ff1b0be0201396cef0ae431bf1b5666d5336ff6a330965b8ee