Pith. sign in

Paper Citation Record · LEDGER

Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models

As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 39 inbound Pith citation observations for arXiv:2502.17387.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.17387 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 39 of 39 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:39:31.109183Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T16:39:58.415508Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 26485edc-3072-4b24-925c-5d34a4ec5208 · inbound

DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning cites this paper.

DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:31:04.818649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T10:31:04.728005Z digest=sha256:7c64ecd0b42c86404a1984f64eba2fd806bc6e14adba491a19b4f904a34e24b1

Observation ee83f16d-1103-4596-9b5b-2018b82c3374 · inbound

DeepDistill: Enhancing LLM Reasoning Capabilities via Large-Scale Difficulty-Graded Data Training cites this paper.

DeepDistill: Enhancing LLM Reasoning Capabilities via Large-Scale Difficulty-Graded Data Training Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T10:39:31.109183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:39:31.109183Z digest=sha256:b01a6968f280e6fc4c142e130466479cdb7041eeecfd20a2a21463e4d65fc41e

Observation 33a3da34-7ee6-46b7-a0f3-3e436693dc6f · inbound

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models cites this paper.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:43.670016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:43.670016Z digest=sha256:e60b595658cd2ff70da1440e64632669403f86849c46603a73a607792c46eb12

Observation cfc6ffb7-03ba-4eba-9fd6-71af80d16ba8 · inbound

Exploring the Potential of Offline RL for Reasoning in LLMs: A Preliminary Study cites this paper.

Exploring the Potential of Offline RL for Reasoning in LLMs: A Preliminary Study Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:38.684571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:38.684571Z digest=sha256:1d7da01bdfdc89988cfae73087f82878d455cce56b4174132fee947893f00f26

Observation 8caaf516-6c3a-4be0-be69-13f6632377a8 · inbound

Nature's Insight: A Novel Framework and Comprehensive Analysis of Agentic Reasoning Through the Lens of Neuroscience cites this paper.

Nature's Insight: A Novel Framework and Comprehensive Analysis of Agentic Reasoning Through the Lens of Neuroscience Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models

Reference 205

Resolution
unresolved
no resolver link, observed 2026-08-15T23:31:12.685186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:31:12.685186Z digest=sha256:54ef7a2d055a0eed36416f06923656257193fd819da471c5e6766f8902fc5461

Observation ef46c4eb-d007-43dd-bcac-af4fe3fbe2f0 · inbound

Reward Reasoning Model cites this paper.

Reward Reasoning Model Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:50.264050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:50.264050Z digest=sha256:f769de7c7e19da3f0dfd48bf6157aa4019e3399a915e4bfde6e0bd027fcea989

Observation 7b167c31-980d-4731-958d-ea1dbf8b670d · inbound

SituatedThinker: Grounding LLM Reasoning with Real-World through Situated Thinking cites this paper.

SituatedThinker: Grounding LLM Reasoning with Real-World through Situated Thinking Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:21:18.385545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:21:18.385545Z digest=sha256:f43e435913c742ec439f186a8aa8fc9a2bbcba7455c75f11fd13388bd704a30f

Observation 587d4f97-b00c-4cbe-a252-20115473630a · inbound

Improving Multilingual Math Reasoning for African Languages cites this paper.

Improving Multilingual Math Reasoning for African Languages Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:10:22.136460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:10:22.136460Z digest=sha256:7f0067cfbe2bdad9f3e26cd55f2b024988f638fd48ed03d9e0309d5a48541426

Observation f97a8a18-1b56-469a-95f7-1f54877170e9 · inbound

Enigmata: Scaling Logical Reasoning in Large Language Models with Synthetic Verifiable Puzzles cites this paper.

Enigmata: Scaling Logical Reasoning in Large Language Models with Synthetic Verifiable Puzzles Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:06.937762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:11:06.937762Z digest=sha256:207e965acd9588e6544509ae40989b6f5ae86d71f4cda5d7f2eb68169b3bceb3

Observation e21e5946-6e2c-4958-aa70-05157d3d6c78 · inbound

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning cites this paper.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:24.089154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:24.089154Z digest=sha256:db1116a25485072a9b6d6193265bce3de18aab8caa93d8e3ec7879d5ed06f00f

Observation 791f41b8-8b5a-4841-9f32-c105a5483f26 · inbound

Ring-lite: Scalable Reasoning via C3PO-Stabilized Reinforcement Learning for LLMs cites this paper.

Ring-lite: Scalable Reasoning via C3PO-Stabilized Reinforcement Learning for LLMs Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:22.426243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:22.426243Z digest=sha256:a49ceccd1456e2e33a09f8da7452c3b27647c3cfe8fb535a42c32a00485cbb26

Observation f41dd2f5-c890-404e-8a61-91943603b5a6 · inbound

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization cites this paper.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.505183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.505183Z digest=sha256:ea4c41ca2edf4f96ecaab717d2af57575f67af48d3c562a15d54236201677433

Observation 92b27c15-f29f-4181-a11c-98ad96e447ed · inbound

Seed-Prover: Deep and Broad Reasoning for Automated Theorem Proving cites this paper.

Seed-Prover: Deep and Broad Reasoning for Automated Theorem Proving Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T10:30:44.585630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:30:44.585630Z digest=sha256:80dc76ed9f463aec89042f22a8b54775a1ab795ebb3172796b842b22a72497a8

Observation 00f51b38-fa8d-4a5c-8994-ee8640e98248 · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:02:24.840613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:745afb14ccca1cbd1c9b9d2cc7a87f21db8516689df15a09334a6861071fc019

Observation 871fff2f-8de7-47c9-8004-2127fc9f1b20 · inbound

EngiBench: A Benchmark for Evaluating Large Language Models on Engineering Problem Solving cites this paper.

EngiBench: A Benchmark for Evaluating Large Language Models on Engineering Problem Solving Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-18T15:11:32.564082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T15:08:03.793304Z digest=sha256:15219441966bf207f4b8efa9f5fa277e245fbc7db95c4f30957ffdd18f7768c1

Observation d02e1bcf-dcbf-4c8c-a7af-9087641b776b · inbound

Breaking the Self-Confirming Loop: Diagnosing and Mitigating Systemic Reward Bias in Self-Rewarding RL cites this paper.

Breaking the Self-Confirming Loop: Diagnosing and Mitigating Systemic Reward Bias in Self-Rewarding RL Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T10:44:29.630216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:44:29.630216Z digest=sha256:a5a750e85073bb4fd3459fa682c11200d99d5e3660643c68e4c5fa12ea276e6a

Observation e18af844-6888-4e1a-b75e-da20f4075c36 · inbound

SPHINX: A Synthetic Environment for Visual Perception and Reasoning cites this paper.

SPHINX: A Synthetic Environment for Visual Perception and Reasoning Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-17T04:21:30.756673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T04:19:26.808804Z digest=sha256:c9632ceed4fc9651fc75dcfad409ed7e4153a95a35cfb7dde16b6cff8918546c

Observation f42e5f60-9a6c-4cfc-a001-1d14b14a6a21 · inbound

GradAlign: Gradient-Aligned Data Selection for LLM Reinforcement Learning cites this paper.

GradAlign: Gradient-Aligned Data Selection for LLM Reinforcement Learning Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-02T21:05:24.244471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:05:24.244471Z digest=sha256:f0eca7e012013883e8b8b8eeece13a918ef0fda5f6110d4b9c8669fdf83aa305

Observation 255142bb-7241-47ae-aaf3-15fc3d45213d · inbound

On the Role of Reasoning Patterns in the Generalization Discrepancy of Long Chain-of-Thought Supervised Fine-Tuning cites this paper.

On the Role of Reasoning Patterns in the Generalization Discrepancy of Long Chain-of-Thought Supervised Fine-Tuning Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T22:08:20.651495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-13T22:08:14.367227Z digest=sha256:0798e9ec6b121560e277619f469f1d5f95fbe7d3cf3dd03e7b121d95ac7962df

Observation 45719082-7f4e-4f94-98cd-19f9444c2ce3 · inbound

Detecting and Suppressing Reward Hacking with Gradient Fingerprints cites this paper.

Detecting and Suppressing Reward Hacking with Gradient Fingerprints Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T08:22:37.577093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T08:18:37.665350Z digest=sha256:2d60f46b8cda31ace849cd297b2b76dad0b2563315fdce82e15d0d83dbae9616

Observation 69b63e41-8a87-47a3-8f37-4ce773ee594c · inbound

MathNet: a Global Multimodal Benchmark for Mathematical Reasoning and Retrieval cites this paper.

MathNet: a Global Multimodal Benchmark for Mathematical Reasoning and Retrieval Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T12:06:04.085088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-10T04:09:46.616019Z digest=sha256:246aba4ae5573df970ee9a8b07ec1868e2dc2cd346bd913570e246734ab2b66b

Observation 4116eac9-a3c3-40b1-8972-e9d431b9f50c · inbound

Evaluating Answer Leakage Robustness of LLM Tutors against Adversarial Student Attacks cites this paper.

Evaluating Answer Leakage Robustness of LLM Tutors against Adversarial Student Attacks Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:01:02.985291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-10T04:22:31.566907Z digest=sha256:e3ae2b8d4e7f60fddc9574e3e5daad049528bee4a95dbcc97b8c49c9cb332987

Observation c422f2e8-8fcc-4627-85cc-0db327271bfd · inbound

$R^2$-dLLM: Accelerating Diffusion Large Language Models via Spatio-Temporal Redundancy Reduction cites this paper.

$R^2$-dLLM: Accelerating Diffusion Large Language Models via Spatio-Temporal Redundancy Reduction Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:46:25.484379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T02:56:32.367404Z digest=sha256:84bdd3ff62bc21d8fc7e1f4cb220e8d5f5bc99bb18f73c49f356cf170b55d8d6

Observation 65af2582-dfcd-4132-a8b4-7b49bf4b24cd · inbound

Forge: Quality-Aware Reinforcement Learning for NP-Hard Optimization in LLMs cites this paper.

Forge: Quality-Aware Reinforcement Learning for NP-Hard Optimization in LLMs Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T02:46:18.885671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-12T02:44:33.143247Z digest=sha256:8dff25d6ea1d369077928c74db8a2a09c265b6776ea013d2ec0a4e110aa33784

Observation 64990604-f0ab-4334-9870-fad57353e7d4 · inbound

Scalable Token-Level Hallucination Detection in Large Language Models cites this paper.

Scalable Token-Level Hallucination Detection in Large Language Models Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:52:22.428508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-13T05:49:23.534294Z digest=sha256:aeb9312bd7dea14d8e2ac220481975e3ea1b16229c483d00c621504509997782

Observation 2107cf75-4187-41e9-931b-7d103dc5108d · inbound

Stress-Testing the Reasoning Competence of LLMs With Proofs Under Minimal Formalism cites this paper.

Stress-Testing the Reasoning Competence of LLMs With Proofs Under Minimal Formalism Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models

Reference 57

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T21:19:29.208078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-14T21:03:13.407192Z digest=sha256:4a73c79ca5d5abf63310466e4b3759bcdd5e3213acdde51097a37633171ccf58

Observation ae0853b1-1f32-48ed-a850-59f1bd637ca3 · inbound

Boosting Reinforcement Learning with Verifiable Rewards via Randomly Selected Few-Shot Guidance cites this paper.

Boosting Reinforcement Learning with Verifiable Rewards via Randomly Selected Few-Shot Guidance Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-15T03:19:43.023601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T03:18:26.590871Z digest=sha256:c3ca5d2d51748a3b2546001f424455dddb7235e4532463e59bfb3c56bdd99897

Observation 1dbdf28d-a2cf-4b39-84ba-34d5fe2097de · inbound

MONA: Muon Optimizer with Nesterov Acceleration for Scalable Language Model Training cites this paper.

MONA: Muon Optimizer with Nesterov Acceleration for Scalable Language Model Training Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-29T19:23:54.125590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-29T19:15:49.229099Z digest=sha256:87eec14dfc326be4eafcb6567a4bed641149227ec2e30506a01178418c1574aa

Observation 9c301df9-2be6-498b-a8b6-79e6b47f0dc2 · inbound

RLVR Datasets and Where to Find Them: Tracing Data Lineage for Better Training Data cites this paper.

RLVR Datasets and Where to Find Them: Tracing Data Lineage for Better Training Data Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-06-29T19:43:55.056489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-29T19:34:12.081362Z digest=sha256:3f8a2f34843d1b64e5bdb19dc24836810b7f06bf0f848efe0adc6319a414babe

Observation 1bf67cc3-4465-4359-93aa-6d8cc8eba11f · inbound

LLMs Are Already Good Tutors: Training-Free Prompt Optimization for Pedagogical Math Tutoring cites this paper.

LLMs Are Already Good Tutors: Training-Free Prompt Optimization for Pedagogical Math Tutoring Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-29T17:53:47.042681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-29T17:50:46.072234Z digest=sha256:12286efa71e462d2b0113fb7d66dd37a5999713447eeb301c9742400b5b759aa

Observation ae0ae862-e140-4be1-996d-8f5b4db0e93b · inbound

Where Rollouts Begin: Low-Load, High-Leverage First-Token Diversification for RLVR cites this paper.

Where Rollouts Begin: Low-Load, High-Leverage First-Token Diversification for RLVR Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-06-29T12:03:24.400886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T11:55:13.135221Z digest=sha256:305e4809c9fff706ff71e8b4560a6026f7c87b1f71e8dababa90699dec072049

Observation 7d7824bb-7a69-4480-a017-63393c37a913 · inbound

The Tutoring Effectiveness Index: Predicting LLM Math Tutor Quality from Four Conversation Signals cites this paper.

The Tutoring Effectiveness Index: Predicting LLM Math Tutor Quality from Four Conversation Signals Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-06-29T00:12:50.001320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T00:08:57.360402Z digest=sha256:2a21f59b5a84866fa591880e899e624683818507ff875880dd9ed92e345417dd

Observation 4229a3b1-0471-4e0b-b1b2-0def05f79229 · inbound

EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning cites this paper.

EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:56:29.149115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-28T10:30:42.057301Z digest=sha256:59483f51123eec00d57a6639967d7c9d1246f8bab508a19829e2e0017b1bd166

Observation 553c98fc-cd8b-438a-a4f5-94e298da6728 · inbound

CALIBER: Calibrating Confidence Before and After Reasoning in Language Models cites this paper.

CALIBER: Calibrating Confidence Before and After Reasoning in Language Models Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:39:58.417294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-26T00:19:10.687624Z digest=sha256:fd0e1e2c19aa3f5f1859d3466ec21a50f3b7ab52a1263708ffdeb2a91c532329

Observation a8d52ec6-e740-45ba-8bee-e16094dcb35e · inbound

GRPO, Dr. GRPO, and DAPO Are Three Operations on One Number: The Group-Standard-Deviation Identity cites this paper.

GRPO, Dr. GRPO, and DAPO Are Three Operations on One Number: The Group-Standard-Deviation Identity Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:57:19.065830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-02T19:52:43.147416Z digest=sha256:64a5a02c21a5252b90486addbbe3e6f19c5374c6cfa190523d90dd91f5fe8288

Observation 6a4d4a1d-dfc8-46ab-93ae-5bb9458538c6 · inbound

Aligning Language Models with Selective Prediction cites this paper.

Aligning Language Models with Selective Prediction Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-12T01:51:25.883463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:51:25.883463Z digest=sha256:5cc715550710e8e25612913404f1d9f538343aba99d3b3327c1a8eaba53ff138

Observation b88bc4ac-058b-458a-800b-2418edf48972 · inbound

On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models cites this paper.

On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-11T20:01:06.623978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:01:06.623978Z digest=sha256:6dc6e415e531cc15716d75f3dadca56b28cbd2bf274241684cdd76e7dbb1e0bd

Observation b37ddf84-76f4-437b-9745-7ec94e741c6a · inbound

Cost of Reasoning in non-English Languages: A Case Study on Japanese cites this paper.

Cost of Reasoning in non-English Languages: A Case Study on Japanese Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-14T14:12:44.286699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T14:12:44.286699Z digest=sha256:9b6be4ab858ed98ac111f051d680b523ac6f7e10caa3a4d2945db512d15c1e93

Observation 4f6994d2-75e0-4909-a4a6-0a4901aee4ef · inbound

AngelSpec: Towards Real-World High Performance Inference with Speculative Decoding cites this paper.

AngelSpec: Towards Real-World High Performance Inference with Speculative Decoding Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-01T01:22:12.434089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T01:22:12.434089Z digest=sha256:3c868be41ef928ea8395e58e860b07f05d7f85e43217792189c9e4a9dbfa6ec4