Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-16T10:31:04.728005Z
Paper Citation Record · LEDGER
As of 4 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 70 inbound Pith citation observations for arXiv:2504.11456.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-16T10:31:04.728005Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T22:56:02.509530Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-09T03:25:57.840663Z
22 of 22 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 26485edc-3072-4b24-925c-5d34a4ec5208 · outbound
DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b931d842-000d-4550-aadb-f6fe85a48a11 · outbound
DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1689582a-6927-475c-8f48-ed77d2dfcab2 · outbound
DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning doi: 10.18653/v1/2023.emnlp-main.468
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2d5c5678-3006-4d48-a05d-14ba7ab3a382 · outbound
DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d4bb8b0b-0c52-479f-810f-09d643e3c0e6 · outbound
DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning Training Verifiers to Solve Math Word Problems
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 46dd5051-9a31-473e-b041-bdf3603445f9 · outbound
DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning Process Reinforcement through Implicit Rewards
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation adfc2564-27b3-4004-9acd-8578b2f4171b · outbound
DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning Reinforcement learning for reasoning in small llms: What works and what doesn’t
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d7fd0904-dc6b-4b76-8db5-42c463690aca · outbound
DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning MathOdyssey: Benchmarking Mathematical Problem-Solving Skills in Large Language Models Using Odyssey Math Data
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2033577b-b2b8-4cda-a571-46374563bd24 · outbound
DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 85cff00c-7d90-46d5-a7ff-42656da1cbae · outbound
DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 995d756a-08af-4938-8ce5-992406dc104c · outbound
DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 630c41d4-ba8c-462e-826c-c081c0fff7ab · outbound
DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning OlympicArena: Benchmarking Multi-discipline Cognitive Reasoning for Superintelligent AI
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 447177ca-6f1b-4e07-a775-d9c9af612eb1 · outbound
DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning Unresolved cited work
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d5894fd5-b3aa-497a-bb74-22d163c92241 · outbound
DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning Augmenting Math Word Problems via Iterative Question Composing
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation da5f633b-de98-4e75-bf85-fc411aeb19d8 · outbound
DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning AIMO-2 Winning Solution: Building State-of-the-Art Mathematical Reasoning Models with OpenMathReasoning dataset
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e3eba730-9185-4187-8c01-ad0fe37424e1 · outbound
DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning Sentence-bert: Sentence embeddings using siamese bert-networks
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8a257143-90ed-4c71-aeae-4cc1c5ca0954 · outbound
DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8564cbcf-f582-49d1-a4be-81653f58aff0 · outbound
DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5e1f518d-fd3c-4fe2-a780-f461cb2b8e55 · outbound
DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning OpenMathInstruct-2: Accelerating AI for Math with Massive Open-Source Instruction Data
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9cc0c66f-1531-4f89-8ac8-883264aa5fc7 · outbound
DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 629dd6a8-b5e8-4559-b592-026a21ba435b · outbound
DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ae8240d8-0d89-450b-876c-8366c45afb75 · outbound
DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning MegaMath: Pushing the Limits of Open Math Corpora
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 47bbb0d6-6ebb-4c6e-a885-d781984ddf9a · inbound
OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 058bade5-4ca3-4d40-aae4-2ec52444afe0 · inbound
Beyond Two-Stage Training: Cooperative SFT and RL for LLM Reasoning DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a60124cb-8aa4-426d-9994-f1153aa919bd · inbound
A Survey of Reinforcement Learning for Large Reasoning Models DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 193
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 723679c4-7e2d-4f29-aa74-cb04aaf5aadf · inbound
Plan Then Action:High-Level Planning Guidance Reinforcement Learning for LLM Reasoning DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation efedbebc-b4da-49ff-8023-bf8b809b4272 · inbound
Probing the Difficulty Perception Mechanism of Large Language Models DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 967a59f4-c332-409d-a416-48d6fa20ece1 · inbound
Rewarding Structural Conformance of Reasoning using Process Mining DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7eff8bab-4ce0-4f8b-b813-c430af938f21 · inbound
RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e0f3afa-9e22-4072-a047-c2e47d2f7927 · inbound
SPHINX: A Synthetic Environment for Visual Perception and Reasoning DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c73c02c1-8d22-4ee2-bebf-4d4972f38c20 · inbound
SCALER:Synthetic Scalable Adaptive Learning Environment for Reasoning DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation df99729a-bc3a-4ae8-9baa-f3fa3296fdb8 · inbound
Conversation for Non-verifiable Learning: Self-Evolving LLMs through Meta-Evaluation DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation cf38665a-e33a-400c-b7a5-84870c2e7f42 · inbound
On the Overscaling Curse of Parallel Thinking: System Efficacy Contradicts Sample Efficiency DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 860ebc19-c46d-4c0c-8fb0-6d1bd6a6f7f8 · inbound
GradAlign: Gradient-Aligned Data Selection for LLM Reinforcement Learning DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9d9e73e-2c42-43b6-88a5-c55fc99f6445 · inbound
Compress the Easy, Explore the Hard: Difficulty-Aware Entropy Regularization for Efficient LLM Reasoning DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 690dc376-9abc-4168-8dd7-8c8e5a7cf793 · inbound
Your Model Diversity, Not Method, Determines Reasoning Strategy DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 07c5dbba-d228-4d05-aeae-fa4ae0032d28 · inbound
PubSwap: Public-Data Off-Policy Coordination for Federated RLVR DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c4d001c6-05c1-4863-a9cb-572442a58a36 · inbound
Rethinking the Comparison Unit in Sequence-Level Reinforcement Learning: An Equal-Length Paired Training Framework from Loss Correction to Sample Construction DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 94fced4b-444c-4351-a118-0b4e0b12381f · inbound
Learning from Less: Measuring the Effectiveness of RLVR in Low Data and Compute Regimes DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0d3b114b-7072-4f2c-83fe-8eb9bab0386a · inbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8997fca0-e118-4973-a178-c56d88e2a3c4 · inbound
Length Value Model: Scalable Value Pretraining for Token-Level Length Modeling DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3e7b28f8-4b77-43c5-b8f0-cb93f6526045 · inbound
Length Value Model: Scalable Value Pretraining for Token-Level Length Modeling DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84e34f6c-4458-468d-9d27-fca1959a941f · inbound
Internalizing Outcome Supervision into Process Supervision: A New Paradigm for Reinforcement Learning for Reasoning DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation bfebc584-9595-40c8-a585-a60125dd2e88 · inbound
Asymmetric On-Policy Distillation: Bridging Exploitation and Imitation at the Token Level DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 32a3b2f2-5629-4d01-ab47-dc53fee74265 · inbound
Asymmetric On-Policy Distillation: Bridging Exploitation and Imitation at the Token Level DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0515b3fd-630d-4969-bf5d-2338303bdbb1 · inbound
Asymmetric On-Policy Distillation: Bridging Exploitation and Imitation at the Token Level DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation efa33de1-b97f-4433-a238-25112f42335b · inbound
On the Implicit Reward Overfitting and the Low-rank Dynamics in RLVR DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 685b4bac-0184-4bf8-99b8-1ceab8937d93 · inbound
AIPO: Learning to Reason from Active Interaction DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 47286ca3-2199-4373-9bf6-b33efca7cdb3 · inbound
AIPO: Learning to Reason from Active Interaction DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b35cf74e-1a9d-4ee5-ada3-20189188e3c4 · inbound
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 894a6985-5629-4100-99a9-42a827d32dfa · inbound
Training with Harnesses: On-Policy Harness Self-Distillation for Complex Reasoning DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7472f34e-0af5-472c-885f-591472bf5d2a · inbound
Forge: Quality-Aware Reinforcement Learning for NP-Hard Optimization in LLMs DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 22d33486-cbbf-4c6a-95c9-20ae25de4097 · inbound
DARE: Difficulty-Adaptive Reinforcement Learning with Co-Evolved Difficulty Estimation DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e0eb2663-a354-47ec-91a2-097eb8ad1ec1 · inbound
Nice Fold or Hero Call: Learning Budget-Efficient Thinking for Adaptive Reasoning DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 63e4046a-1b0f-43a9-99ce-44fb76fbc868 · inbound
Learning to Foresee: Unveiling the Unlocking Efficiency of On-Policy Distillation DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 23eeed74-71e7-412f-b5eb-958a2afb8867 · inbound
Pion: A Spectrum-Preserving Optimizer via Orthogonal Equivalence Transformation DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ad9b5f8d-ce88-4956-b600-102f09a6f772 · inbound
Multi-Rollout On-Policy Distillation via Peer Successes and Failures DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 550971ca-a461-4aae-96f8-1821c3e2c580 · inbound
Taming the Thinker: Conditional Entropy Shaping for Adaptive LLM Reasoning DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e4b2e5c6-750a-41d0-b97e-86ef77a03cc4 · inbound
Beyond Mode Collapse: Distribution Matching for Diverse Reasoning DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 971dcc5f-8264-42b0-8b6c-464b61374a60 · inbound
DEL: Digit Entropy Loss for Numerical Learning of Large Language Models DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation bd7ea9ff-da6f-420d-bdd9-6ee31bad6352 · inbound
DelTA: Discriminative Token Credit Assignment for Reinforcement Learning from Verifiable Rewards DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8bb3ca51-c57b-4648-895e-466727ea8ea2 · inbound
LANG: Reinforcement Learning for Multilingual Reasoning with Language-Adaptive Hint Guidance DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e880cc08-4db1-45fc-9a62-cd7c46fb811b · inbound
Not only where, But when: Temporal Scheduling for RLVR DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ea838eab-8e97-476f-8d80-a2736669e7f2 · inbound
Spend Your Rollouts Where It Counts: Rollout Allocation for Group-Based RL Post-Training DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation bda730de-94b9-4c7c-8f93-9450c7624124 · inbound
RLVR Datasets and Where to Find Them: Tracing Data Lineage for Better Training Data DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9bee5b38-0f96-4678-a3e3-9faac621eb5c · inbound
DenoiseRL: Bootstrapping Reasoning Models to Recover from Noisy Prefixes DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d7258fbd-1837-4aa7-8859-85c5e3b47204 · inbound
DenoiseRL: Bootstrapping Reasoning Models to Recover from Noisy Prefixes DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a769edab-07f2-4b0d-9d8a-e2a919c54835 · inbound
Trust Region On-Policy Distillation DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 287
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a757bd81-2922-43b1-bea3-55363c01d36a · inbound
Filter, Then Reweight: Rethinking Optimization Granularity in On-Policy Distillation DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7fb7b005-ff50-4cc2-a597-31b394148e56 · inbound
ThoughtFold: Folding Reasoning Chains via Introspective Preference Learning DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c980f253-16a5-419b-9818-56453ac27ce8 · inbound
GeoMin: Data-Efficient Semi-Supervised RLVR via Geometric Distribution Modeling DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9a80c423-7e11-43b9-8251-bb8b81bd5859 · inbound
GRAIL: Gradient-Reweighted Advantages for Reinforcement Learning with Verifiable Rewards DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b425e865-08fb-4d09-aafc-324d5546c6da · inbound
State commitment learning: training language models to distinguish computation from memory DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 457f3a84-162b-4381-8e5e-f7a65fda66f9 · inbound
ConSteer-RL: Steering Reasoning Capabilities in Large Language Models via Confidence-Aware Reinforcement Learning DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8249fa7b-1492-4cc3-a41a-584ac4c82c17 · inbound
SG-OPD: Sign-Gated On-Policy Distillation via Sign-Consistency Gating and Phased Teacher Sampling DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8c83603f-2553-43f7-9c8a-8d0ba9286330 · inbound
RLCSD: Reinforcement Learning with Contrastive On-Policy Self-Distillation DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 365dd4e4-44fb-4a6a-980e-69e8631d0689 · inbound
Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 74c5ef6d-7dda-442c-b49a-f823e2885d23 · inbound
Finding the Evidence: Discovering Decision-Supporting Tokens for On-Policy Reasoning Distillation DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1cc6cde3-6c36-4222-87cd-4a0aa0896f7b · inbound
AsyncOPD: How Stale Can On-Policy Distillation Be? DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2832645c-535e-4fe9-8480-08a3aa464086 · inbound
EntroRouter: Learning Efficient Model Routing via Entropy Regulation DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ea29cd36-3041-4a6c-aa58-ad151fe1e294 · inbound
The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0cd46127-b12a-4cf3-9c6a-f76ce529268f · inbound
Beyond Compilation: Evaluating Faithful Natural-Language-to-Lean Statement Formalization DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f640058f-4066-40c0-b8b7-8541b8dd9730 · inbound
On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78c2130d-9121-4dd8-9f88-afced1b0401f · inbound
Weak-to-Strong Generalization via Direct On-Policy Distillation DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6cf49e12-3bfb-4301-9504-f6cef34cf187 · inbound
Weak-to-Strong Generalization via Direct On-Policy Distillation DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2dea5867-8b5f-49dc-9414-306428d24929 · inbound
Max Out GRPO Signal: Adaptive Trace Prefix Control for Hard Reasoning Problems DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8b390d5b-bf7a-42a2-b48e-ce75570111dc · inbound
Agon: Competitive Cross-Model RL with Implicit Rival Grading of Reasoning DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8c12d294-4409-4f4e-a72f-2f576ca10288 · inbound
Proxy Exploration and Reusable Guidance: A Modular LLM Post-Training Paradigm via Proxy-Guided Update Signals DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b5f5768-87a2-432a-88c7-92682d12d1ef · inbound
Cross-Tokenizer On-Policy Distillation via Byte-Prefix Marginalization DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c36225b-6de3-4a63-8363-09d4bb8883ad · inbound
Weak-to-Strong On-Policy Distillation DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72fd49f8-e1e1-40ae-961e-20159b36095b · inbound
LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b63273ff-1c3c-404e-80a4-b52983f59eaf · inbound
ThinkReset: Learnable Intermediate Interface Construction for Bounded-Context Long-Horizon Reasoning DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.