Pith. sign in

Paper Citation Record · LEDGER

AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 29 inbound Pith citation observations for arXiv:2505.16400.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.16400 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 29 of 29 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:04:14.050518Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation fe7e1c32-886c-47e8-8499-7d151890b587 · inbound

Flow-GRPO: Training Flow Matching Models via Online RL cites this paper.

Flow-GRPO: Training Flow Matching Models via Online RL AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:45:16.862524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T18:45:16.641012Z digest=sha256:d01622ec97fdb57ae19d5a8b20573f375d920b2cabfb0a1a978382bd68d83f40

Observation 106c8006-b943-42f6-89be-79b3fc2895ba · inbound

MiMo-VL Technical Report cites this paper.

MiMo-VL Technical Report AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.050518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.050518Z digest=sha256:c3599e0c7568702991c8fa8237b4d461cbabbaa5943d130c06e7874670f4bd28

Observation c2edf7f1-ed2c-4a8f-80a6-07974803d571 · inbound

AceReason-Nemotron 1.1: Advancing Math and Code Reasoning through SFT and RL Synergy cites this paper.

AceReason-Nemotron 1.1: Advancing Math and Code Reasoning through SFT and RL Synergy AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T00:41:33.940080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:41:33.940080Z digest=sha256:31a15508cb546b6e6f224d125967107d969baec263607c6b66b1b0931f39df86

Observation 2a0e87ad-35f9-427c-8aae-3ec1f1dfbbb9 · inbound

Ring-lite: Scalable Reasoning via C3PO-Stabilized Reinforcement Learning for LLMs cites this paper.

Ring-lite: Scalable Reasoning via C3PO-Stabilized Reinforcement Learning for LLMs AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:22.519933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:22.519933Z digest=sha256:74a60a3935572bcbb014d4ab7e416753d666572f22ba6a96d7348b61cab0a439

Observation 162c4416-8a67-4ce8-9d07-a7313e31f590 · inbound

Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model cites this paper.

Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T18:57:38.595495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:57:38.595495Z digest=sha256:fb103b2ae7562104a0e9beab19f057107d85fb087f9bdf8fd101439fb584d603

Observation 6712258d-8282-40f6-ad1f-31d0bc269b16 · inbound

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks cites this paper.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T17:32:30.634578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:32:30.634578Z digest=sha256:ecfccd463ad600311c62dc948fd3e284cc16bc2bf95dbdc5374d337f9b8038c6

Observation 7ea75ee1-7bc6-480f-8cec-9b1aa1fdea48 · inbound

Stabilizing Knowledge, Promoting Reasoning: Dual-Token Constraints for RLVR cites this paper.

Stabilizing Knowledge, Promoting Reasoning: Dual-Token Constraints for RLVR AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-21T23:24:26.284917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-21T23:20:45.685446Z digest=sha256:cc4ba91cf3714b227394a7a309d6913fd9f9b23ef37df24e211b466fded28d37

Observation 9a6aba26-2b8e-416f-827a-99da7e9707bb · inbound

JT-Math: A Multi-Stage Framework for Advanced Mathematical Reasoning in Large Language Models cites this paper.

JT-Math: A Multi-Stage Framework for Advanced Mathematical Reasoning in Large Language Models AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T14:11:51.663587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:11:51.663587Z digest=sha256:93e87a27cafae166eb512ca4b518201a51bcb84c4a54ce92f23aa6a62a8e3d01

Observation 77c19d91-9dd0-49e4-bf54-bcd3b4dcf963 · inbound

PaVeRL-SQL: Text-to-SQL via Partial-Match Rewards and Verbal Reinforcement Learning cites this paper.

PaVeRL-SQL: Text-to-SQL via Partial-Match Rewards and Verbal Reinforcement Learning AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T22:48:32.672774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:48:32.672774Z digest=sha256:872d176cf07ca8d3a5c925a5ba5172b2d45ea463e32fd2d60b4a9e19a988a0e3

Observation 2500b2d5-6c02-4916-99b1-335cb24575aa · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:02:25.126142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:f36be2ab3ce53707f05a60847e1915e2775a1c07b72813afdd2780e565df9a45

Observation cdc9989a-6d34-43aa-8d34-8c94d1d57ac4 · inbound

CE-GPPO: Coordinating Entropy via Gradient-Preserving Clipping Policy Optimization in Reinforcement Learning cites this paper.

CE-GPPO: Coordinating Entropy via Gradient-Preserving Clipping Policy Optimization in Reinforcement Learning AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-18T15:11:32.270447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-18T15:06:33.991929Z digest=sha256:d45c021d2ef090c096882e4c502360a559d7c32b599e372ade1f8a4468f3ec69

Observation 3a280a8b-e968-43f2-8ec5-90bbbe690e94 · inbound

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation cites this paper.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.550417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.550417Z digest=sha256:d64b8338ce3c3671622dde79bdfa81cfe3c2f7f8b45fd1ee9bf06df0447d02f6

Observation 8b08caa5-d6d7-401f-be09-9adfd1e9b9a5 · inbound

OckBench: Measuring the Efficiency of LLM Reasoning cites this paper.

OckBench: Measuring the Efficiency of LLM Reasoning AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T23:27:36.778635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:27:36.778635Z digest=sha256:50e5d34b7c01522f79accdc045cb6204dcf3a7db4be8b54e1e7c1d3a8d081b97

Observation 051254d6-6815-46c0-8514-79096d9ba67c · inbound

Entropy Ratio Clipping as a Soft Global Constraint for Stable Reinforcement Learning cites this paper.

Entropy Ratio Clipping as a Soft Global Constraint for Stable Reinforcement Learning AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:58:46.311837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-17T00:53:51.251900Z digest=sha256:31a3f003f5a435a39768712243e03ddd62b8971006078a070542a4040f3fa30d

Observation 604cc654-1611-494a-a9f5-b122ba7c947b · inbound

Efficient Reasoning on the Edge cites this paper.

Efficient Reasoning on the Edge AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-13T23:28:12.790404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:28:12.790404Z digest=sha256:74a748e680f6ac551965addda10a9dc5d8f788c6d6c082c4a06446860acfbd7a

Observation 30ceaf37-843f-45f8-8949-c6f54325ab76 · inbound

SUPERNOVA: Eliciting General Reasoning in LLMs with Reinforcement Learning on Natural Instructions cites this paper.

SUPERNOVA: Eliciting General Reasoning in LLMs with Reinforcement Learning on Natural Instructions AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:11:01.629864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T17:16:29.695213Z digest=sha256:ec019f76425d5219341d91df6b23a15d50a553bd19fa70433a35a39f1e212e5c

Observation 1eac8f59-c2c5-4a84-aa95-f9e582fcd24c · inbound

Characterizing Model-Native Skills cites this paper.

Characterizing Model-Native Skills AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T06:06:19.468811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T05:42:49.694715Z digest=sha256:5e444318783165d21f3303f10577a822b23d072f9f6688c236529f26dde64b14

Observation 963255fd-4918-4ab8-a1be-4b5e861860c7 · inbound

Train Separately, Merge Together: Modular Post-Training with Mixture-of-Experts cites this paper.

Train Separately, Merge Together: Modular Post-Training with Mixture-of-Experts AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:56:11.305553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T05:52:28.822723Z digest=sha256:45ab74022afed3f0754f17480621475ba9eef20fe58f9fb2b7c1f71aa255bf51

Observation 993bcdb0-e451-4784-979e-d0e3b9e60613 · inbound

Understanding and Mitigating Spurious Signal Amplification in Test-Time Reinforcement Learning for Math Reasoning cites this paper.

Understanding and Mitigating Spurious Signal Amplification in Test-Time Reinforcement Learning for Math Reasoning AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T14:16:07.278352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-09T22:55:19.471307Z digest=sha256:f716b213b11660b9e832222ae657dab84304faf80071a390f3c17cf7ef13628b

Observation edc3109a-db52-4c8f-b539-c362704567e8 · inbound

Delay, Plateau, or Collapse: Evaluating the Impact of Systematic Verification Error on RLVR cites this paper.

Delay, Plateau, or Collapse: Evaluating the Impact of Systematic Verification Error on RLVR AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:45:53.101099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T18:52:52.969408Z digest=sha256:19c3b935ab7d42a1e49ca08e3b35e445c4a43f4821ad994f72567f1fa83de4ec

Observation 8cc6f019-6d8f-482f-924f-841fba0a9cdf · inbound

Towards Order Fairness: Mitigating LLMs Order Sensitivity through Dual Group Advantage Optimization cites this paper.

Towards Order Fairness: Mitigating LLMs Order Sensitivity through Dual Group Advantage Optimization AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T07:37:29.841604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-13T07:32:58.404947Z digest=sha256:c0cc02f654a04a35b01d3bbdc842096c832eab51e0697716e890c0c8072f9006

Observation 2c103f40-2966-4ba2-9c42-93d845d2059d · inbound

Scalable Token-Level Hallucination Detection in Large Language Models cites this paper.

Scalable Token-Level Hallucination Detection in Large Language Models AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:52:22.421954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T05:49:23.534294Z digest=sha256:1489eda11514c14197cdba3be233175368013fb78c1ba154f5ac4025df5af0a7

Observation 4e0abeda-e4b8-4968-8653-cb0b3fd38774 · inbound

Respecting Self-Uncertainty in On-Policy Self-Distillation for Efficient LLM Reasoning cites this paper.

Respecting Self-Uncertainty in On-Policy Self-Distillation for Efficient LLM Reasoning AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:52:52.411736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T19:51:17.595707Z digest=sha256:356f90156ce826759e944d4642aa9b983b6f80db83d3849ab5770eae33db8c5b

Observation 4d9f336e-c8ca-4339-b799-afc4c8a24a24 · inbound

When Self-Belief Misleads: Active Label Acquisition for Reinforcement Learning with Verifiable Rewards cites this paper.

When Self-Belief Misleads: Active Label Acquisition for Reinforcement Learning with Verifiable Rewards AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-29T23:04:01.322678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T22:58:46.313028Z digest=sha256:318ef522f2799b1779a14f996ed7424dcee5d7369b90f786c08f99ebfa776633

Observation 49c626dd-add2-429b-a250-05b5da73c1f4 · inbound

RLVR Datasets and Where to Find Them: Tracing Data Lineage for Better Training Data cites this paper.

RLVR Datasets and Where to Find Them: Tracing Data Lineage for Better Training Data AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-06-29T19:43:55.053657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-29T19:34:12.081362Z digest=sha256:332d316e34c8d3c6762dcae04171e74ab11fc9f79b55ed4683039bd2ca67a550

Observation cae54978-a304-4bb7-9065-517d124851da · inbound

From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning cites this paper.

From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning

Reference 74

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T20:38:55.942362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T01:13:11.483599Z digest=sha256:d4529478fb52839d78c31988fa272b6fa35284f4abed3b7dc4bb5c3a59b5a67f

Observation f080444e-7265-48f3-befc-cf754f9647fc · inbound

Beyond Entropy: Learning from Token-Level Distributional Deviations for LLM Reasoning cites this paper.

Beyond Entropy: Learning from Token-Level Distributional Deviations for LLM Reasoning AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:49:30.159779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T17:37:43.856000Z digest=sha256:14dde44e33e5d1e147c8cca861cc8c5e9d33189bb116390bf78cfd2ebb9ebfaa

Observation 4afb7a0a-b48a-4c0a-ab97-a7d8e16d28ed · inbound

What are Key Factors for Updates in RL for LLM Reasoning? cites this paper.

What are Key Factors for Updates in RL for LLM Reasoning? AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T09:09:43.581253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-26T10:24:53.245739Z digest=sha256:c464385b51fc77f7654d7ec31801f5d8f7644b2e3ae6ce6b8030cf9ef0ad9ba7

Observation 70ba9640-73f3-4a19-bd19-009a3d40a001 · inbound

Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning cites this paper.

Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T06:37:36.445162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T06:37:36.445162Z digest=sha256:e05589418b2d3c6ae7f00cd99bf7adc0846c8c9324aeaad693ab0432a667db23