Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 29 inbound Pith citation observations for arXiv:2505.16400.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:04:14.050518Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation fe7e1c32-886c-47e8-8499-7d151890b587 · inbound
Flow-GRPO: Training Flow Matching Models via Online RL AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 106c8006-b943-42f6-89be-79b3fc2895ba · inbound
MiMo-VL Technical Report AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2edf7f1-ed2c-4a8f-80a6-07974803d571 · inbound
AceReason-Nemotron 1.1: Advancing Math and Code Reasoning through SFT and RL Synergy AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a0e87ad-35f9-427c-8aae-3ec1f1dfbbb9 · inbound
Ring-lite: Scalable Reasoning via C3PO-Stabilized Reinforcement Learning for LLMs AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 162c4416-8a67-4ce8-9d07-a7313e31f590 · inbound
Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6712258d-8282-40f6-ad1f-31d0bc269b16 · inbound
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ea75ee1-7bc6-480f-8cec-9b1aa1fdea48 · inbound
Stabilizing Knowledge, Promoting Reasoning: Dual-Token Constraints for RLVR AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9a6aba26-2b8e-416f-827a-99da7e9707bb · inbound
JT-Math: A Multi-Stage Framework for Advanced Mathematical Reasoning in Large Language Models AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77c19d91-9dd0-49e4-bf54-bcd3b4dcf963 · inbound
PaVeRL-SQL: Text-to-SQL via Partial-Match Rewards and Verbal Reinforcement Learning AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2500b2d5-6c02-4916-99b1-335cb24575aa · inbound
A Survey of Reinforcement Learning for Large Reasoning Models AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cdc9989a-6d34-43aa-8d34-8c94d1d57ac4 · inbound
CE-GPPO: Coordinating Entropy via Gradient-Preserving Clipping Policy Optimization in Reinforcement Learning AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3a280a8b-e968-43f2-8ec5-90bbbe690e94 · inbound
MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b08caa5-d6d7-401f-be09-9adfd1e9b9a5 · inbound
OckBench: Measuring the Efficiency of LLM Reasoning AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 051254d6-6815-46c0-8514-79096d9ba67c · inbound
Entropy Ratio Clipping as a Soft Global Constraint for Stable Reinforcement Learning AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 604cc654-1611-494a-a9f5-b122ba7c947b · inbound
Efficient Reasoning on the Edge AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30ceaf37-843f-45f8-8949-c6f54325ab76 · inbound
SUPERNOVA: Eliciting General Reasoning in LLMs with Reinforcement Learning on Natural Instructions AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1eac8f59-c2c5-4a84-aa95-f9e582fcd24c · inbound
Characterizing Model-Native Skills AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 963255fd-4918-4ab8-a1be-4b5e861860c7 · inbound
Train Separately, Merge Together: Modular Post-Training with Mixture-of-Experts AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 993bcdb0-e451-4784-979e-d0e3b9e60613 · inbound
Understanding and Mitigating Spurious Signal Amplification in Test-Time Reinforcement Learning for Math Reasoning AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation edc3109a-db52-4c8f-b539-c362704567e8 · inbound
Delay, Plateau, or Collapse: Evaluating the Impact of Systematic Verification Error on RLVR AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8cc6f019-6d8f-482f-924f-841fba0a9cdf · inbound
Towards Order Fairness: Mitigating LLMs Order Sensitivity through Dual Group Advantage Optimization AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2c103f40-2966-4ba2-9c42-93d845d2059d · inbound
Scalable Token-Level Hallucination Detection in Large Language Models AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4e0abeda-e4b8-4968-8653-cb0b3fd38774 · inbound
Respecting Self-Uncertainty in On-Policy Self-Distillation for Efficient LLM Reasoning AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4d9f336e-c8ca-4339-b799-afc4c8a24a24 · inbound
When Self-Belief Misleads: Active Label Acquisition for Reinforcement Learning with Verifiable Rewards AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 49c626dd-add2-429b-a250-05b5da73c1f4 · inbound
RLVR Datasets and Where to Find Them: Tracing Data Lineage for Better Training Data AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cae54978-a304-4bb7-9065-517d124851da · inbound
From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f080444e-7265-48f3-befc-cf754f9647fc · inbound
Beyond Entropy: Learning from Token-Level Distributional Deviations for LLM Reasoning AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4afb7a0a-b48a-4c0a-ab97-a7d8e16d28ed · inbound
What are Key Factors for Updates in RL for LLM Reasoning? AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 70ba9640-73f3-4a19-bd19-009a3d40a001 · inbound
Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.