Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-31T23:13:14.375921Z
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 0 inbound Pith citation observations for arXiv:2607.24062.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-31T23:13:14.375921Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
15 of 15 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation ec4ad778-ee12-4ae2-b4db-51d16583f4f9 · outbound
ACRL: Adaptive Control of Training-Inference Discrepancy for Stable Reinforcement Learning 20250910
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4efd088-3dff-45ee-b721-f54c8c9608af · outbound
ACRL: Adaptive Control of Training-Inference Discrepancy for Stable Reinforcement Learning Gemma 3 Technical Report
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83cd4246-c8ff-4614-a9be-7bd40368e28b · outbound
ACRL: Adaptive Control of Training-Inference Discrepancy for Stable Reinforcement Learning Accessed: 2025-12-
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3184fe9-07d5-42f7-b163-a7fd134e9fbb · outbound
ACRL: Adaptive Control of Training-Inference Discrepancy for Stable Reinforcement Learning S., and Lin, M
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b1aa996-826f-4e23-a6c6-47053801ddb4 · outbound
ACRL: Adaptive Control of Training-Inference Discrepancy for Stable Reinforcement Learning Proximal Policy Optimization Algorithms
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 473d0a73-b70c-4606-9fe8-7154193eec71 · outbound
ACRL: Adaptive Control of Training-Inference Discrepancy for Stable Reinforcement Learning Klear-reasoner: Advanc- ing reasoning capability via gradient-preserving clipping policy optimization.arXiv preprint arXiv:2508.07629, 2025a
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c94e3fd4-8af7-4eb4-b8ec-2d73af51c093 · outbound
ACRL: Adaptive Control of Training-Inference Discrepancy for Stable Reinforcement Learning Qwen3 Technical Report
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fad91dce-ee09-47e0-8d6b-be8fa7d98e43 · outbound
ACRL: Adaptive Control of Training-Inference Discrepancy for Stable Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72e3c3ce-ef91-431a-a99a-f97d4ba0215f · outbound
ACRL: Adaptive Control of Training-Inference Discrepancy for Stable Reinforcement Learning Group Sequence Policy Optimization
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8071c120-4352-4e08-b569-89463d71fdb8 · outbound
ACRL: Adaptive Control of Training-Inference Discrepancy for Stable Reinforcement Learning Hyper-parameters used for experiments training
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35ce8b5d-d972-4296-9c31-5c3f40a2dbc8 · outbound
ACRL: Adaptive Control of Training-Inference Discrepancy for Stable Reinforcement Learning Recipes for Pre-training LLMs with MXFP8
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a333f25-2142-4057-a7e8-0798944cba1c · outbound
ACRL: Adaptive Control of Training-Inference Discrepancy for Stable Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df6fa51d-b4a9-4b64-b654-5698894129eb · outbound
ACRL: Adaptive Control of Training-Inference Discrepancy for Stable Reinforcement Learning MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 988833e9-f7d6-4b85-9772-714e6920cdae · outbound
ACRL: Adaptive Control of Training-Inference Discrepancy for Stable Reinforcement Learning Intrinsically Interpretable Attention via Sparse Post-Training
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be703209-73b9-41d0-ad95-efbea120787d · outbound
ACRL: Adaptive Control of Training-Inference Discrepancy for Stable Reinforcement Learning Training Verifiers to Solve Math Word Problems
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.