Pith. sign in

Paper Citation Record · LEDGER

ACRL: Adaptive Control of Training-Inference Discrepancy for Stable Reinforcement Learning

As of 14 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 0 inbound Pith citation observations for arXiv:2607.24062.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.24062 v1

Coverage vector

measured 15 of 15 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-31T23:13:14.375921Z

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

15 of 15 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved14
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ec4ad778-ee12-4ae2-b4db-51d16583f4f9 · outbound

This paper cites 20250910.

ACRL: Adaptive Control of Training-Inference Discrepancy for Stable Reinforcement Learning 20250910

Reference 4

Resolution
malformed identifier
no resolver link, observed 2026-07-31T23:13:13.608682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:13:13.608682Z digest=sha256:4deb7207cfd7c3791fce295cd27432da43108dad5f8d83c2c566a8d0ad869a75

Observation b4efd088-3dff-45ee-b721-f54c8c9608af · outbound

This paper cites Gemma 3 Technical Report.

ACRL: Adaptive Control of Training-Inference Discrepancy for Stable Reinforcement Learning Gemma 3 Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-31T23:13:13.668099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:13:13.668099Z digest=sha256:63daacb0a423dcc6a40d0af30b6c7b2a02bf3b3a5e09024cce629849c5d98a23

Observation 83cd4246-c8ff-4614-a9be-7bd40368e28b · outbound

This paper cites Accessed: 2025-12-.

ACRL: Adaptive Control of Training-Inference Discrepancy for Stable Reinforcement Learning Accessed: 2025-12-

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-31T23:13:13.702195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:13:13.702195Z digest=sha256:e6629283d6bb919a2039e8b06f9e0493fa251200893e876bf97f05e3ecedadfb

Observation f3184fe9-07d5-42f7-b163-a7fd134e9fbb · outbound

This paper cites S., and Lin, M.

ACRL: Adaptive Control of Training-Inference Discrepancy for Stable Reinforcement Learning S., and Lin, M

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-31T23:13:13.839022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:13:13.839022Z digest=sha256:1e5d67a5f1f8d98f470ee0945ef6d1ccb8737f0a9b1b07f478a8940742ae77a1

Observation 0b1aa996-826f-4e23-a6c6-47053801ddb4 · outbound

This paper cites Proximal Policy Optimization Algorithms.

ACRL: Adaptive Control of Training-Inference Discrepancy for Stable Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-31T23:13:13.924241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:13:13.924241Z digest=sha256:e9745dcea6785b99bc51e7e3613c02f03ba814ccfcd0b667e6d7ac6e585769e8

Observation 473d0a73-b70c-4606-9fe8-7154193eec71 · outbound

This paper cites Klear-reasoner: Advanc- ing reasoning capability via gradient-preserving clipping policy optimization.arXiv preprint arXiv:2508.07629, 2025a.

ACRL: Adaptive Control of Training-Inference Discrepancy for Stable Reinforcement Learning Klear-reasoner: Advanc- ing reasoning capability via gradient-preserving clipping policy optimization.arXiv preprint arXiv:2508.07629, 2025a

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-31T23:13:14.071857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:13:14.071857Z digest=sha256:66019003ff7d58ae10d13c66e0f3785d9b6953e6be718c13316a6e9a129d6eb7

Observation c94e3fd4-8af7-4eb4-b8ec-2d73af51c093 · outbound

This paper cites Qwen3 Technical Report.

ACRL: Adaptive Control of Training-Inference Discrepancy for Stable Reinforcement Learning Qwen3 Technical Report

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-31T23:13:14.145390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:13:14.145390Z digest=sha256:179b22315cd2f304ca51f23b6a5c55f5eb19785193a51c228709c11f81583b46

Observation fad91dce-ee09-47e0-8d6b-be8fa7d98e43 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

ACRL: Adaptive Control of Training-Inference Discrepancy for Stable Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-31T23:13:14.222985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:13:14.222985Z digest=sha256:aae7d4ecfb188c901c669b347b4e649c081e699bcdc5c4a15ef632b015a298f1

Observation 72e3c3ce-ef91-431a-a99a-f97d4ba0215f · outbound

This paper cites Group Sequence Policy Optimization.

ACRL: Adaptive Control of Training-Inference Discrepancy for Stable Reinforcement Learning Group Sequence Policy Optimization

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-31T23:13:14.307024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:13:14.307024Z digest=sha256:b5538578465dbd2acd25f8f5180f8fbae0588b8bc619bd0a7e8b6242ad953ba8

Observation 8071c120-4352-4e08-b569-89463d71fdb8 · outbound

This paper cites Hyper-parameters used for experiments training.

ACRL: Adaptive Control of Training-Inference Discrepancy for Stable Reinforcement Learning Hyper-parameters used for experiments training

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-31T23:13:14.375921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:13:14.375921Z digest=sha256:3c12a7aecaad787af5bc38c7de81dafac35e9ca5b840d4943d6da587f102a40c

Observation 35ce8b5d-d972-4296-9c31-5c3f40a2dbc8 · outbound

This paper cites Recipes for Pre-training LLMs with MXFP8.

ACRL: Adaptive Control of Training-Inference Discrepancy for Stable Reinforcement Learning Recipes for Pre-training LLMs with MXFP8

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-31T23:13:13.763909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:13:13.763909Z digest=sha256:058e33f376873f93c36bcf6f68ba76e46010914fed4f4d03d3582467ee3decb2

Observation 7a333f25-2142-4057-a7e8-0798944cba1c · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

ACRL: Adaptive Control of Training-Inference Discrepancy for Stable Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-07-31T23:13:13.990903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:13:13.990903Z digest=sha256:d74e0d2036be9024366396475b0dcb8d58f27b46ebd38b9b8aab09c1bd7ba73f

Observation df6fa51d-b4a9-4b64-b654-5698894129eb · outbound

This paper cites MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention.

ACRL: Adaptive Control of Training-Inference Discrepancy for Stable Reinforcement Learning MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-07-31T23:13:13.419211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:13:13.419211Z digest=sha256:d76d64d0589d1104a079520a90de697bf840a29ff4026969d803c41acec1b773

Observation 988833e9-f7d6-4b85-9772-714e6920cdae · outbound

This paper cites Intrinsically Interpretable Attention via Sparse Post-Training.

ACRL: Adaptive Control of Training-Inference Discrepancy for Stable Reinforcement Learning Intrinsically Interpretable Attention via Sparse Post-Training

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-07-31T23:13:13.542369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:13:13.542369Z digest=sha256:fc70e79746d1d40cc917a61d5ed8602076d37d014410a624af770fe5c750a878

Observation be703209-73b9-41d0-ad95-efbea120787d · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

ACRL: Adaptive Control of Training-Inference Discrepancy for Stable Reinforcement Learning Training Verifiers to Solve Math Word Problems

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-07-31T23:13:13.476505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:13:13.476505Z digest=sha256:e9ce7c8f12eee95d0e53ed8fb902d437566816b7a80ec034a75718e6cc325beb

Pith citing papers

No inbound Pith citation observations are available.