Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-11T14:30:20.959431Z
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2607.04728.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-11T14:30:20.959431Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
31 of 31 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation a06f6d39-a2dd-4355-ac07-fd68f37a7eb4 · outbound
Turning Off-Policy Tokens On-Policy: A Plug-in Approach for Improving LLM Alignment MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c1e416c-a0c4-4052-a3bd-b5380f63b673 · outbound
Turning Off-Policy Tokens On-Policy: A Plug-in Approach for Improving LLM Alignment The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 273e1317-374f-4713-9a8f-598c35730718 · outbound
Turning Off-Policy Tokens On-Policy: A Plug-in Approach for Improving LLM Alignment Soft Adaptive Policy Optimization
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5893bd3b-9568-4881-ab4e-d9ba90796236 · outbound
Turning Off-Policy Tokens On-Policy: A Plug-in Approach for Improving LLM Alignment DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 401e2898-b4b6-464c-8117-91fcc19a8df3 · outbound
Turning Off-Policy Tokens On-Policy: A Plug-in Approach for Improving LLM Alignment CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7449652-6eb0-425e-845c-9361c74c850b · outbound
Turning Off-Policy Tokens On-Policy: A Plug-in Approach for Improving LLM Alignment Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18a73839-087e-4128-9a07-439a972ec5c8 · outbound
Turning Off-Policy Tokens On-Policy: A Plug-in Approach for Improving LLM Alignment Dense passage retrieval for open-domain question answering
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9284340e-1586-48fc-89bd-4f886c07094d · outbound
Turning Off-Policy Tokens On-Policy: A Plug-in Approach for Improving LLM Alignment A step back: Prefix importance ratio stabilizes policy optimization
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 654a09ac-c475-40c3-9571-0ff876ad20da · outbound
Turning Off-Policy Tokens On-Policy: A Plug-in Approach for Improving LLM Alignment Trust Region Masking for Long-Horizon LLM Reinforcement Learning
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2187083-7152-45e6-b97d-b56bcbd30d48 · outbound
Turning Off-Policy Tokens On-Policy: A Plug-in Approach for Improving LLM Alignment Reasoning and Tool-use Compete in Agentic RL:From Quantifying Interference to Disentangled Tuning
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8da5dd53-26a6-43f2-9975-071ed232866d · outbound
Turning Off-Policy Tokens On-Policy: A Plug-in Approach for Improving LLM Alignment Sparse-rl: Breaking the memory wall in llm reinforcement learning via stable sparse rollouts
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 966550d4-8039-4999-9ce3-377c0d88fc17 · outbound
Turning Off-Policy Tokens On-Policy: A Plug-in Approach for Improving LLM Alignment Stabilizing moe reinforcement learning by aligning training and inference routers
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14d476a9-6048-4777-8f0c-4d58f9306b4d · outbound
Turning Off-Policy Tokens On-Policy: A Plug-in Approach for Improving LLM Alignment When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3493a053-8c96-43b8-b22b-37a0f47598f3 · outbound
Turning Off-Policy Tokens On-Policy: A Plug-in Approach for Improving LLM Alignment Measuring and narrowing the compositionality gap in language models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f545beb-a004-4881-8812-af5042f968b1 · outbound
Turning Off-Policy Tokens On-Policy: A Plug-in Approach for Improving LLM Alignment Proximal Policy Optimization Algorithms
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a0481f4-1ce8-43fe-b90b-75ad9b522765 · outbound
Turning Off-Policy Tokens On-Policy: A Plug-in Approach for Improving LLM Alignment DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5313c16c-70fc-4fa9-9043-601d12e9eda1 · outbound
Turning Off-Policy Tokens On-Policy: A Plug-in Approach for Improving LLM Alignment Maximum likelihood reinforcement learning
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38afff03-0a3d-44dc-bed8-9dc255baf7c6 · outbound
Turning Off-Policy Tokens On-Policy: A Plug-in Approach for Improving LLM Alignment Every step evolves: Scaling reinforcement learning for trillion-scale thinking model
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c595e6b-63bb-4269-b8e1-862c8c53534b · outbound
Turning Off-Policy Tokens On-Policy: A Plug-in Approach for Improving LLM Alignment UCOB: Learning to Utilize and Evolve Agentic Skills via Credit-Aware On-Policy Bidirectional Self-Distillation
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 919c82d0-d2ae-4405-a3b5-08700456d47f · outbound
Turning Off-Policy Tokens On-Policy: A Plug-in Approach for Improving LLM Alignment Text Embeddings by Weakly-Supervised Contrastive Pre-training
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab5248c1-5f39-4517-8972-f24a69e3cd93 · outbound
Turning Off-Policy Tokens On-Policy: A Plug-in Approach for Improving LLM Alignment Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b01057bc-de5b-4b1e-a26e-000787aa8b3b · outbound
Turning Off-Policy Tokens On-Policy: A Plug-in Approach for Improving LLM Alignment Qwen3 Technical Report
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation acdb543a-83ee-4fcb-99bc-792c49b9609a · outbound
Turning Off-Policy Tokens On-Policy: A Plug-in Approach for Improving LLM Alignment DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0884fa46-92f7-4a1a-8e31-ddf7993ad560 · outbound
Turning Off-Policy Tokens On-Policy: A Plug-in Approach for Improving LLM Alignment HTAM: Hierarchical Transition-Attended Memory for Operator Optimization
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e61a3fb6-412d-4014-b12c-9fc26347f03b · outbound
Turning Off-Policy Tokens On-Policy: A Plug-in Approach for Improving LLM Alignment PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4dee2c79-c12f-4200-8af0-9f02725ea7f1 · outbound
Turning Off-Policy Tokens On-Policy: A Plug-in Approach for Improving LLM Alignment Stabilizing reinforcement learning with llms: For- mulation and practices
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cafb0fc6-8c8c-44af-8bce-8e71a9cdd387 · outbound
Turning Off-Policy Tokens On-Policy: A Plug-in Approach for Improving LLM Alignment Unresolved cited work
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2b2833c-1eec-46c7-a7b0-59f3de53656c · outbound
Turning Off-Policy Tokens On-Policy: A Plug-in Approach for Improving LLM Alignment Unresolved cited work
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b32333c9-0f0b-4a8e-b95d-a2dfe061f2bc · outbound
Turning Off-Policy Tokens On-Policy: A Plug-in Approach for Improving LLM Alignment Table 4: Hyperparameters used in math experi- ments
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e5ad6d9-0f95-4f0a-9bb4-6e3ff8f0b025 · outbound
Turning Off-Policy Tokens On-Policy: A Plug-in Approach for Improving LLM Alignment Unresolved cited work
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8ee70f5-9673-485d-917a-f0db747c5ddc · outbound
Turning Off-Policy Tokens On-Policy: A Plug-in Approach for Improving LLM Alignment Table 8: Performance of SIS under varying policy staleness on Qwen3-8B-Base, where stalenessN is the number of mini-batch updates per rollout
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.