Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T07:23:30.344213Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 0 inbound Pith citation observations for arXiv:2607.10481.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T07:23:30.344213Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
63 of 63 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 0e4e57c1-b755-49ca-a60a-fa13eeb3b440 · outbound
ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples OpenAI o1 System Card
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 450c0916-6b95-4f93-9157-9fdf49aee130 · outbound
ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples Nature , volume =
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e10c48c1-b611-4d29-88b3-ecf63eb507da · outbound
ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2340d31-664f-444e-90bb-5f6515277157 · outbound
ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64a91aa6-9641-4b7c-8eed-7f521c90c960 · outbound
ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples Qwen3 Technical Report
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 360b6a8f-6196-442f-b298-6d016242c44c · outbound
ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7e73dca-5877-4ff9-be43-1fb21fe231ba · outbound
ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples 2024 , eprint=
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 104fd3d2-3d38-4978-896d-7c756d512d2f · outbound
ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples Tulu 3: Pushing Frontiers in Open Language Model Post-Training
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac960a03-5beb-4333-ae8a-a92078c66deb · outbound
ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd2cf552-6097-47ec-87ef-e3049d793cd4 · outbound
ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples Proximal Policy Optimization Algorithms
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 986de566-4ebb-4d99-add4-7fd0c747c883 · outbound
ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples International conference on machine learning , pages=
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1def997a-e4da-48fa-b1cd-a2048e2d81e7 · outbound
ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples High-Dimensional Continuous Control Using Generalized Advantage Estimation
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf3897f7-3cf1-4dd2-8641-827a86680b12 · outbound
ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9dca93f1-90c3-4970-bf38-5bce292e677c · outbound
ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d398004d-dca2-4410-b88d-32b70387c35b · outbound
ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples Understanding R1-Zero-Like Training: A Critical Perspective
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56514245-048e-4c88-a373-e180751db1be · outbound
ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples arXiv preprint arXiv:2507.20673 , year=
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8522327a-21c8-42ba-924a-983938f2330c · outbound
ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples 2025 , eprint=
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 786d679e-0833-4e47-a0a1-ed845b8e4855 · outbound
ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples Group Sequence Policy Optimization
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1de4cc21-8ae0-4c05-b957-4778999137c0 · outbound
ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples Your Efficient RL Framework Secretly Brings You Off-Policy RL Training , url =
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4c92e99-8e6f-4e18-a5d5-b5e39c27a706 · outbound
ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples When Speed Kills Stability: Demystifying
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8aaa4bc6-f4aa-421c-9d84-a588d0c4a3fb · outbound
ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples 2025 , eprint=
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7960eeb-f82d-4099-8439-dedf72049e49 · outbound
ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples 2025 , eprint=
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d271953-69c3-468e-ae90-1b7a0bc3e4af · outbound
ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples Small Leak Can Sink a Great Ship--Boost RL Training on MoE with IcePop! , url =
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2bbb3002-cd38-4609-962a-455a163933af · outbound
ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples Reward Hacking in Reinforcement Learning
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce7b66e2-fe00-47f2-9e65-f8d5409ac127 · outbound
ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples International Conference on Machine Learning , pages=
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5181479d-b2d4-4764-b6a3-63147c4bdc3e · outbound
ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples Unresolved cited work
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75e1e575-ceae-4910-b4be-8f1b294b5501 · outbound
ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples arXiv preprint arXiv:2510.22543 , year=
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6ee98a5-88fd-4b88-afd1-bd1fff59d1b6 · outbound
ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples Why Language Models Hallucinate
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 644ec384-f034-42a5-b532-3af159802735 · outbound
ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples arXiv preprint arXiv:2509.09177 , year=
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1f20914-edae-452f-bf14-017970f5e5bf · outbound
ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples arXiv preprint arXiv:2508.17850 , year=
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c05e314c-dd90-4047-aaac-b1b841762423 · outbound
ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples Advances in Neural Information Processing Systems (NeurIPS) , volume=
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a741d53-da2c-4f08-9894-e3d11f428091 · outbound
ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples Advances in neural information processing systems , volume=
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52d4cd12-08db-4a46-9185-101a8416eb73 · outbound
ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples On a few pitfalls in KL divergence gradient estimation for RL
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8268c740-42ef-4ac1-8371-7bfda21e2a1e · outbound
ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples arXiv preprint arXiv:2512.21852 , year=
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2250deb4-8107-49ac-9686-a8c2265794d8 · outbound
ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples arXiv preprint arXiv:2510.20817 , year=
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 791ae013-e99d-4e5c-9fe0-14095100ee33 · outbound
ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples Beyond Reverse
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a94f418-4dc3-47e1-900f-7c6878ac2504 · outbound
ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples First Conference on Language Modeling , year=
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e73d123-5477-4d87-b305-e231911add6d · outbound
ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples Advances in Neural Information Processing Systems , volume=
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf03e8f2-4be6-42ff-8570-0a3750bf6877 · outbound
ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples 2021 , eprint=
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0af955a8-16ca-4abe-8d0d-10a7dafbf55a · outbound
ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples Does Reinforcement Learning Really Incentivize Reasoning Capacity in
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d932ad8-5103-4abf-8d8e-3a68b7e1d25b · outbound
ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6f435ce-a6d7-4e87-9064-dc2720359b4f · outbound
ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12a94e09-3547-486a-90a6-3ad997509fb6 · outbound
ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples 2023 , cdate=
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cfdfc1eb-6ad9-477e-8379-337302f35f1e · outbound
ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53f87756-ad26-448a-93e4-ff1191eadeb9 · outbound
ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples Learning to Reason under Off-Policy Guidance
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7de04c6c-5efd-4ba3-ab2a-02f49bb48455 · outbound
ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples arXiv preprint arXiv:2506.07527 , year=
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7b7907b-ccc5-4b9b-be3f-b7943ba9f25f · outbound
ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ebad1479-c453-443e-9776-30c06c72f2f6 · outbound
ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples arXiv preprint arXiv:2509.04419 , year=
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35b8c770-1527-4917-8fb3-60869d0916a8 · outbound
ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples RePO: Replay-Enhanced Policy Optimization
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6ba5694-badf-4fad-8793-20ff8e40f96b · outbound
ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb6a0f3c-58e8-49d2-94f5-7c88251d8adf · outbound
ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples arXiv preprint arXiv:2510.02245 , year=
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e769e426-7214-4d05-bd84-4105242bdd2e · outbound
ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7fe8cac-6531-4408-99c4-1bd10dd168f0 · outbound
ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples Selective Expert Guidance for Effective and Diverse Exploration in Reinforcement Learning of LLMs
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8577ebe5-249b-4800-bb18-c975f86e57a1 · outbound
ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples arXiv preprint arXiv:2510.03865 , year=
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44e25a45-8a51-4105-a098-e97129606755 · outbound
ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples arXiv preprint arXiv:2509.07430 , year=
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd617188-aaca-41c7-8cff-f6e816320646 · outbound
ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples The Twelfth International Conference on Learning Representations , year=
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b022685-e7d2-4fbb-aa66-842bfd501e22 · outbound
ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples MiMo-V2-Flash Technical Report
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28cdf359-2843-4333-b14d-b21f1a92d30f · outbound
ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples arXiv preprint arXiv:2603.22117 , year=
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 682775fd-3a03-44a7-aa4c-dab4e66ddea5 · outbound
ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples arXiv preprint arXiv:2603.19835 , year=
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06234128-6f8c-46a4-a1cf-4929522df424 · outbound
ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples arXiv preprint arXiv:2603.22446 , year=
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79e6168e-577f-4be5-ba8a-bb1cfab2dcca · outbound
ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples Experience Augmented Policy Optimization for LLM Reasoning
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b5031bc-8e80-4294-bedd-c73a5983a727 · outbound
ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples One-Way Policy Optimization for Self-Evolving LLMs
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50e9359a-fb92-42cc-a5c2-a388db996a82 · outbound
ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples Clipping Bottleneck: Stabilizing RLVR via Stochastic Recovery of Near-Boundary Signals
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.