Pith. sign in

Paper Citation Record · LEDGER

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples

As of 5 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 0 inbound Pith citation observations for arXiv:2607.10481.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.10481 v2

Coverage vector

measured 63 of 63 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T07:23:30.344213Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

63 of 63 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved63
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0e4e57c1-b755-49ca-a60a-fa13eeb3b440 · outbound

This paper cites OpenAI o1 System Card.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples OpenAI o1 System Card

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:23.928177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:23.928177Z digest=sha256:3643b426849327db33f21380da8f0f28f5e26089220623bf4cd55ad646a95b35

Observation 450c0916-6b95-4f93-9157-9fdf49aee130 · outbound

This paper cites Nature , volume =.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples Nature , volume =

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:24.006590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:24.006590Z digest=sha256:11dd8454239892cdd371e339d58c3261f6e6a1e3f7c2e40936e07d47f64751dc

Observation e10c48c1-b611-4d29-88b3-ecf63eb507da · outbound

This paper cites DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:24.088146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:24.088146Z digest=sha256:ab1b94b471865d2fc8c52db457a133e0fd3006b9aa47450b70fa805ffd4cd220

Observation b2340d31-664f-444e-90bb-5f6515277157 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:24.173789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:24.173789Z digest=sha256:8997dab75f25f4e8e3a661b8228598ba142a34550e635a25ef69913d89eccb01

Observation 64a91aa6-9641-4b7c-8eed-7f521c90c960 · outbound

This paper cites Qwen3 Technical Report.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples Qwen3 Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:24.240938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:24.240938Z digest=sha256:32c1a6948761d73e16c8b0a767375706648fbec4d8fea945f6904922ca40c978

Observation 360b6a8f-6196-442f-b298-6d016242c44c · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:24.296105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:24.296105Z digest=sha256:c4de14445f076abf3319ba9ffccb5100792c3f988ae60a8816a9f877b76f6672

Observation e7e73dca-5877-4ff9-be43-1fb21fe231ba · outbound

This paper cites 2024 , eprint=.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples 2024 , eprint=

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:24.367087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:24.367087Z digest=sha256:f8b84a1ac5c1c16b71e4d46bc2b1d21159c24ed8d568f29862c1565ffa91d438

Observation 104fd3d2-3d38-4978-896d-7c756d512d2f · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:24.456293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:24.456293Z digest=sha256:ecacd511142d6a05e27e652be5ab890859a5a7020fc1ebde37b2d19e7bfc6b78

Observation ac960a03-5beb-4333-ae8a-a92078c66deb · outbound

This paper cites Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:24.544737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:24.544737Z digest=sha256:8f0849c2da935a45250e52a14ea1c868990e606c53c941656e737b37ffa05ac7

Observation cd2cf552-6097-47ec-87ef-e3049d793cd4 · outbound

This paper cites Proximal Policy Optimization Algorithms.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples Proximal Policy Optimization Algorithms

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:24.608784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:24.608784Z digest=sha256:2505edaba66e6dfd17371d018c49934ed421d66a16765d9bc5244cd5dbc763ce

Observation 986de566-4ebb-4d99-add4-7fd0c747c883 · outbound

This paper cites International conference on machine learning , pages=.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples International conference on machine learning , pages=

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:24.718817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:24.718817Z digest=sha256:e554f49dbea9fe6b3a7909fa608f63e0f1cb1a908ceaa7f9c9aaaf30470645db

Observation 1def997a-e4da-48fa-b1cd-a2048e2d81e7 · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:24.802584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:24.802584Z digest=sha256:147a7eda037bafb59feea29a55e96f003f2c15677608f623d00d2f2428daef86

Observation bf3897f7-3cf1-4dd2-8641-827a86680b12 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:24.906088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:24.906088Z digest=sha256:732f3cf0c5c2c6af130f7776efa13d9eb729559486f5978bbfc9d5d4ec487727

Observation 9dca93f1-90c3-4970-bf38-5bce292e677c · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:24.982944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:24.982944Z digest=sha256:48066240bb13b457e5d8b5f647bca5a4501cf9da5b93066e25a2700a003d7e24

Observation d398004d-dca2-4410-b88d-32b70387c35b · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples Understanding R1-Zero-Like Training: A Critical Perspective

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:25.052567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:25.052567Z digest=sha256:a481a5ba87da213d82a4e436b7a855c33bc0628b34986a76a3e56173a3e9ca4a

Observation 56514245-048e-4c88-a373-e180751db1be · outbound

This paper cites arXiv preprint arXiv:2507.20673 , year=.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples arXiv preprint arXiv:2507.20673 , year=

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:25.055939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:25.055939Z digest=sha256:0a0c012f737e7b28cff4a5374375c3989668fb0f948338b69ca6069b83c67d82

Observation 8522327a-21c8-42ba-924a-983938f2330c · outbound

This paper cites 2025 , eprint=.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples 2025 , eprint=

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:25.058589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:25.058589Z digest=sha256:bdfeebb9122fad649eb8771717be76082242d2653224a168e24d98ff709d836a

Observation 786d679e-0833-4e47-a0a1-ed845b8e4855 · outbound

This paper cites Group Sequence Policy Optimization.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples Group Sequence Policy Optimization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:25.134121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:25.134121Z digest=sha256:0d29cf2774b3c4f3e0dc725740ea09b808b4c8c3457bb10a4339e6401c974691

Observation 1de4cc21-8ae0-4c05-b957-4778999137c0 · outbound

This paper cites Your Efficient RL Framework Secretly Brings You Off-Policy RL Training , url =.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples Your Efficient RL Framework Secretly Brings You Off-Policy RL Training , url =

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:25.298585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:25.298585Z digest=sha256:7e6b436bb38efa2c824bf7a6517c8012b9a2f9f9d0069fd9f99605543328ca8c

Observation e4c92e99-8e6f-4e18-a5d5-b5e39c27a706 · outbound

This paper cites When Speed Kills Stability: Demystifying.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples When Speed Kills Stability: Demystifying

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:25.446955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:25.446955Z digest=sha256:3e52ca2de83bde47b2c0915e21a23f27a63ba02e87086d1208dab62e655ad6eb

Observation 8aaa4bc6-f4aa-421c-9d84-a588d0c4a3fb · outbound

This paper cites 2025 , eprint=.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples 2025 , eprint=

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:25.588238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:25.588238Z digest=sha256:7aa5819930ad74bf4d4acb9bcde822eea1aa4d46e20fa9d7827b077004c37f30

Observation c7960eeb-f82d-4099-8439-dedf72049e49 · outbound

This paper cites 2025 , eprint=.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples 2025 , eprint=

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:25.770620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:25.770620Z digest=sha256:90854d4e7e6931001f0dd1ecd105d26613624917f436946f6d387578fadf43ac

Observation 4d271953-69c3-468e-ae90-1b7a0bc3e4af · outbound

This paper cites Small Leak Can Sink a Great Ship--Boost RL Training on MoE with IcePop! , url =.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples Small Leak Can Sink a Great Ship--Boost RL Training on MoE with IcePop! , url =

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:25.894447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:25.894447Z digest=sha256:302ae13fda1427b37c065ebe0f3306f1a33df48d0cf6357c507e1d671f5d3311

Observation 2bbb3002-cd38-4609-962a-455a163933af · outbound

This paper cites Reward Hacking in Reinforcement Learning.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples Reward Hacking in Reinforcement Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:26.043996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:26.043996Z digest=sha256:3b664fa8f6def5254e4afac18e8dd0b3b3fe7a26415332bc630af7b03c9cb41f

Observation ce7b66e2-fe00-47f2-9e65-f8d5409ac127 · outbound

This paper cites International Conference on Machine Learning , pages=.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples International Conference on Machine Learning , pages=

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:26.169322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:26.169322Z digest=sha256:e2837cca3e66f3d49e2fb3f14a60a405acc4c3b864c4994f09a185b349fcab53

Observation 5181479d-b2d4-4764-b6a3-63147c4bdc3e · outbound

This paper cites an unresolved cited work.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:26.324565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:26.324565Z digest=sha256:229e877febe55a139c6a0942fbdd4e69a22f94ef5dd85c8b0fc46b4a4a2fe019

Observation 75e1e575-ceae-4910-b4be-8f1b294b5501 · outbound

This paper cites arXiv preprint arXiv:2510.22543 , year=.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples arXiv preprint arXiv:2510.22543 , year=

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:26.485479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:26.485479Z digest=sha256:0eceb88b94c7a15c39a635f755750745daff61e6ed067b0b1d0f9596f8798dc7

Observation f6ee98a5-88fd-4b88-afd1-bd1fff59d1b6 · outbound

This paper cites Why Language Models Hallucinate.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples Why Language Models Hallucinate

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:26.610302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:26.610302Z digest=sha256:e4b4157291ef9dd9c232d1ae565bcc28e2a6d844f78949a7d81be71c46bc652d

Observation 644ec384-f034-42a5-b532-3af159802735 · outbound

This paper cites arXiv preprint arXiv:2509.09177 , year=.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples arXiv preprint arXiv:2509.09177 , year=

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:26.753190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:26.753190Z digest=sha256:61c9885dc8a0276435a4ea1d81999b72503c53ef9118762bb178e463ce9f3ba8

Observation a1f20914-edae-452f-bf14-017970f5e5bf · outbound

This paper cites arXiv preprint arXiv:2508.17850 , year=.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples arXiv preprint arXiv:2508.17850 , year=

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:26.879211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:26.879211Z digest=sha256:54761893d1c200282743d232f668057f4a7049d29dbb55957ad3b93fdf6a80da

Observation c05e314c-dd90-4047-aaac-b1b841762423 · outbound

This paper cites Advances in Neural Information Processing Systems (NeurIPS) , volume=.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples Advances in Neural Information Processing Systems (NeurIPS) , volume=

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:27.022813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:27.022813Z digest=sha256:f0644ad347d197b32279dfe64b3d8250b87b637cf14147d22267be1b9eabfa96

Observation 0a741d53-da2c-4f08-9894-e3d11f428091 · outbound

This paper cites Advances in neural information processing systems , volume=.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples Advances in neural information processing systems , volume=

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:27.135881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:27.135881Z digest=sha256:557abd3f6924a224888d51ae9ee55a55d16c388b6ed60b5e35d132d60292a0b3

Observation 52d4cd12-08db-4a46-9185-101a8416eb73 · outbound

This paper cites On a few pitfalls in KL divergence gradient estimation for RL.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples On a few pitfalls in KL divergence gradient estimation for RL

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:27.338911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:27.338911Z digest=sha256:603b66e0f0fb01e476b173f252152815cc2960c0fcb9eaee5e658521ee3b0220

Observation 8268c740-42ef-4ac1-8371-7bfda21e2a1e · outbound

This paper cites arXiv preprint arXiv:2512.21852 , year=.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples arXiv preprint arXiv:2512.21852 , year=

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:27.495012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:27.495012Z digest=sha256:a8a0ac1ac9986fb931dd5df165691c726c5c071b292371de1a087a5133c2d8d4

Observation 2250deb4-8107-49ac-9686-a8c2265794d8 · outbound

This paper cites arXiv preprint arXiv:2510.20817 , year=.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples arXiv preprint arXiv:2510.20817 , year=

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:27.584045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:27.584045Z digest=sha256:22df3111cfc43ec87ba21d80dbe2460b3aeffcd75a99dfbe67ce3dfbd3b4f9ad

Observation 791ae013-e99d-4e5c-9fe0-14095100ee33 · outbound

This paper cites Beyond Reverse.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples Beyond Reverse

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:27.635901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:27.635901Z digest=sha256:062d87542403c9c1f06f364dcb8164a7aaaddf13ff14e1181d9ff22f699e9da3

Observation 2a94f418-4dc3-47e1-900f-7c6878ac2504 · outbound

This paper cites First Conference on Language Modeling , year=.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples First Conference on Language Modeling , year=

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:27.740294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:27.740294Z digest=sha256:3caf0f88d50c3dbf5dc11235581903db24a7f0b44295436fdefe40a727ff5da6

Observation 4e73d123-5477-4d87-b305-e231911add6d · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples Advances in Neural Information Processing Systems , volume=

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:27.805767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:27.805767Z digest=sha256:5c24d7a49227f33774a3befd22e1d9d0fed878efee811fb680a998bc026bd1d9

Observation bf03e8f2-4be6-42ff-8570-0a3750bf6877 · outbound

This paper cites 2021 , eprint=.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples 2021 , eprint=

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:27.878462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:27.878462Z digest=sha256:3e121bc9141d67911f9dcae270aabcdcebdef14a0889072544df7a145c9df9f1

Observation 0af955a8-16ca-4abe-8d0d-10a7dafbf55a · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples Does Reinforcement Learning Really Incentivize Reasoning Capacity in

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:27.968153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:27.968153Z digest=sha256:bdd6a43840e2dbea3a6b766d122bdd4c890df0c4ebeebae8b985bc9f210154ea

Observation 7d932ad8-5103-4abf-8d8e-3a68b7e1d25b · outbound

This paper cites The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:28.043198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:28.043198Z digest=sha256:30e8df4dceaba63eda7e317cc08322d62d59d2d7f59b88e1481d99f1d9e88d81

Observation a6f435ce-a6d7-4e87-9064-dc2720359b4f · outbound

This paper cites Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:28.120890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:28.120890Z digest=sha256:70d85a5f054497f4db3eb6a7b94656a6aad8f88136229aa47e88c66882783940

Observation 12a94e09-3547-486a-90a6-3ad997509fb6 · outbound

This paper cites 2023 , cdate=.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples 2023 , cdate=

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:28.195301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:28.195301Z digest=sha256:328e8bee68d848802c21554fb885802ed56c7867eebf7acf446e6147dcb1563c

Observation cfdfc1eb-6ad9-477e-8379-337302f35f1e · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:28.266130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:28.266130Z digest=sha256:9614e7e17330b88c6f904ca917f091d20d493ab2000a234c2e7c895709d875de

Observation 53f87756-ad26-448a-93e4-ff1191eadeb9 · outbound

This paper cites Learning to Reason under Off-Policy Guidance.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples Learning to Reason under Off-Policy Guidance

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:28.362049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:28.362049Z digest=sha256:2270c4bb42a43c8d51754dc411a36dce833f255f3df3b84921fb38da5a7ae022

Observation 7de04c6c-5efd-4ba3-ab2a-02f49bb48455 · outbound

This paper cites arXiv preprint arXiv:2506.07527 , year=.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples arXiv preprint arXiv:2506.07527 , year=

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:28.452592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:28.452592Z digest=sha256:ccb33cff7760f1badae964316bda87926cbc39b134bdf680eea4167e55137c51

Observation e7b7907b-ccc5-4b9b-be3f-b7943ba9f25f · outbound

This paper cites SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:28.545828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:28.545828Z digest=sha256:aa388911c219788934ebd6fb40b0dc7c9e17d9bd2863791add21bf5dfb26b7da

Observation ebad1479-c453-443e-9776-30c06c72f2f6 · outbound

This paper cites arXiv preprint arXiv:2509.04419 , year=.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples arXiv preprint arXiv:2509.04419 , year=

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:28.622105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:28.622105Z digest=sha256:324533e705d50adff7067941c85f8797c9d72abde587e96e6a674c2079ff9bdb

Observation 35b8c770-1527-4917-8fb3-60869d0916a8 · outbound

This paper cites RePO: Replay-Enhanced Policy Optimization.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples RePO: Replay-Enhanced Policy Optimization

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:28.717603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:28.717603Z digest=sha256:cba275fbff85096708b932078f0738b59c8984185d01ffe51ee0fc365a397ae8

Observation a6ba5694-badf-4fad-8793-20ff8e40f96b · outbound

This paper cites Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:28.818300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:28.818300Z digest=sha256:1c51cd50dbbd711bb3c309a4e941af764da19f69e0cd8a68f84e596f36d305e3

Observation eb6a0f3c-58e8-49d2-94f5-7c88251d8adf · outbound

This paper cites arXiv preprint arXiv:2510.02245 , year=.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples arXiv preprint arXiv:2510.02245 , year=

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:28.913789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:28.913789Z digest=sha256:03e10f1e3fb79de7dc8f3331f6225f6b97b92fe01b465008094c9e98b768efc1

Observation e769e426-7214-4d05-bd84-4105242bdd2e · outbound

This paper cites RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:29.013207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:29.013207Z digest=sha256:763be778e48143744569e799475abae15a83c5c90fab3c0c07d469f800a58041

Observation f7fe8cac-6531-4408-99c4-1bd10dd168f0 · outbound

This paper cites Selective Expert Guidance for Effective and Diverse Exploration in Reinforcement Learning of LLMs.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples Selective Expert Guidance for Effective and Diverse Exploration in Reinforcement Learning of LLMs

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:29.108343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:29.108343Z digest=sha256:908e21197a01f750ec2d1566ef329c4edcfb5588b13b5be7536a1087fd9940a1

Observation 8577ebe5-249b-4800-bb18-c975f86e57a1 · outbound

This paper cites arXiv preprint arXiv:2510.03865 , year=.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples arXiv preprint arXiv:2510.03865 , year=

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:29.154734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:29.154734Z digest=sha256:f5a3ea4f7301e27e64da0c38464c7da33c4e46df8dfdeeefd36f86d0e4762ee1

Observation 44e25a45-8a51-4105-a098-e97129606755 · outbound

This paper cites arXiv preprint arXiv:2509.07430 , year=.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples arXiv preprint arXiv:2509.07430 , year=

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:29.245333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:29.245333Z digest=sha256:0e8ada96de5323531458c34ab0404760b33518e70812eff8f71992ce3a90db71

Observation dd617188-aaca-41c7-8cff-f6e816320646 · outbound

This paper cites The Twelfth International Conference on Learning Representations , year=.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples The Twelfth International Conference on Learning Representations , year=

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:29.404192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:29.404192Z digest=sha256:2ab5d62fc85134345d131ff1e942cfe2b7070b3ff82cb396d7ea645de91964fb

Observation 9b022685-e7d2-4fbb-aa66-842bfd501e22 · outbound

This paper cites MiMo-V2-Flash Technical Report.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples MiMo-V2-Flash Technical Report

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:29.526424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:29.526424Z digest=sha256:5e0bbd12ef11b8d3a4710b7986d5d21d39da16ac0b873b087636ef2dbd80b02f

Observation 28cdf359-2843-4333-b14d-b21f1a92d30f · outbound

This paper cites arXiv preprint arXiv:2603.22117 , year=.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples arXiv preprint arXiv:2603.22117 , year=

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:29.693770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:29.693770Z digest=sha256:b487ce05bc7eb21f90b2c6b579dc1133a2c217d473e95e55bb39615393e0a75a

Observation 682775fd-3a03-44a7-aa4c-dab4e66ddea5 · outbound

This paper cites arXiv preprint arXiv:2603.19835 , year=.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples arXiv preprint arXiv:2603.19835 , year=

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:29.817926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:29.817926Z digest=sha256:d8df0870c1586a417f6ebd06fcb2b7d8f2c50b97f3ea3e938a61be0a63b8d98b

Observation 06234128-6f8c-46a4-a1cf-4929522df424 · outbound

This paper cites arXiv preprint arXiv:2603.22446 , year=.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples arXiv preprint arXiv:2603.22446 , year=

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:29.943795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:29.943795Z digest=sha256:ba0583b650a243e5ad3d47b248badc2a0eacc19a63811f2bad17fd76fc1fbee6

Observation 79e6168e-577f-4be5-ba8a-bb1cfab2dcca · outbound

This paper cites Experience Augmented Policy Optimization for LLM Reasoning.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples Experience Augmented Policy Optimization for LLM Reasoning

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:30.079906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:30.079906Z digest=sha256:eaf6019037c7eda414e36b27d49b99d3b7d91ab5d423985e3959706336c30997

Observation 9b5031bc-8e80-4294-bedd-c73a5983a727 · outbound

This paper cites One-Way Policy Optimization for Self-Evolving LLMs.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples One-Way Policy Optimization for Self-Evolving LLMs

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:30.215870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:30.215870Z digest=sha256:4ff8b299fefd96f258105874e74576bcf5d917b59542d1ca58ad5b3e1d5f327c

Observation 50e9359a-fb92-42cc-a5c2-a388db996a82 · outbound

This paper cites Clipping Bottleneck: Stabilizing RLVR via Stochastic Recovery of Near-Boundary Signals.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples Clipping Bottleneck: Stabilizing RLVR via Stochastic Recovery of Near-Boundary Signals

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:30.344213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:30.344213Z digest=sha256:bdebf3b5c63e31f369b07230e13a699066bebd70a2492dda66160ba36267bd3c

Pith citing papers

No inbound Pith citation observations are available.