Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T12:52:03.527092Z
Paper Citation Record · LEDGER
As of 15 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 0 inbound Pith citation observations for arXiv:2607.19331.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T12:52:03.527092Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
44 of 44 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 9fb79cc2-c404-47a3-b915-cf07e0226282 · outbound
ISO: An RLVR-Native Optimization Stack The path not taken: Rlvr provably learns off the principals.arXiv preprint arXiv:2511.08567, 2025
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d155ade-fccd-44a4-a9f3-e6e4bf9c828d · outbound
ISO: An RLVR-Native Optimization Stack Grok: Ai assistant, 2025
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 109c1178-5c5a-4d7c-8532-d21f95ce3024 · outbound
ISO: An RLVR-Native Optimization Stack The next frontier of data training: Rl environments, February 2026
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c496e67-3f2d-42a3-88c9-d35d896f1301 · outbound
ISO: An RLVR-Native Optimization Stack DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6b2143b-7ea9-4528-98d3-66cbad8e05e9 · outbound
ISO: An RLVR-Native Optimization Stack DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5411b60-45a5-48a4-a6fb-6f5e5c341293 · outbound
ISO: An RLVR-Native Optimization Stack REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee2d3756-c0ec-4833-b6b9-811ffea78cc5 · outbound
ISO: An RLVR-Native Optimization Stack Maximum likelihood reinforce- ment learning.arXiv preprint arXiv:2602.02710, 2026
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7028889a-d730-47ee-a0c6-7e4a7169f6c0 · outbound
ISO: An RLVR-Native Optimization Stack slime: An llm post-training framework for rl scaling.https://github.com/THUDM/slime, 2025
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31f01d3d-e689-4049-8649-e893b182bb2a · outbound
ISO: An RLVR-Native Optimization Stack Hybridflow: A flexible and efficient rlhf framework
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b11721a-81b1-4a91-8b56-0e64c293fca5 · outbound
ISO: An RLVR-Native Optimization Stack Decoupled weight decay regularization
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c3ccaa9-27b3-4a29-aebb-400c1592ea16 · outbound
ISO: An RLVR-Native Optimization Stack Muon is Scalable for LLM Training
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52816f9f-79f5-4533-a45c-f73a39cf46d7 · outbound
ISO: An RLVR-Native Optimization Stack Apollo: Sgd-like memory, adamw-level performance
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f7eb4a2-f1b8-49dd-b36e-69ef81faf30f · outbound
ISO: An RLVR-Native Optimization Stack Fantastic Pretraining Optimizers and Where to Find Them
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9458d855-f1b3-4ecc-afb2-a657f835b91e · outbound
ISO: An RLVR-Native Optimization Stack RL's Razor: Why Online Reinforcement Learning Forgets Less
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 119f322d-a943-40f9-943d-fad6988d6e7d · outbound
ISO: An RLVR-Native Optimization Stack Reinforcement learning finetunes small subnetworks in large language models.arXiv preprint arXiv:2505.11711, 2025
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0780e322-c512-4258-ab23-4810c06b97a4 · outbound
ISO: An RLVR-Native Optimization Stack On-policy distillation.Thinking Machines Lab: Connec- tionism, 2025
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e76d293-0dad-4615-b52f-2ae0ba5f85be · outbound
ISO: An RLVR-Native Optimization Stack GLM-5: from Vibe Coding to Agentic Engineering
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bacfbee4-d40a-47c5-8201-102e767cf400 · outbound
ISO: An RLVR-Native Optimization Stack The invisible leash: Why rlvr may or may not escape its origin.arXiv preprint arXiv:2507.14843, 2025
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 486c41ab-4829-414a-898a-1eb672dc75eb · outbound
ISO: An RLVR-Native Optimization Stack ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38134479-33df-4628-9f80-57bffd7229d1 · outbound
ISO: An RLVR-Native Optimization Stack Brorl: Scaling reinforcement learning via broadened exploration.arXiv preprint arXiv:2510.01180, 2025
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79f025dc-e300-41aa-9d6c-67e0ed74d483 · outbound
ISO: An RLVR-Native Optimization Stack Embodied-R1: Reinforced Embodied Reasoning for General Robotic Manipulation
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af66c004-4399-4349-bfd9-b33bfccb439d · outbound
ISO: An RLVR-Native Optimization Stack Dion: Distributed orthonormalized updates.arXiv preprint arXiv:2504.05295, 2025
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2609e633-3506-4d5d-90b5-15426bcbe039 · outbound
ISO: An RLVR-Native Optimization Stack Qwen2.5 technical report, 2025
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76ef3918-82ce-4323-a643-fbef981af767 · outbound
ISO: An RLVR-Native Optimization Stack Co-evolving llm coder and unit tester via reinforcement learning.arXiv preprint arXiv:2506.03136, 2025
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47e2f468-b945-49e0-a61a-5378f017a1f6 · outbound
ISO: An RLVR-Native Optimization Stack ToolRL: Reward is All Tool Learning Needs
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5077b1a6-2bb5-4b3b-b98c-d3dd32cf2384 · outbound
ISO: An RLVR-Native Optimization Stack MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a50e5f76-aae4-422d-b6ae-c3b0d1a79069 · outbound
ISO: An RLVR-Native Optimization Stack Editing Models with Task Arithmetic
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be3ca050-58fa-49a9-88f8-e3dfa851b4e1 · outbound
ISO: An RLVR-Native Optimization Stack Ties-merging: Resolving interference when merging models.Advances in neural information processing systems, 36:7093–7115, 2023
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b3ef784-c894-4896-9559-8102ca1605ff · outbound
ISO: An RLVR-Native Optimization Stack Task singular vectors: Reducing task interference in model merging
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2453583-a213-4771-b9ea-2ec7471e6c20 · outbound
ISO: An RLVR-Native Optimization Stack Behavior knowledge merge in reinforced agentic models, 2026
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a69c3de-09a4-47c5-904d-dbc4404c5eab · outbound
ISO: An RLVR-Native Optimization Stack Orthogonal model merging.arXiv preprint arXiv:2602.05943, 2026
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 089bb966-f1d1-43a3-9cb9-6e6e94077479 · outbound
ISO: An RLVR-Native Optimization Stack When Importance Sampling Misallocates Credit: Asymmetric Ratios for Outcome-Supervised RL
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5626877-54be-43a9-b64e-65c5bfc77123 · outbound
ISO: An RLVR-Native Optimization Stack Justrl: Scaling a 1.5 b llm with a simple rl recipe
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3acdd35c-cce4-4afc-9028-589217449f8d · outbound
ISO: An RLVR-Native Optimization Stack Qwen3 technical report, 2025
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 672af4c9-5981-4490-ac04-e3502bb241ac · outbound
ISO: An RLVR-Native Optimization Stack Deepmath-103k: A large-scale, challenging, decontaminated, and verifiable mathematical dataset for advancing reasoning
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ae03c82-301e-4cb1-acf4-0f022b6fed8f · outbound
ISO: An RLVR-Native Optimization Stack Modular manifolds.Thinking Machines Lab: Connectionism, 2025
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 756f4711-e414-4ca6-bac4-8e5ae5ad7934 · outbound
ISO: An RLVR-Native Optimization Stack Reparameterized llm training via orthogonal equivalence transformation.Advances in Neural Information Processing Systems, 38:140775–140821, 2026
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc02de12-8664-4b8f-9ce7-dc09e30ff14e · outbound
ISO: An RLVR-Native Optimization Stack POET-X: Memory-efficient LLM Training by Scaling Orthogonal Transformation
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e2b093c-30a2-4fdb-bee3-d02a1923b038 · outbound
ISO: An RLVR-Native Optimization Stack Spectral Adapter: Fine-Tuning in Spectral Space
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7754d884-2a61-488c-8e1c-9b72b9aab03c · outbound
ISO: An RLVR-Native Optimization Stack Stella: Subspace learning in low-rank adaptation using stiefel manifold.Advances in Neural Information Processing Systems, 38:75066–75092, 2026
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0665ddc-670f-4eb3-8722-a88110719496 · outbound
ISO: An RLVR-Native Optimization Stack Pion: A Spectrum-Preserving Optimizer via Orthogonal Equivalence Transformation
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94062c08-7628-4134-afc7-7e51bdb34c4d · outbound
ISO: An RLVR-Native Optimization Stack Lora without regret.Thinking Machines Lab: Connectionism, 2025
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd033691-a5f8-4659-b63a-036f71bf2c21 · outbound
ISO: An RLVR-Native Optimization Stack Stabilizing Knowledge, Promoting Reasoning: Dual-Token Constraints for RLVR
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca4803f1-3782-4a8a-9bec-5f89986cd66a · outbound
ISO: An RLVR-Native Optimization Stack LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.