Pith. sign in

Paper Citation Record · LEDGER

ISO: An RLVR-Native Optimization Stack

As of 15 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 0 inbound Pith citation observations for arXiv:2607.19331.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.19331 v1

Coverage vector

measured 44 of 44 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T12:52:03.527092Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

44 of 44 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved43
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9fb79cc2-c404-47a3-b915-cf07e0226282 · outbound

This paper cites The path not taken: Rlvr provably learns off the principals.arXiv preprint arXiv:2511.08567, 2025.

ISO: An RLVR-Native Optimization Stack The path not taken: Rlvr provably learns off the principals.arXiv preprint arXiv:2511.08567, 2025

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T12:51:58.381326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:51:58.381326Z digest=sha256:f7d5c9023c95ad9491ba7fe80d7fd61d4d1facd547b5e345744d9c59108d1a9d

Observation 2d155ade-fccd-44a4-a9f3-e6e4bf9c828d · outbound

This paper cites Grok: Ai assistant, 2025.

ISO: An RLVR-Native Optimization Stack Grok: Ai assistant, 2025

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T12:51:58.456075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:51:58.456075Z digest=sha256:7f0f71072be96ff69da6205265f18e6dc438283be6843483275833aa99ee6c81

Observation 109c1178-5c5a-4d7c-8532-d21f95ce3024 · outbound

This paper cites The next frontier of data training: Rl environments, February 2026.

ISO: An RLVR-Native Optimization Stack The next frontier of data training: Rl environments, February 2026

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T12:51:58.547495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:51:58.547495Z digest=sha256:0ca3f192d6a37c6d07d7634b2379c8e5df7f46ebb17f1438fab1ce23b94c0b59

Observation 8c496e67-3f2d-42a3-88c9-d35d896f1301 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

ISO: An RLVR-Native Optimization Stack DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T12:51:58.615937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:51:58.615937Z digest=sha256:c5575fdcecdb6fd98a8b42b5713bff6c51c025d1d6d557df9733edbf87afafaf

Observation e6b2143b-7ea9-4528-98d3-66cbad8e05e9 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

ISO: An RLVR-Native Optimization Stack DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T12:51:58.724095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:51:58.724095Z digest=sha256:6ae9d6324db7c40b136dd474a6f6f688fadfbfff2453d41494260653576988b1

Observation f5411b60-45a5-48a4-a6fb-6f5e5c341293 · outbound

This paper cites REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization.

ISO: An RLVR-Native Optimization Stack REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T12:51:58.848244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:51:58.848244Z digest=sha256:5a6e6f73188b98b77e8d42ad1ed038424282edc469ac00dde1a2613aac6c398e

Observation ee2d3756-c0ec-4833-b6b9-811ffea78cc5 · outbound

This paper cites Maximum likelihood reinforce- ment learning.arXiv preprint arXiv:2602.02710, 2026.

ISO: An RLVR-Native Optimization Stack Maximum likelihood reinforce- ment learning.arXiv preprint arXiv:2602.02710, 2026

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T12:51:58.987971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:51:58.987971Z digest=sha256:387f490fe9665b7ef6ad0e3e5cd900d55cf5833c8b79ca8ac96b628f0afd4f2b

Observation 7028889a-d730-47ee-a0c6-7e4a7169f6c0 · outbound

This paper cites slime: An llm post-training framework for rl scaling.https://github.com/THUDM/slime, 2025.

ISO: An RLVR-Native Optimization Stack slime: An llm post-training framework for rl scaling.https://github.com/THUDM/slime, 2025

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T12:51:59.150161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:51:59.150161Z digest=sha256:8847a86d6d74da602f65d77c0c9decdecca34dd4172a5e196966be7d1005f022

Observation 31f01d3d-e689-4049-8649-e893b182bb2a · outbound

This paper cites Hybridflow: A flexible and efficient rlhf framework.

ISO: An RLVR-Native Optimization Stack Hybridflow: A flexible and efficient rlhf framework

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T12:51:59.263916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:51:59.263916Z digest=sha256:10b9d453ad6640c2ed92383765f2b4b9b95819b9a559d75f2ef08b975c942816

Observation 6b11721a-81b1-4a91-8b56-0e64c293fca5 · outbound

This paper cites Decoupled weight decay regularization.

ISO: An RLVR-Native Optimization Stack Decoupled weight decay regularization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T12:51:59.385713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:51:59.385713Z digest=sha256:c316a09ba79ce273058da6b0bb45c8f9bd27b6a9e775ea657fb635fb66910e04

Observation 0c3ccaa9-27b3-4a29-aebb-400c1592ea16 · outbound

This paper cites Muon is Scalable for LLM Training.

ISO: An RLVR-Native Optimization Stack Muon is Scalable for LLM Training

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T12:51:59.490848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:51:59.490848Z digest=sha256:fda6af384b6951a93d54153ef4d417e56acba2cd406e39bbad0c2fb207e67d4e

Observation 52816f9f-79f5-4533-a45c-f73a39cf46d7 · outbound

This paper cites Apollo: Sgd-like memory, adamw-level performance.

ISO: An RLVR-Native Optimization Stack Apollo: Sgd-like memory, adamw-level performance

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T12:51:59.572851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:51:59.572851Z digest=sha256:e83d4e7e248ffd5cac2559f95df0fed9a1af2baf6d2c37927abba9ee2e4c5cc4

Observation 3f7eb4a2-f1b8-49dd-b36e-69ef81faf30f · outbound

This paper cites Fantastic Pretraining Optimizers and Where to Find Them.

ISO: An RLVR-Native Optimization Stack Fantastic Pretraining Optimizers and Where to Find Them

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T12:51:59.644271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:51:59.644271Z digest=sha256:029427857679dbcfe4b8613013718084553428ef09d6b8a0722ff51678ba41d9

Observation 9458d855-f1b3-4ecc-afb2-a657f835b91e · outbound

This paper cites RL's Razor: Why Online Reinforcement Learning Forgets Less.

ISO: An RLVR-Native Optimization Stack RL's Razor: Why Online Reinforcement Learning Forgets Less

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T12:51:59.749347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:51:59.749347Z digest=sha256:0286c4dc951a9fd9b418a19d8d7ae78b3756215a4ccfc88afbe17c4e959b4dca

Observation 119f322d-a943-40f9-943d-fad6988d6e7d · outbound

This paper cites Reinforcement learning finetunes small subnetworks in large language models.arXiv preprint arXiv:2505.11711, 2025.

ISO: An RLVR-Native Optimization Stack Reinforcement learning finetunes small subnetworks in large language models.arXiv preprint arXiv:2505.11711, 2025

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T12:51:59.860730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:51:59.860730Z digest=sha256:45f8937d290365e168da2a96fddb6a2f4b378ab8c65695181b60c92f35e700dd

Observation 0780e322-c512-4258-ab23-4810c06b97a4 · outbound

This paper cites On-policy distillation.Thinking Machines Lab: Connec- tionism, 2025.

ISO: An RLVR-Native Optimization Stack On-policy distillation.Thinking Machines Lab: Connec- tionism, 2025

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T12:51:59.927168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:51:59.927168Z digest=sha256:e57c2433e5b5d3f3ebbbeb3e7f7dfcf3e067c1956bfd491f24f11f251267e098

Observation 9e76d293-0dad-4615-b52f-2ae0ba5f85be · outbound

This paper cites GLM-5: from Vibe Coding to Agentic Engineering.

ISO: An RLVR-Native Optimization Stack GLM-5: from Vibe Coding to Agentic Engineering

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T12:52:00.004361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:52:00.004361Z digest=sha256:5e5d8a7de2d6fdfdcf469423896f8d38f9bde1535f91e07eeab5ac653b661fb4

Observation bacfbee4-d40a-47c5-8201-102e767cf400 · outbound

This paper cites The invisible leash: Why rlvr may or may not escape its origin.arXiv preprint arXiv:2507.14843, 2025.

ISO: An RLVR-Native Optimization Stack The invisible leash: Why rlvr may or may not escape its origin.arXiv preprint arXiv:2507.14843, 2025

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T12:52:00.067015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:52:00.067015Z digest=sha256:2c8b9655880e44c53ac35d4590adb2a6f268dbdf18a2fda7edb41f7a67ba18d5

Observation 486c41ab-4829-414a-898a-1eb672dc75eb · outbound

This paper cites ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models.

ISO: An RLVR-Native Optimization Stack ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T12:52:00.138146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:52:00.138146Z digest=sha256:2f8f55dc1777fc49e99242deb03449a26adf0f32d0dc7fcaa6076e862f3ce249

Observation 38134479-33df-4628-9f80-57bffd7229d1 · outbound

This paper cites Brorl: Scaling reinforcement learning via broadened exploration.arXiv preprint arXiv:2510.01180, 2025.

ISO: An RLVR-Native Optimization Stack Brorl: Scaling reinforcement learning via broadened exploration.arXiv preprint arXiv:2510.01180, 2025

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T12:52:00.283718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:52:00.283718Z digest=sha256:7bc626354e6d718d5812e93dd37605bb233a51f3dd3f20a8325bf2ef3cfbf4a6

Observation 79f025dc-e300-41aa-9d6c-67e0ed74d483 · outbound

This paper cites Embodied-R1: Reinforced Embodied Reasoning for General Robotic Manipulation.

ISO: An RLVR-Native Optimization Stack Embodied-R1: Reinforced Embodied Reasoning for General Robotic Manipulation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T12:52:00.380279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:52:00.380279Z digest=sha256:50b98aad4329aad4743d0af2f0ed0782df2b0a108a20577e825bc6c673f605b5

Observation af66c004-4399-4349-bfd9-b33bfccb439d · outbound

This paper cites Dion: Distributed orthonormalized updates.arXiv preprint arXiv:2504.05295, 2025.

ISO: An RLVR-Native Optimization Stack Dion: Distributed orthonormalized updates.arXiv preprint arXiv:2504.05295, 2025

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T12:52:00.445811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:52:00.445811Z digest=sha256:d56070bc39e7f028e2b39c723d3af2501fffc219b43b2f8a4531e04015b0e0d8

Observation 2609e633-3506-4d5d-90b5-15426bcbe039 · outbound

This paper cites Qwen2.5 technical report, 2025.

ISO: An RLVR-Native Optimization Stack Qwen2.5 technical report, 2025

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T12:52:00.529969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:52:00.529969Z digest=sha256:687f4f7e58ccb31d5eb6c91f930332eb7502a45a38b2717813d85d5ce8ac8b6f

Observation 76ef3918-82ce-4323-a643-fbef981af767 · outbound

This paper cites Co-evolving llm coder and unit tester via reinforcement learning.arXiv preprint arXiv:2506.03136, 2025.

ISO: An RLVR-Native Optimization Stack Co-evolving llm coder and unit tester via reinforcement learning.arXiv preprint arXiv:2506.03136, 2025

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T12:52:00.627803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:52:00.627803Z digest=sha256:1e902162c40ea17c7b17593584a2509e95557d2f1c4b270fca584caffaf9c135

Observation 47e2f468-b945-49e0-a61a-5378f017a1f6 · outbound

This paper cites ToolRL: Reward is All Tool Learning Needs.

ISO: An RLVR-Native Optimization Stack ToolRL: Reward is All Tool Learning Needs

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T12:52:00.702419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:52:00.702419Z digest=sha256:846bfc37d73aadc21446d84f21d704932d3bc4ea175ccf3a80dc0fb3cbca3584

Observation 5077b1a6-2bb5-4b3b-b98c-d3dd32cf2384 · outbound

This paper cites MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent.

ISO: An RLVR-Native Optimization Stack MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T12:52:00.784884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:52:00.784884Z digest=sha256:8f401328a3b9fb88c07c7643f4da0d9477a87bf539b340eee73707a5d083fa05

Observation a50e5f76-aae4-422d-b6ae-c3b0d1a79069 · outbound

This paper cites Editing Models with Task Arithmetic.

ISO: An RLVR-Native Optimization Stack Editing Models with Task Arithmetic

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T12:52:00.879258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:52:00.879258Z digest=sha256:b64ad51ed99d31bc0a06746a13bf0953745c6908cd7c80b2503f81491ba34edc

Observation be3ca050-58fa-49a9-88f8-e3dfa851b4e1 · outbound

This paper cites Ties-merging: Resolving interference when merging models.Advances in neural information processing systems, 36:7093–7115, 2023.

ISO: An RLVR-Native Optimization Stack Ties-merging: Resolving interference when merging models.Advances in neural information processing systems, 36:7093–7115, 2023

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T12:52:00.979104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:52:00.979104Z digest=sha256:66c52f321e3e4009ba2a7008678afcaeee16269d2c544a70ac1f05eef8159b18

Observation 9b3ef784-c894-4896-9559-8102ca1605ff · outbound

This paper cites Task singular vectors: Reducing task interference in model merging.

ISO: An RLVR-Native Optimization Stack Task singular vectors: Reducing task interference in model merging

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T12:52:01.188253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:52:01.188253Z digest=sha256:232e471ce4ba50a20a41603ba0fb9f57b7dc817f95ea08ed23011bbeaa5b2a88

Observation e2453583-a213-4771-b9ea-2ec7471e6c20 · outbound

This paper cites Behavior knowledge merge in reinforced agentic models, 2026.

ISO: An RLVR-Native Optimization Stack Behavior knowledge merge in reinforced agentic models, 2026

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T12:52:01.359422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:52:01.359422Z digest=sha256:32380d9fa382a8250f0734a1c31065abfb27b3c525cea50f2e1ff0144c68878f

Observation 4a69c3de-09a4-47c5-904d-dbc4404c5eab · outbound

This paper cites Orthogonal model merging.arXiv preprint arXiv:2602.05943, 2026.

ISO: An RLVR-Native Optimization Stack Orthogonal model merging.arXiv preprint arXiv:2602.05943, 2026

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T12:52:01.568383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:52:01.568383Z digest=sha256:5f908e8d8ba27db9316ce97639aff2656c08a06557554f896b4f242eb876f626

Observation 089bb966-f1d1-43a3-9cb9-6e6e94077479 · outbound

This paper cites When Importance Sampling Misallocates Credit: Asymmetric Ratios for Outcome-Supervised RL.

ISO: An RLVR-Native Optimization Stack When Importance Sampling Misallocates Credit: Asymmetric Ratios for Outcome-Supervised RL

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T12:52:01.770385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:52:01.770385Z digest=sha256:81b92f51db7675a77c46ff762af4cb74b5a298b42317dcd734f9f84f65550d62

Observation a5626877-54be-43a9-b64e-65c5bfc77123 · outbound

This paper cites Justrl: Scaling a 1.5 b llm with a simple rl recipe.

ISO: An RLVR-Native Optimization Stack Justrl: Scaling a 1.5 b llm with a simple rl recipe

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T12:52:01.995877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:52:01.995877Z digest=sha256:cb7df1fab339fd3fe6c7671d55b7d74f2ee8b162498a2312a10aa6450fbef235

Observation 3acdd35c-cce4-4afc-9028-589217449f8d · outbound

This paper cites Qwen3 technical report, 2025.

ISO: An RLVR-Native Optimization Stack Qwen3 technical report, 2025

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T12:52:02.184137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:52:02.184137Z digest=sha256:e904a9bcc363b260eac83a6bc58378e0e2a743e070da0111eb34607f30e5fa1b

Observation 672af4c9-5981-4490-ac04-e3502bb241ac · outbound

This paper cites Deepmath-103k: A large-scale, challenging, decontaminated, and verifiable mathematical dataset for advancing reasoning.

ISO: An RLVR-Native Optimization Stack Deepmath-103k: A large-scale, challenging, decontaminated, and verifiable mathematical dataset for advancing reasoning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T12:52:02.283198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:52:02.283198Z digest=sha256:1cb11e440c12939d997fa7fe48fea8bae499b1981f24d622cf1a2d5cbdfc1f22

Observation 9ae03c82-301e-4cb1-acf4-0f022b6fed8f · outbound

This paper cites Modular manifolds.Thinking Machines Lab: Connectionism, 2025.

ISO: An RLVR-Native Optimization Stack Modular manifolds.Thinking Machines Lab: Connectionism, 2025

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T12:52:02.396885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:52:02.396885Z digest=sha256:60fc7420876b9551a12c80f41fb3a540511b00d0433f10f0d4afba985ec1824f

Observation 756f4711-e414-4ca6-bac4-8e5ae5ad7934 · outbound

This paper cites Reparameterized llm training via orthogonal equivalence transformation.Advances in Neural Information Processing Systems, 38:140775–140821, 2026.

ISO: An RLVR-Native Optimization Stack Reparameterized llm training via orthogonal equivalence transformation.Advances in Neural Information Processing Systems, 38:140775–140821, 2026

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T12:52:02.495025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:52:02.495025Z digest=sha256:ac903846ca37de578651be816fd824cf2e1e6c9265e304c67a2f9a2dd5b909f2

Observation bc02de12-8664-4b8f-9ce7-dc09e30ff14e · outbound

This paper cites POET-X: Memory-efficient LLM Training by Scaling Orthogonal Transformation.

ISO: An RLVR-Native Optimization Stack POET-X: Memory-efficient LLM Training by Scaling Orthogonal Transformation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T12:52:02.679503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:52:02.679503Z digest=sha256:9b6c55689d1867b252d394ec68e0301e9c3ce7efce4cc0fb82b53e9c02768311

Observation 1e2b093c-30a2-4fdb-bee3-d02a1923b038 · outbound

This paper cites Spectral Adapter: Fine-Tuning in Spectral Space.

ISO: An RLVR-Native Optimization Stack Spectral Adapter: Fine-Tuning in Spectral Space

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T12:52:02.857528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:52:02.857528Z digest=sha256:50cba016a6268cebb0917c1271fb4536cd3dfde5499d20e44f0df6da39215174

Observation 7754d884-2a61-488c-8e1c-9b72b9aab03c · outbound

This paper cites Stella: Subspace learning in low-rank adaptation using stiefel manifold.Advances in Neural Information Processing Systems, 38:75066–75092, 2026.

ISO: An RLVR-Native Optimization Stack Stella: Subspace learning in low-rank adaptation using stiefel manifold.Advances in Neural Information Processing Systems, 38:75066–75092, 2026

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T12:52:02.991752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:52:02.991752Z digest=sha256:8d615fabc8e305acdaf0ba94752ff3aac7d636525df9501eb9f643cd56448430

Observation c0665ddc-670f-4eb3-8722-a88110719496 · outbound

This paper cites Pion: A Spectrum-Preserving Optimizer via Orthogonal Equivalence Transformation.

ISO: An RLVR-Native Optimization Stack Pion: A Spectrum-Preserving Optimizer via Orthogonal Equivalence Transformation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T12:52:03.085095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:52:03.085095Z digest=sha256:fceaf4d7ce2451ef646a6198ba2def5a053bb0800de696528f2977d67253c5ea

Observation 94062c08-7628-4134-afc7-7e51bdb34c4d · outbound

This paper cites Lora without regret.Thinking Machines Lab: Connectionism, 2025.

ISO: An RLVR-Native Optimization Stack Lora without regret.Thinking Machines Lab: Connectionism, 2025

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T12:52:03.224941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:52:03.224941Z digest=sha256:112d86a6a00fe4f59afa28a1e4d56295e84d1e7849fcee3f7cfc6f125dfd6867

Observation dd033691-a5f8-4659-b63a-036f71bf2c21 · outbound

This paper cites Stabilizing Knowledge, Promoting Reasoning: Dual-Token Constraints for RLVR.

ISO: An RLVR-Native Optimization Stack Stabilizing Knowledge, Promoting Reasoning: Dual-Token Constraints for RLVR

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T12:52:03.418484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:52:03.418484Z digest=sha256:34d30508b04abe6cc88f6446bf4a91e95b43a26b7532c256de1ff35557ed5006

Observation ca4803f1-3782-4a8a-9bec-5f89986cd66a · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

ISO: An RLVR-Native Optimization Stack LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 44

Resolution
malformed identifier
no resolver link, observed 2026-08-01T12:52:03.527092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:52:03.527092Z digest=sha256:c9cbdb2b2ac8ed8e4e8efbd77b1eb33766921a790eada56d7a4f27b0ce8cea2c

Pith citing papers

No inbound Pith citation observations are available.