Pith. sign in

Paper Citation Record · LEDGER

When Does Muon Help Agentic Reinforcement Learning?

As of 7 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 0 inbound Pith citation observations for arXiv:2607.16169.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.16169 v4

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T04:19:39.717440Z

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

22 of 22 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a1c2244b-635d-41eb-bdd2-695eb4c5fb18 · outbound

This paper cites InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics.

When Does Muon Help Agentic Reinforcement Learning? InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T04:19:37.337225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:19:37.337225Z digest=sha256:2318c9d8d2ee4f530ceb6bf9677a7501ec942d5bca9706f96705fd30d63df0b3

Observation 9ab76eb5-1731-4888-82ef-7badbef85f51 · outbound

This paper cites ArXiv:2602.22817.

When Does Muon Help Agentic Reinforcement Learning? ArXiv:2602.22817

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T04:19:37.610753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:19:37.610753Z digest=sha256:33a2baa7da06e92eccb02bd4cc62ad49e455668a02a2c2e41cbe8961a8f55fc7

Observation 480bcf6f-b36b-4529-b349-a8c5530ef28b · outbound

This paper cites MiMuon: Mixed Muon Optimizer with Improved Generalization for Large Models.

When Does Muon Help Agentic Reinforcement Learning? MiMuon: Mixed Muon Optimizer with Improved Generalization for Large Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T04:19:37.724140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:19:37.724140Z digest=sha256:4c25dce18b513aa6c28ba8ebbcdb8353dd370df0e3be9824174ec2a5e7a9eac5

Observation f0a3e926-f1ad-4016-8abd-f1b2aeac5ec5 · outbound

This paper cites SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales.

When Does Muon Help Agentic Reinforcement Learning? SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T04:19:37.877565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:19:37.877565Z digest=sha256:6877e5e0f1fcb5a10b91480224d7a9da424d00305fe285a9d9e8b16c5bf6562c

Observation 945767c7-a8f1-4104-888d-16bc9aaa08bf · outbound

This paper cites Lion, K.; Hübler, F.; Li, B.; Orvieto, A.; and He, N.

When Does Muon Help Agentic Reinforcement Learning? Lion, K.; Hübler, F.; Li, B.; Orvieto, A.; and He, N

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T04:19:38.180888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:19:38.180888Z digest=sha256:ca7bf37e173a6b6150baa6b5899decacd6ccb35f77341f4c69d6724ee52b46e0

Observation 5360b176-75a2-4f1a-99f7-d08c326dad6e · outbound

This paper cites Muown: Row-Norm Control for Muon Optimization.

When Does Muon Help Agentic Reinforcement Learning? Muown: Row-Norm Control for Muon Optimization

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T04:19:38.316139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:19:38.316139Z digest=sha256:7c7ab9bc2507d2ac9fdca9ebc1bf4e6a9f39d34a71c885c9799d7317d532a25e

Observation 31f5b1af-31b0-4aac-9ee3-6833d9bf061e · outbound

This paper cites Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less.

When Does Muon Help Agentic Reinforcement Learning? Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T04:19:38.424519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:19:38.424519Z digest=sha256:329b50110475abbe9eb508faa6d5ad7d9a8b28f859cbf74c669741303051ee08

Observation 53346f2d-09e9-4009-bfcd-963fb0f8f16b · outbound

This paper cites Muon$^2$: Boosting Muon via Adaptive Second-Moment Preconditioning.

When Does Muon Help Agentic Reinforcement Learning? Muon$^2$: Boosting Muon via Adaptive Second-Moment Preconditioning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T04:19:38.514523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:19:38.514523Z digest=sha256:98c606542e901c7f7e09b1bf26f1f72d3198320c427a011312bc2ce92926840e

Observation d7029920-6001-4cca-a5b5-91c07e9f138a · outbound

This paper cites Meng, Z.; and Chen, K.

When Does Muon Help Agentic Reinforcement Learning? Meng, Z.; and Chen, K

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T04:19:38.715893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:19:38.715893Z digest=sha256:36612de11988e2a80c42fb5cc983c8f6bf7db35eaade7d2d5e3f042eb69a8aff

Observation ddda07d6-d500-4e11-9822-d809b12448e7 · outbound

This paper cites CRAFT: Counterfactual Credit Assignment from Free Sibling Rollouts for Self-Distilled Agentic Reinforcement Learning.

When Does Muon Help Agentic Reinforcement Learning? CRAFT: Counterfactual Credit Assignment from Free Sibling Rollouts for Self-Distilled Agentic Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T04:19:38.920308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:19:38.920308Z digest=sha256:2cf6a6849ea5fe6bcf9492c9682340332a1678e8f5a7bd4fddebbda58d0d3e14

Observation 9a2e4e88-102b-4c41-97e4-78687ea8ba3b · outbound

This paper cites HTMuon: Improving Muon via Heavy-Tailed Spectral Correction.

When Does Muon Help Agentic Reinforcement Learning? HTMuon: Improving Muon via Heavy-Tailed Spectral Correction

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T04:19:39.069342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:19:39.069342Z digest=sha256:a7d1344478b2a3ca1ab7c9934548f79aa23d6ee9865af607239c83142fd371d8

Observation b1eade41-a668-4039-9b4c-41f2e10aca10 · outbound

This paper cites Qu, X.; Huang, P.; and Horvath, S.

When Does Muon Help Agentic Reinforcement Learning? Qu, X.; Huang, P.; and Horvath, S

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T04:19:39.222709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:19:39.222709Z digest=sha256:7e473498bcbab9d93a9f3a6960b6d08e6e048ed502fb290c8364735829972dd3

Observation 734d5f15-22eb-4cf3-ab53-819cfdfae2a3 · outbound

This paper cites Can Muon Fine-tune Adam-Pretrained Models?.

When Does Muon Help Agentic Reinforcement Learning? Can Muon Fine-tune Adam-Pretrained Models?

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T04:19:39.354599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:19:39.354599Z digest=sha256:e8bc2f5131240e4ab6a8ca5c3e8ae2e1b2b60c4a6ba50200d7e8e270ce02b745

Observation cffaf9cb-68a8-49bc-8848-b2a9d3007c82 · outbound

This paper cites EnvRL: Learn from Environment Dynamics in Agentic Reinforcement Learning.

When Does Muon Help Agentic Reinforcement Learning? EnvRL: Learn from Environment Dynamics in Agentic Reinforcement Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T04:19:39.533735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:19:39.533735Z digest=sha256:9d4c6c24ef39dead118964e3787087fff685899771e15d88cd8280ee953132bc

Observation 7240a8e5-bb55-401b-8e5d-251fa0ed7eac · outbound

This paper cites Semantic Consistency Policy Optimization for Reinforcement Learning of LLM Agents.

When Does Muon Help Agentic Reinforcement Learning? Semantic Consistency Policy Optimization for Reinforcement Learning of LLM Agents

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T04:19:39.560422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:19:39.560422Z digest=sha256:b2eb348087287046e91f4196fa8c651531f99db20f75ef17c97711da0dc2cb75

Observation 9c92c0b1-3975-449c-ad1e-73d115cea13c · outbound

This paper cites StraTA: Incentivizing Agentic Reinforcement Learning with Strategic Trajectory Abstraction.

When Does Muon Help Agentic Reinforcement Learning? StraTA: Incentivizing Agentic Reinforcement Learning with Strategic Trajectory Abstraction

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T04:19:39.591203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:19:39.591203Z digest=sha256:85b86b91c06a0362a3428574ec39d79a3b4548ab3a76b11b385ca0d638ea6954

Observation 1da60b44-fa0d-46f8-af08-cc2b65a7413e · outbound

This paper cites Qwen2.5 Technical Report.

When Does Muon Help Agentic Reinforcement Learning? Qwen2.5 Technical Report

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T04:19:39.628532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:19:39.628532Z digest=sha256:8f612cbdffa2b4613b70375ff7e23c7b5c86d8640196fe1aa321b14fcc7b33b8

Observation a8fccb79-8244-48ba-b63a-72a5f0c9bda9 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

When Does Muon Help Agentic Reinforcement Learning? DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T04:19:39.663587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:19:39.663587Z digest=sha256:b8dfd756e2ca0391a51d9499bae9f0a773830dde58e2386b56931e4877ac647f

Observation e22bd75c-120a-4213-9c03-f5f540262eab · outbound

This paper cites AMO: Adaptive Muon Orthogonalization.

When Does Muon Help Agentic Reinforcement Learning? AMO: Adaptive Muon Orthogonalization

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T04:19:39.717440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:19:39.717440Z digest=sha256:397c46617318f35851d3cc2032029d0766951bb312819a53b2be455296f13550

Observation 9f85e988-4666-48a2-8404-79b41df5d528 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

When Does Muon Help Agentic Reinforcement Learning? DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T04:19:39.483270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:19:39.483270Z digest=sha256:061383b05049f11ad08a08c55422db6fce780e5f2376f1565f7f738d3d8594cf

Observation 19c581b5-bdb8-4dea-8192-c779640ef290 · outbound

This paper cites Kimi K2: Open Agentic Intelligence.

When Does Muon Help Agentic Reinforcement Learning? Kimi K2: Open Agentic Intelligence

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T04:19:38.028406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:19:38.028406Z digest=sha256:6d10792801834d3ad3681819b70c88f24145be2c96156e887b692d5bbfa95957

Observation bbe6a394-062b-4864-a496-6a5b3836f254 · outbound

This paper cites Rethinking Muon Beyond Pretraining: Spectral Failures and High-Pass Remedies for VLA and RLVR.

When Does Muon Help Agentic Reinforcement Learning? Rethinking Muon Beyond Pretraining: Spectral Failures and High-Pass Remedies for VLA and RLVR

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-04T04:19:37.465048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:19:37.465048Z digest=sha256:749d1cad6ba4430627aa41ad8811e740fc02c3ca62e45c5a03b9e7c1519e585e

Pith citing papers

No inbound Pith citation observations are available.