Pith. sign in

Paper Citation Record · LEDGER

Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning

As of 22 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2608.02149.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.02149 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T13:44:42.889956Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

31 of 31 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0d5209f4-e2fb-458e-992f-b0a770131b5b · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T13:44:40.090481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:44:40.090481Z digest=sha256:5c974f4021b0b1fe2a67fe8dd381e9edd1f31834775ff107db8dc1de7dfdc945

Observation 6d2b32d9-87b1-4bba-84ed-fc3f4ef25269 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning Advances in Neural Information Processing Systems , volume=

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T13:44:40.125359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:44:40.125359Z digest=sha256:9f773cbef1d1ce45b30341a9e928d88520a2f4481154d506304f1c16c53fd9a9

Observation 53186b75-65a6-4686-bc35-4da457bbb35e · outbound

This paper cites Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models.

Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T13:44:40.174330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:44:40.174330Z digest=sha256:d786b47c774d3017fdd46df9fbfde17695837ff1602ab32a92115a6fab3e03f6

Observation 90afc498-6351-41c4-9777-6ca0d0fdedbd · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning Advances in Neural Information Processing Systems , volume=

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T13:44:40.218780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:44:40.218780Z digest=sha256:50d067f1c766afc523c7e444273d6717206978babd35b8022fa9a3fcf16c1c38

Observation 323d704e-600f-4940-865d-198f5a908e14 · outbound

This paper cites Beyond the Sampled Token: Preserving Candidate Support in RLVR.

Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning Beyond the Sampled Token: Preserving Candidate Support in RLVR

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T13:44:40.295338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:44:40.295338Z digest=sha256:92aabb821f5be38dbd52cd9a973d22324b052770eb25239d52b53095461fecd2

Observation 729733ae-9987-4e8c-aca0-688845e6a566 · outbound

This paper cites Machine learning , volume=.

Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning Machine learning , volume=

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T13:44:40.372508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:44:40.372508Z digest=sha256:8bf6deea6c5fc28d8c5d690d891e50e09bb04091a3f9a55bd26140322b647ff6

Observation ce21e18e-e14c-41af-a3f4-b49e12f096d5 · outbound

This paper cites arXiv preprint arXiv:2602.02710 , year=.

Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning arXiv preprint arXiv:2602.02710 , year=

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T13:44:40.376806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:44:40.376806Z digest=sha256:07352ef878d3f7906016e81bd9fc7ed1145b05110b8b0f0f5fa3d4b59e68ff3f

Observation 8ab93098-64be-4742-aa57-a9f4fd05c4fd · outbound

This paper cites Proximal Policy Optimization Algorithms.

Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning Proximal Policy Optimization Algorithms

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T13:44:40.380372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:44:40.380372Z digest=sha256:c8727af0adeb4c79d40949273071baa1e49ad471cd884d0cfa997325e504b32e

Observation 1dbc430a-0bee-4cfc-8b31-1a00ca4d3f7d · outbound

This paper cites Transactions of the American Mathematical Society , volume=.

Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning Transactions of the American Mathematical Society , volume=

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T13:44:40.424271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:44:40.424271Z digest=sha256:13d761cc3b1247a8658bcff44b436549cf6c672b74bc5ce6a91547fb3b1458b5

Observation a75913b7-3fc9-4ce1-ba32-f567b45c445b · outbound

This paper cites Statistics & probability letters , volume=.

Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning Statistics & probability letters , volume=

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T13:44:40.546665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:44:40.546665Z digest=sha256:6636644163a2c34175935884ea144159491c1f2de58cc8ad48d3f5a2919fd62f

Observation 867b1ec8-7068-4172-a7ee-f67adb1efc2b · outbound

This paper cites Canadian Journal of Mathematics , volume=.

Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning Canadian Journal of Mathematics , volume=

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T13:44:40.655152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:44:40.655152Z digest=sha256:0c0449fa4cc232973a0436e356b16b6ee7b6b00abe0057d6ebacf96a9d531c33

Observation 9f0d7f77-bff0-4442-9fec-17be97ef1551 · outbound

This paper cites arXiv preprint arXiv:2601.18779 , year=.

Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning arXiv preprint arXiv:2601.18779 , year=

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T13:44:40.813207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:44:40.813207Z digest=sha256:18f16bd93c0525fb19bf24d0aba794e300017ee175dfc1d3a9d22a4726fc2d0b

Observation 451ef461-c10a-47a3-8426-8130383d516c · outbound

This paper cites arXiv preprint arXiv:2602.21189 , year=.

Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning arXiv preprint arXiv:2602.21189 , year=

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T13:44:40.932110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:44:40.932110Z digest=sha256:4feb98db5f36f81c229c17dc521aeafe94a4721f8947c5902a150d57c6fd3eb9

Observation 22a0bc3d-0cdd-49bb-a5aa-83258e00ca71 · outbound

This paper cites Qwen3 Technical Report.

Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning Qwen3 Technical Report

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T13:44:41.090156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:44:41.090156Z digest=sha256:d2d817dbd8193cb85f5cf88d9d1f5dc5e343bdb2c72bd510c97baa57fe46773c

Observation 511f50f6-471c-4206-bb31-8a581321acdf · outbound

This paper cites Proceedings of the Twentieth European Conference on Computer Systems , pages=.

Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning Proceedings of the Twentieth European Conference on Computer Systems , pages=

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T13:44:41.231524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:44:41.231524Z digest=sha256:d5dcb615fe9ed6fd6793b685f7f13909a512d3831bc5afb3ef874c368c3965b1

Observation ed68086f-f966-48d3-8a3f-f15ffe3fafd5 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning Measuring Mathematical Problem Solving With the MATH Dataset

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T13:44:41.310367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:44:41.310367Z digest=sha256:c3583342cd010fd0822e531ad8cc752d43d17c537090643184b4d024f3901c8b

Observation c376b6d3-0b92-41b0-a44d-bc5e95ab786f · outbound

This paper cites 2025 , howpublished =.

Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning 2025 , howpublished =

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T13:44:41.515791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:44:41.515791Z digest=sha256:ce24ff901999294e9b918730d03c98d8fc9cdd63447c298b1c6a538a722b5e03

Observation 50d1c619-e285-4906-a690-0e82a7612cc1 · outbound

This paper cites Beyond Mode Collapse: Distribution Matching for Diverse Reasoning.

Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning Beyond Mode Collapse: Distribution Matching for Diverse Reasoning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T13:44:41.612646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:44:41.612646Z digest=sha256:b425c9a05ecd602d2a547ea23c9fd5c7d2b133af946f0d26983c551eea45f2e6

Observation e18ef7a0-5499-4bff-8516-6b38465b8313 · outbound

This paper cites Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T13:44:41.707471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:44:41.707471Z digest=sha256:ad5161e214fff30a7e02ede083135396f9272f3a90eff50a9b4069b7ff10ad46

Observation eb3425a5-3729-4f41-a508-4b9dfa134d56 · outbound

This paper cites Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs.

Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T13:44:41.794435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:44:41.794435Z digest=sha256:327811862df75111620bc0f88bb4251a6c9a1d977738d4959c56e28dc79e7f6e

Observation dd81640a-f32f-4349-8699-8fd4c34fdef8 · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T13:44:41.925937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:44:41.925937Z digest=sha256:01b4cf70bfccbd56d659d56de4e91b2443917a44452a8053973a64859cbd5b34

Observation 76bb9b35-0e89-426d-aa2d-8445e1964eca · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T13:44:42.094603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:44:42.094603Z digest=sha256:cdc6d728e8dac072e53802d7e3bdc1e1dfb58cc89179773aa5c0d3984048f323

Observation 601c01fc-f827-46be-80ab-f4ef28a8bd95 · outbound

This paper cites Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents.

Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T13:44:42.254343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:44:42.254343Z digest=sha256:72cad8ce928c7db16ff6d673d7ee95eb35ee08722ea4f891efcfe9210c4bfb98

Observation c5e6ef5a-aab2-45de-8d24-68d3f06d044f · outbound

This paper cites arXiv preprint arXiv:2509.25133 , year=.

Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning arXiv preprint arXiv:2509.25133 , year=

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T13:44:42.441366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:44:42.441366Z digest=sha256:8f532a86e97a419f1f5fd65dcae3f1ee55f6ae24569606e0695d17d3922fa36c

Observation 1616428c-6970-468b-a01c-d481697f6cfd · outbound

This paper cites Findings of the Association for Computational Linguistics: ACL 2026 , pages=.

Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning Findings of the Association for Computational Linguistics: ACL 2026 , pages=

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T13:44:42.515466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:44:42.515466Z digest=sha256:0d7af504372a691d65b2a9ed4d7917cbc3951649807f0d45b5f470f5d7bb9cb1

Observation 33f4a1b7-835d-41e6-a3a1-1ad31fd9f8fa · outbound

This paper cites arXiv preprint arXiv:2509.15207 , year=.

Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning arXiv preprint arXiv:2509.15207 , year=

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T13:44:42.573566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:44:42.573566Z digest=sha256:6ffa0ef8fc5cfa0b10e5c9fce99cde9cf2d2c79beefd7bca09022bdbf3a0cd8a

Observation ab809372-8108-4b96-a612-56dd41d4016c · outbound

This paper cites arXiv preprint arXiv:2509.26209 , year=.

Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning arXiv preprint arXiv:2509.26209 , year=

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T13:44:42.659322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:44:42.659322Z digest=sha256:b68ef5b921934339607ae94e1e00d190484b63cf92da83f5dbe3d750ddd826af

Observation f465618d-2642-4c03-99f4-85b6e8c5f630 · outbound

This paper cites TreePO: Bridging the Gap of Policy Optimization and Efficacy and Inference Efficiency with Heuristic Tree-based Modeling.

Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning TreePO: Bridging the Gap of Policy Optimization and Efficacy and Inference Efficiency with Heuristic Tree-based Modeling

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T13:44:42.721995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:44:42.721995Z digest=sha256:d3c15ef7509964de4f219d401e4590942428a742705f953581ac3046db1c0dbc

Observation 74c03c94-b637-4681-8f32-77dc7cceaec1 · outbound

This paper cites Springer Series in Statistics ( , year=.

Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning Springer Series in Statistics ( , year=

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T13:44:42.764654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:44:42.764654Z digest=sha256:426fe7705df304941ab737fa8b723ab977e244092c63461e0cf5c54f8aaa5e18

Observation d97f1f9c-884e-43d8-b2bb-2ee87fa170d5 · outbound

This paper cites International conference on machine learning , pages=.

Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning International conference on machine learning , pages=

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T13:44:42.817011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:44:42.817011Z digest=sha256:07997cde82b60f09c5b232bcef9b6e74645eddf6a2661b725be130ce93184d5e

Observation d3f7cfcf-88a9-4fe6-9cbb-5f4d42b469ec · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T13:44:42.889956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:44:42.889956Z digest=sha256:b601aad16150976f03f9be9b29c8daeb174be141eebd17d3c7bf5b0810f9db9b

Pith citing papers

No inbound Pith citation observations are available.