Pith. sign in

Paper Citation Record · LEDGER

Predictive Divergence Masks for LLM RL

As of 22 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2607.10848.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.10848 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-14T08:49:37.615924Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 68551c4e-dc65-4aed-bdb1-f151fabf8187 · outbound

This paper cites Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs.

Predictive Divergence Masks for LLM RL Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-14T08:49:37.615924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T08:49:37.615924Z digest=sha256:6335fdbb2c4645fd9b34bc1ea46858dbda1c71dbfa0d463ad3c20825c8186e34

Observation 52ac8a75-1c16-4bfe-a1ca-1fa17084e47c · outbound

This paper cites MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention.

Predictive Divergence Masks for LLM RL MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-14T08:49:37.615924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T08:49:37.615924Z digest=sha256:3121b43f3586b3afc9d9d67ae3afe7e1050b432a6a5272b8243b6a339553bec0

Observation 906cac7a-da9c-4681-b459-2d42a2f87604 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Predictive Divergence Masks for LLM RL DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-14T08:49:37.615924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T08:49:37.615924Z digest=sha256:7e3c233e327e6632784c1f8a953b3bfc11c10bbcf362e7a16f2b825d9814bc74

Observation a6deee50-f3d8-42a3-8e35-22a2bb1d132e · outbound

This paper cites https://thinkingmachines.ai/blog/defeating-nondeterminism-in- llm-inference/.

Predictive Divergence Masks for LLM RL https://thinkingmachines.ai/blog/defeating-nondeterminism-in- llm-inference/

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-14T08:49:37.615924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T08:49:37.615924Z digest=sha256:d271bd01d6b1f831d2f1d1ed13e224094a1c48318ac22133b47a86a838f28518

Observation f6c6f435-b655-48f1-9fec-eb2b48ee8b53 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

Predictive Divergence Masks for LLM RL Understanding R1-Zero-Like Training: A Critical Perspective

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-14T08:49:37.615924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T08:49:37.615924Z digest=sha256:871e6f4b9349a0554575bfe72b4666c59ddb97deeed8eb5eb430dff804a4561e

Observation 0eb8cef3-6fd1-42ae-b224-2feaf23564c4 · outbound

This paper cites Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning.

Predictive Divergence Masks for LLM RL Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-14T08:49:37.615924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T08:49:37.615924Z digest=sha256:b0852ce7e020974fde30f27cc3f24427af538010f206e1f8169e25220147cdc4

Observation 70c6814d-7b83-431a-8c00-a43aeddf6ae4 · outbound

This paper cites Defeating the training-inference mismatch via fp16.arXiv preprint arXiv:2510.26788,.

Predictive Divergence Masks for LLM RL Defeating the training-inference mismatch via fp16.arXiv preprint arXiv:2510.26788,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-14T08:49:37.615924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T08:49:37.615924Z digest=sha256:41599efc22852d185aa5df17bedd2ddf4aa86ce7ed84f719c6f86da8badde6e7

Observation f0215ada-3e65-4c29-8cd0-763b780cc13d · outbound

This paper cites Rethinking the Trust Region in LLM Reinforcement Learning.

Predictive Divergence Masks for LLM RL Rethinking the Trust Region in LLM Reinforcement Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-14T08:49:37.615924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T08:49:37.615924Z digest=sha256:0b446383f47e269ad6c2ef7c4e1f838132b7f49c348d2dd93415d4e080d7e6aa

Observation a31a58da-a056-41e9-8ed3-7d959e170055 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Predictive Divergence Masks for LLM RL Proximal Policy Optimization Algorithms

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-14T08:49:37.615924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T08:49:37.615924Z digest=sha256:6ea4bc8445ccf709e2bb01b84b9f8544bb7066a21758002bf462919d0b0f7f16

Observation 9b62590e-aeb5-47dc-8912-7f00e983fdd1 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Predictive Divergence Masks for LLM RL DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-14T08:49:37.615924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T08:49:37.615924Z digest=sha256:2c03b96282e94126c2a6498b6008dcf7972c0dd39e5783f78054a7211f76b78f

Observation 5423e8eb-5d03-4520-bb39-c284be433c82 · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

Predictive Divergence Masks for LLM RL HybridFlow: A Flexible and Efficient RLHF Framework

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-14T08:49:37.615924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T08:49:37.615924Z digest=sha256:0d86d8045312cf9f3bf2ace78451f834c07d5297a5c30890c78cc2f0d0d5f495

Observation 7fa69b9c-a5e3-4df8-9ea9-5222a7380d00 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

Predictive Divergence Masks for LLM RL Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-14T08:49:37.615924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T08:49:37.615924Z digest=sha256:8afa7dbc9a1c8516b9553145dfe328b345a70832359445b931889989bb2ec594

Observation d28266e3-e8ba-4885-91e6-b672dc761455 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Predictive Divergence Masks for LLM RL Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-14T08:49:37.615924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T08:49:37.615924Z digest=sha256:e26a75c5e99915c6b618e6e7dadeeadff0493408b5254fe73d6b196f4e4cc5b7

Observation ff94ef2b-28cd-49bc-9af9-e63f2b4be31b · outbound

This paper cites Simple Policy Optimization.

Predictive Divergence Masks for LLM RL Simple Policy Optimization

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-14T08:49:37.615924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T08:49:37.615924Z digest=sha256:ace3b29affccf234a6ee8406a63a257bb09ae41985fb3e3d37ba13f7bb19ef6d

Observation 45e72984-97cd-420a-8b54-9113b30ba513 · outbound

This paper cites Qwen3 Technical Report.

Predictive Divergence Masks for LLM RL Qwen3 Technical Report

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-14T08:49:37.615924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T08:49:37.615924Z digest=sha256:c06b1400972de8265f40b413ebf6149d1623a52d81f368a9db0c33410bdbfd5f

Observation 5861436c-563c-4e29-9d3e-c224264a6331 · outbound

This paper cites Rethinking the Divergence Regularization in LLM RL.

Predictive Divergence Masks for LLM RL Rethinking the Divergence Regularization in LLM RL

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-14T08:49:37.615924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T08:49:37.615924Z digest=sha256:778b321bae726285bd0d9b47c2bfc12aca5dd356e1220b5f67d94c39358a5a2f

Observation 335fd98d-8447-453d-adf1-5664ee146f3d · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Predictive Divergence Masks for LLM RL DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-14T08:49:37.615924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T08:49:37.615924Z digest=sha256:b2b40fff3182cb1de1066011f1963bcd061a8e90f8454f5e11bcb0ac873a1795

Observation a80c74fc-c51f-4766-a16c-9c8c9ee4f54c · outbound

This paper cites Stabilizing reinforcement learning with llms: Formulation and practices.arXiv preprint arXiv:2512.01374,.

Predictive Divergence Masks for LLM RL Stabilizing reinforcement learning with llms: Formulation and practices.arXiv preprint arXiv:2512.01374,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-14T08:49:37.615924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T08:49:37.615924Z digest=sha256:790fdd96002c80f97594151752df48e74d009e5735df685bb2f94b6a08857971

Observation 1c040be7-4644-4a71-ad7a-08c77ae88e6e · outbound

This paper cites Reinforcing General Reasoning without Verifiers.

Predictive Divergence Masks for LLM RL Reinforcing General Reasoning without Verifiers

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-14T08:49:37.615924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T08:49:37.615924Z digest=sha256:1d0ffc8ba003cb74109200b9a607c0c871813a2769b32ed70a6d3d27923b72a5

Observation d701038e-e666-4992-8d4a-dd92dcb8800c · outbound

This paper cites an unresolved cited work.

Predictive Divergence Masks for LLM RL Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-14T08:49:37.615924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T08:49:37.615924Z digest=sha256:779f0077bfbb0a894772cbb28ee9594e62e98791a05d0f5cbaaf6cc5e8401287

Observation 722c673d-87a4-41cb-9887-8a4da861ba9a · outbound

This paper cites GRPO (Shao et al.,.

Predictive Divergence Masks for LLM RL GRPO (Shao et al.,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-14T08:49:37.615924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T08:49:37.615924Z digest=sha256:4ec9483a1069a43b1d0a17c9fb9b3173da09b006c238c867b8ee1ab96a2185b6

Observation 6cfebd66-a1dd-4f1a-b6ab-4a7422b3d5ea · outbound

This paper cites an unresolved cited work.

Predictive Divergence Masks for LLM RL Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-14T08:49:37.615924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T08:49:37.615924Z digest=sha256:b9d4a09e2dfc42475e5a6bbc93375b10a0ced7116bc3340e822d3e1306fb9291

Observation 7dd4f3d8-062a-4429-92f4-e9cca86232e2 · outbound

This paper cites These methods modify the objective or the clipping rule, but they still decide the update direction from the sampled importance ratio.

Predictive Divergence Masks for LLM RL These methods modify the objective or the clipping rule, but they still decide the update direction from the sampled importance ratio

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-14T08:49:37.615924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T08:49:37.615924Z digest=sha256:28bd39694d0deea17febb1ac019dc2a0d4f005a667a55e3f3077c708995fb97c

Observation 0bb1750d-f617-4116-8345-fc7dadec84cf · outbound

This paper cites (5) that we build on.

Predictive Divergence Masks for LLM RL (5) that we build on

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-14T08:49:37.615924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T08:49:37.615924Z digest=sha256:8180fa719f0bba8b5ac0361ef7a0a8b30b5151a0a2c4a12abb609ee6e27cdf0d

Observation 38072ebf-d0b6-44c5-bc05-9d4b3b38beb3 · outbound

This paper cites Aggregated-tail estimator.Under the top-K aggregated-tail construction, the support contains the retained tokensK and one tail bucket.

Predictive Divergence Masks for LLM RL Aggregated-tail estimator.Under the top-K aggregated-tail construction, the support contains the retained tokensK and one tail bucket

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-14T08:49:37.615924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T08:49:37.615924Z digest=sha256:526715ba227eb26aef8d4f9d6fc6969aee285a763c0ac915986a205c6eb4f6e4

Observation 7ee75953-0c84-48f1-b679-d164f04cf47b · outbound

This paper cites By default, both training and rollout use BF16.

Predictive Divergence Masks for LLM RL By default, both training and rollout use BF16

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-14T08:49:37.615924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T08:49:37.615924Z digest=sha256:fbbb03b0dbd6dd4875ea7a4a91da4b1044b54132353503ecf61f9ae1a97c7949

Observation a6836379-670c-4122-b8d1-2462534f9de2 · outbound

This paper cites an unresolved cited work.

Predictive Divergence Masks for LLM RL Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-14T08:49:37.615924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T08:49:37.615924Z digest=sha256:34a1a09605d4f5e6abb4ebee27897a6e322ceececd435f6476c213bf6c21feec

Pith citing papers

No inbound Pith citation observations are available.