Pith. sign in

Paper Citation Record · LEDGER

Turning Off-Policy Tokens On-Policy: A Plug-in Approach for Improving LLM Alignment

As of 14 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2607.04728.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.04728 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-11T14:30:20.959431Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

31 of 31 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a06f6d39-a2dd-4355-ac07-fd68f37a7eb4 · outbound

This paper cites MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention.

Turning Off-Policy Tokens On-Policy: A Plug-in Approach for Improving LLM Alignment MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-11T14:30:20.959431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:30:20.959431Z digest=sha256:5a3c334ae475535558d2d7b54bdef47c644bff243fad4315b2f6342529bb2b27

Observation 3c1e416c-a0c4-4052-a3bd-b5380f63b673 · outbound

This paper cites The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models.

Turning Off-Policy Tokens On-Policy: A Plug-in Approach for Improving LLM Alignment The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-11T14:30:20.959431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:30:20.959431Z digest=sha256:dc796eaeec47b7f7287ac4cb02f658c93152d823b7c11309b2aba26ebf869edd

Observation 273e1317-374f-4713-9a8f-598c35730718 · outbound

This paper cites Soft Adaptive Policy Optimization.

Turning Off-Policy Tokens On-Policy: A Plug-in Approach for Improving LLM Alignment Soft Adaptive Policy Optimization

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-11T14:30:20.959431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:30:20.959431Z digest=sha256:4d1a740a24cc04387436bd033b2cf7f5bf36172e5c51b77c64b7c34b048104f8

Observation 5893bd3b-9568-4881-ab4e-d9ba90796236 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Turning Off-Policy Tokens On-Policy: A Plug-in Approach for Improving LLM Alignment DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-11T14:30:20.959431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:30:20.959431Z digest=sha256:e9e2faaa03f00aadf47fd686710cb87387d9af12210a1438a174426852ae1150

Observation 401e2898-b4b6-464c-8117-91fcc19a8df3 · outbound

This paper cites CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment.

Turning Off-Policy Tokens On-Policy: A Plug-in Approach for Improving LLM Alignment CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-11T14:30:20.959431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:30:20.959431Z digest=sha256:66f6a139bc4d521f04874a9b6600bcaafddf0ff9211b0909a144ef76d92a2b5f

Observation b7449652-6eb0-425e-845c-9361c74c850b · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

Turning Off-Policy Tokens On-Policy: A Plug-in Approach for Improving LLM Alignment Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-11T14:30:20.959431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:30:20.959431Z digest=sha256:80dadc7f42d28721099086d8d6a6b85a5f9daa2efa337d4e7126f84616492201

Observation 18a73839-087e-4128-9a07-439a972ec5c8 · outbound

This paper cites Dense passage retrieval for open-domain question answering.

Turning Off-Policy Tokens On-Policy: A Plug-in Approach for Improving LLM Alignment Dense passage retrieval for open-domain question answering

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-11T14:30:20.959431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:30:20.959431Z digest=sha256:6d7c90792534b40fd6fecc8cd9a5802ef749042c2a9937cdd354a790135a1ccd

Observation 9284340e-1586-48fc-89bd-4f886c07094d · outbound

This paper cites A step back: Prefix importance ratio stabilizes policy optimization.

Turning Off-Policy Tokens On-Policy: A Plug-in Approach for Improving LLM Alignment A step back: Prefix importance ratio stabilizes policy optimization

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-11T14:30:20.959431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:30:20.959431Z digest=sha256:e03f1c25a1186414936e2985e55ed775393f4abfccdc8eeef8a23be1e5714585

Observation 654a09ac-c475-40c3-9571-0ff876ad20da · outbound

This paper cites Trust Region Masking for Long-Horizon LLM Reinforcement Learning.

Turning Off-Policy Tokens On-Policy: A Plug-in Approach for Improving LLM Alignment Trust Region Masking for Long-Horizon LLM Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-11T14:30:20.959431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:30:20.959431Z digest=sha256:f74213e9b738f60cbf5494f66cccf048681763cb6f8e9a2a2dfc5c820297e2df

Observation b2187083-7152-45e6-b97d-b56bcbd30d48 · outbound

This paper cites Reasoning and Tool-use Compete in Agentic RL:From Quantifying Interference to Disentangled Tuning.

Turning Off-Policy Tokens On-Policy: A Plug-in Approach for Improving LLM Alignment Reasoning and Tool-use Compete in Agentic RL:From Quantifying Interference to Disentangled Tuning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-11T14:30:20.959431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:30:20.959431Z digest=sha256:1bd5ba09b258dd5e0fd1f85d6325bc152206bdb4648049ed7f07278dc1bd8b6a

Observation 8da5dd53-26a6-43f2-9975-071ed232866d · outbound

This paper cites Sparse-rl: Breaking the memory wall in llm reinforcement learning via stable sparse rollouts.

Turning Off-Policy Tokens On-Policy: A Plug-in Approach for Improving LLM Alignment Sparse-rl: Breaking the memory wall in llm reinforcement learning via stable sparse rollouts

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-11T14:30:20.959431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:30:20.959431Z digest=sha256:5bf338817b5b174154183abe194b9a6c5c88a1200750744b423e769c5b48a8ed

Observation 966550d4-8039-4999-9ce3-377c0d88fc17 · outbound

This paper cites Stabilizing moe reinforcement learning by aligning training and inference routers.

Turning Off-Policy Tokens On-Policy: A Plug-in Approach for Improving LLM Alignment Stabilizing moe reinforcement learning by aligning training and inference routers

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-11T14:30:20.959431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:30:20.959431Z digest=sha256:abe3361ea2d4cbae0730afe9415eee85f33808485fae3493ef86c92021f6cef2

Observation 14d476a9-6048-4777-8f0c-4d58f9306b4d · outbound

This paper cites When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories.

Turning Off-Policy Tokens On-Policy: A Plug-in Approach for Improving LLM Alignment When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-11T14:30:20.959431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:30:20.959431Z digest=sha256:5fe932ad23db97fa93a992c8e5fba3ef985c7833ed9d2f70d96f79fcc6257e33

Observation 3493a053-8c96-43b8-b22b-37a0f47598f3 · outbound

This paper cites Measuring and narrowing the compositionality gap in language models.

Turning Off-Policy Tokens On-Policy: A Plug-in Approach for Improving LLM Alignment Measuring and narrowing the compositionality gap in language models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-11T14:30:20.959431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:30:20.959431Z digest=sha256:561f698c669b911667861e441adcfd511977afdc3157036e5f7c4c53eb3ae81a

Observation 5f545beb-a004-4881-8812-af5042f968b1 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Turning Off-Policy Tokens On-Policy: A Plug-in Approach for Improving LLM Alignment Proximal Policy Optimization Algorithms

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-11T14:30:20.959431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:30:20.959431Z digest=sha256:94a86d5bf14277f73e8ed76cc9c93fb28a5c3ccecdc215dc63374303693d2f6d

Observation 7a0481f4-1ce8-43fe-b90b-75ad9b522765 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Turning Off-Policy Tokens On-Policy: A Plug-in Approach for Improving LLM Alignment DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-11T14:30:20.959431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:30:20.959431Z digest=sha256:72e3fc8562d357f21e92c57208abc19391cf2901174e3c7b076a3b78981db5fa

Observation 5313c16c-70fc-4fa9-9043-601d12e9eda1 · outbound

This paper cites Maximum likelihood reinforcement learning.

Turning Off-Policy Tokens On-Policy: A Plug-in Approach for Improving LLM Alignment Maximum likelihood reinforcement learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-11T14:30:20.959431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:30:20.959431Z digest=sha256:7b0a9fe14ce0670935c35ed722098f6fb8cbd547109b0a2e19df0c85a55fecbd

Observation 38afff03-0a3d-44dc-bed8-9dc255baf7c6 · outbound

This paper cites Every step evolves: Scaling reinforcement learning for trillion-scale thinking model.

Turning Off-Policy Tokens On-Policy: A Plug-in Approach for Improving LLM Alignment Every step evolves: Scaling reinforcement learning for trillion-scale thinking model

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-11T14:30:20.959431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:30:20.959431Z digest=sha256:240deb0b02b628fe8811480b129d2c36ade1598707250c5bd294bdeddd4ab7da

Observation 8c595e6b-63bb-4269-b8e1-862c8c53534b · outbound

This paper cites UCOB: Learning to Utilize and Evolve Agentic Skills via Credit-Aware On-Policy Bidirectional Self-Distillation.

Turning Off-Policy Tokens On-Policy: A Plug-in Approach for Improving LLM Alignment UCOB: Learning to Utilize and Evolve Agentic Skills via Credit-Aware On-Policy Bidirectional Self-Distillation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-11T14:30:20.959431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:30:20.959431Z digest=sha256:019dee135715c00038a7e70b3f7937125db5be35faa3ae73e8245804bdbdc5a3

Observation 919c82d0-d2ae-4405-a3b5-08700456d47f · outbound

This paper cites Text Embeddings by Weakly-Supervised Contrastive Pre-training.

Turning Off-Policy Tokens On-Policy: A Plug-in Approach for Improving LLM Alignment Text Embeddings by Weakly-Supervised Contrastive Pre-training

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-11T14:30:20.959431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:30:20.959431Z digest=sha256:42bd53fa0eb595f1202fb09cbf10e937329498885423ca1b62b7afe16bf1d9b3

Observation ab5248c1-5f39-4517-8972-f24a69e3cd93 · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

Turning Off-Policy Tokens On-Policy: A Plug-in Approach for Improving LLM Alignment Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-11T14:30:20.959431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:30:20.959431Z digest=sha256:4299628150483eb2058d484a788ab57384de2de1c6082f11b52f481164cd45c4

Observation b01057bc-de5b-4b1e-a26e-000787aa8b3b · outbound

This paper cites Qwen3 Technical Report.

Turning Off-Policy Tokens On-Policy: A Plug-in Approach for Improving LLM Alignment Qwen3 Technical Report

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-11T14:30:20.959431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:30:20.959431Z digest=sha256:50dafb7fc595e85210b1119065e7c630597b2e1fd7409e5b8299b7e40b779b5a

Observation acdb543a-83ee-4fcb-99bc-792c49b9609a · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Turning Off-Policy Tokens On-Policy: A Plug-in Approach for Improving LLM Alignment DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-11T14:30:20.959431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:30:20.959431Z digest=sha256:a2d789b72c6b3a0c312afae13ce4cb47a8ec8d64822aeadb385c40cf5a7a32f3

Observation 0884fa46-92f7-4a1a-8e31-ddf7993ad560 · outbound

This paper cites HTAM: Hierarchical Transition-Attended Memory for Operator Optimization.

Turning Off-Policy Tokens On-Policy: A Plug-in Approach for Improving LLM Alignment HTAM: Hierarchical Transition-Attended Memory for Operator Optimization

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-11T14:30:20.959431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:30:20.959431Z digest=sha256:4676edcc31af6de4bb0b47ff4bd3b8d3525d7474f32f16641c01954201ee83ca

Observation e61a3fb6-412d-4014-b12c-9fc26347f03b · outbound

This paper cites PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel.

Turning Off-Policy Tokens On-Policy: A Plug-in Approach for Improving LLM Alignment PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-11T14:30:20.959431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:30:20.959431Z digest=sha256:afc85932adead87241cbce0cd2bf9194b433378ae963bb7ffc96df8cf1df56cc

Observation 4dee2c79-c12f-4200-8af0-9f02725ea7f1 · outbound

This paper cites Stabilizing reinforcement learning with llms: For- mulation and practices.

Turning Off-Policy Tokens On-Policy: A Plug-in Approach for Improving LLM Alignment Stabilizing reinforcement learning with llms: For- mulation and practices

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-11T14:30:20.959431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:30:20.959431Z digest=sha256:1c0f6f5016a35e3810755afcfdcef2225505dea77e5caffa892bbad44e0c070e

Observation cafb0fc6-8c8c-44af-8bce-8e71a9cdd387 · outbound

This paper cites an unresolved cited work.

Turning Off-Policy Tokens On-Policy: A Plug-in Approach for Improving LLM Alignment Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-11T14:30:20.959431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:30:20.959431Z digest=sha256:6f6c8c3ab1e1eb9a12f24139b66cfd5db19a8280d86c411f43512f12f25b7d96

Observation b2b2833c-1eec-46c7-a7b0-59f3de53656c · outbound

This paper cites an unresolved cited work.

Turning Off-Policy Tokens On-Policy: A Plug-in Approach for Improving LLM Alignment Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-11T14:30:20.959431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:30:20.959431Z digest=sha256:3961441bcbce8ccee518e9bd09f2ea5a2e7dbe7d95655cbbe15120ee068781b1

Observation b32333c9-0f0b-4a8e-b95d-a2dfe061f2bc · outbound

This paper cites Table 4: Hyperparameters used in math experi- ments.

Turning Off-Policy Tokens On-Policy: A Plug-in Approach for Improving LLM Alignment Table 4: Hyperparameters used in math experi- ments

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-11T14:30:20.959431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:30:20.959431Z digest=sha256:30edc2457f8db260af865ba04fe632e7fea699ad0c83f8292481a855e6e791f2

Observation 4e5ad6d9-0f95-4f0a-9bb4-6e3ff8f0b025 · outbound

This paper cites an unresolved cited work.

Turning Off-Policy Tokens On-Policy: A Plug-in Approach for Improving LLM Alignment Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-11T14:30:20.959431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:30:20.959431Z digest=sha256:ca41de63507cde4376922cca62c6088fc1c0cf63c0a309cd0b8307561ed39bb0

Observation e8ee70f5-9673-485d-917a-f0db747c5ddc · outbound

This paper cites Table 8: Performance of SIS under varying policy staleness on Qwen3-8B-Base, where stalenessN is the number of mini-batch updates per rollout.

Turning Off-Policy Tokens On-Policy: A Plug-in Approach for Improving LLM Alignment Table 8: Performance of SIS under varying policy staleness on Qwen3-8B-Base, where stalenessN is the number of mini-batch updates per rollout

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-11T14:30:20.959431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:30:20.959431Z digest=sha256:07412461daf4f907d8e31c9581d4a88b8d166565aba45fa338210373e5972a6f

Pith citing papers

No inbound Pith citation observations are available.