Pith. sign in

Paper Citation Record · LEDGER

Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning

As of 4 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 1 inbound Pith citation observation for arXiv:2606.10968.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.10968 v2

Coverage vector

measured 33 of 33 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-27T13:40:39.598985Z

measured 34 of 34 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-14T08:49:37.615924Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

33 of 33 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved24
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch8

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation da798e84-75d3-4c64-ab8b-74fa70eaf887 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T04:47:37.876261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-27T13:40:39.598985Z digest=sha256:757efd5ab5564576365334f53d11ecf995f7b37632dc902f757c38d20b05c53c

Observation 749ea2ef-3962-42c1-a1be-13b71d402f2b · outbound

This paper cites International conference on machine learning , pages=.

Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning International conference on machine learning , pages=

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-27T13:40:39.598985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T13:40:39.598985Z digest=sha256:17c6b8d861c8715e66ded4a33e2399744ed9cdd0a8a8b2282b25d90934485ba2

Observation 7641eaf3-7e1e-49e6-81d9-7b7b932eb043 · outbound

This paper cites Proceedings of the nineteenth international conference on machine learning , pages=.

Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning Proceedings of the nineteenth international conference on machine learning , pages=

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-27T13:40:39.598985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T13:40:39.598985Z digest=sha256:992c425df4372b5752e1cb9ef528999219bb793b94aeaaba23da19346d0019bc

Observation b4e07276-ef90-4037-86e7-8af3931f4821 · outbound

This paper cites International conference on machine learning , pages=.

Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning International conference on machine learning , pages=

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-27T13:40:39.598985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T13:40:39.598985Z digest=sha256:b214ef49a9f097193e47c915640ceca539c63c3dda40418b9a26b3e095680fbf

Observation c9611386-c5dd-4155-af39-025814280182 · outbound

This paper cites Proceedings of the AAAI Conference on Artificial Intelligence , volume=.

Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning Proceedings of the AAAI Conference on Artificial Intelligence , volume=

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-27T13:40:39.598985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T13:40:39.598985Z digest=sha256:b410d9b04cbb6ba52629bab77d7d5e998b0f5efb0372a8a135e644332a6a7ff7

Observation 1c75c9dc-d30b-48e6-bfa1-6ba694efd28b · outbound

This paper cites Riedmiller , title=.

Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning Riedmiller , title=

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-27T13:40:39.598985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T13:40:39.598985Z digest=sha256:8a2c630c3d733915b3e649d28f45782ea76b5261d354a16689ba832584dc4c8f

Observation 2b8f5a20-3205-4698-a8af-7db4de249624 · outbound

This paper cites Francis Song and Abbas Abdolmaleki and Jost Tobias Springenberg and Aidan Clark and Hubert Soyer and Jack W.

Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning Francis Song and Abbas Abdolmaleki and Jost Tobias Springenberg and Aidan Clark and Hubert Soyer and Jack W

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-27T13:40:39.598985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T13:40:39.598985Z digest=sha256:c9abaf391d03dab05b7e9d143d41e1f0971f4a3157d509568f15afaa8dd7f19c

Observation c612ccec-95a9-4578-8cc1-07d7baa69d94 · outbound

This paper cites CoRR , volume=.

Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning CoRR , volume=

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-27T13:40:39.598985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T13:40:39.598985Z digest=sha256:690cc8aaecc1c16b9c90614a5b13694a14cf4a58c9ba8d4430b145c381bc56d4

Observation 7b893bb6-924a-434c-80a2-6da0f348ec9d · outbound

This paper cites 2019 , cdate=.

Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning 2019 , cdate=

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-27T13:40:39.598985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T13:40:39.598985Z digest=sha256:572f39f2f43ad6d8e680111016866f432dfb7702f046ed00be85269ca58f977d

Observation bceba189-06f8-41bf-890a-d888b595aa6e · outbound

This paper cites International Conference on Learning Representations , year=.

Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning International Conference on Learning Representations , year=

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-27T13:40:39.598985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T13:40:39.598985Z digest=sha256:93fa470a1d43da59c9a3707eebd8f1b04e8f1edb4dfe4c91bbe2b94e15f4aaf1

Observation 7e6b9655-0f05-42d3-a70e-9efa17eea6f1 · outbound

This paper cites International Conference on Learning Representations , year=.

Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning International Conference on Learning Representations , year=

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-27T13:40:39.598985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T13:40:39.598985Z digest=sha256:d832ff8ff422406ddafb81e68d7a987a42ab88412fc4588118f1987458a34b5d

Observation 0f58bc07-fe6a-4a4c-873b-61022426dbad · outbound

This paper cites The annals of mathematical statistics , volume=.

Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning The annals of mathematical statistics , volume=

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-27T13:40:39.598985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T13:40:39.598985Z digest=sha256:4cadd9d1dd8192ad322c15c77ba4b88aa1de5fc3daaf27b73487d99ce5ca428a

Observation 6dd911a5-1551-47b9-87ae-1e79f6199b1e · outbound

This paper cites an unresolved cited work.

Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-27T13:40:39.598985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T13:40:39.598985Z digest=sha256:5971e44b996612f989105e313ab7199368fb9687678a319ec6d603f9db76325a

Observation 1422971e-36f9-4ce6-9813-6b991d84c1d9 · outbound

This paper cites Advances in neural information processing systems , volume=.

Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning Advances in neural information processing systems , volume=

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-27T13:40:39.598985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T13:40:39.598985Z digest=sha256:b2d4ab41a9d33cd1f7c3c12314c7cd55d310e1e879c850f2ddb8a32a51b4d1e8

Observation 14e2c108-c115-4b57-abb2-1ce1f1a51311 · outbound

This paper cites an unresolved cited work.

Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-27T13:40:39.598985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T13:40:39.598985Z digest=sha256:262c25df2381d5f7073ec18293ce911d7751748f6b851347bc0b37639e5a62f0

Observation 204eb0be-40c8-4979-8f94-28b87591052b · outbound

This paper cites an unresolved cited work.

Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-27T13:40:39.598985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T13:40:39.598985Z digest=sha256:6ed2b5de0d90bad8a23be2ade014938f103cc8bc678b40586ccc6de8af8ddf27

Observation 721df89d-fe93-443b-aaf9-e63f58b5f088 · outbound

This paper cites 2026 , url=.

Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning 2026 , url=

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-27T13:40:39.598985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T13:40:39.598985Z digest=sha256:1226463a5e69b733a4b01f356acf86ebedecd0d660854682c0e868f4e2b97f57

Observation 37ab2e50-8c7c-4c37-ad8c-ec3a83a62571 · outbound

This paper cites Hybridflow: A flexible and efficient rlhf framework.

Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning Hybridflow: A flexible and efficient rlhf framework

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-06-27T13:40:56.637058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-27T13:40:39.598985Z digest=sha256:a9eb1db81fa0d85a1e01519908a532276b65d0665bd95814a4495ebc5fd287c1

Observation 1b5f9014-2376-492d-b3e7-ae5b5c56e784 · outbound

This paper cites CoRR , volume=.

Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning CoRR , volume=

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-27T13:40:39.598985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T13:40:39.598985Z digest=sha256:742ca6838602f7dd41f6cbc55052072fd77cf29552110ef7586b9a8a7eefe181

Observation 6702c7ff-0e30-4f28-8e6a-45d6766bf672 · outbound

This paper cites CoRR , volume=.

Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning CoRR , volume=

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-27T13:40:39.598985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T13:40:39.598985Z digest=sha256:f4a3623b53d474dc68ed116c97b44082bd8d5c31fd45669ab4675ba8f7f29b5a

Observation 6759555a-8bdd-4595-bba1-86fb34d9f78c · outbound

This paper cites Soft Adaptive Policy Optimization.

Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning Soft Adaptive Policy Optimization

Reference 21

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T04:47:37.871052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-27T13:40:39.598985Z digest=sha256:9e171ab45c7754d795cc5a65e976e32304c691fd10b5ccae67610bfdf7f6bb69

Observation 3e36b273-1c89-4306-8a8a-fcabc4df73ce · outbound

This paper cites arXiv preprint arXiv:2601.22718 , year=.

Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning arXiv preprint arXiv:2601.22718 , year=

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-03T04:47:37.873592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-27T13:40:39.598985Z digest=sha256:919fa440658c2545f405bdd8d65a7b29417ac9fcea449a76700909ac6b787391

Observation 22ee3450-61f3-4f6a-b286-a32905d15bfe · outbound

This paper cites CoRR , volume=.

Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning CoRR , volume=

Reference 23

Resolution
unresolved
no resolver link, observed 2026-06-27T13:40:39.598985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T13:40:39.598985Z digest=sha256:4035ae658f3fd8a0baf7a5e5efbc36dff12b320559b909aecc332a41a8b50ca5

Observation 515e78cc-88b2-4214-ad1b-581355e8abc7 · outbound

This paper cites Second Conference on Language Modeling , year=.

Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning Second Conference on Language Modeling , year=

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-27T13:40:39.598985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T13:40:39.598985Z digest=sha256:3e1d87b22fc29468d88d86c0bfefa60c25951833fce2873ce0ea4592f13bb7b1

Observation ad0c724c-5297-4c94-99a9-550e54b433a9 · outbound

This paper cites 2025 , url=.

Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning 2025 , url=

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-27T13:40:39.598985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T13:40:39.598985Z digest=sha256:5228edbb3e7e682dd1fc19c369c401be482944a395f788f96024049409bbf039

Observation fe439ed9-0450-4c35-bcf8-9aaa0cfb9b18 · outbound

This paper cites CoRR , volume=.

Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning CoRR , volume=

Reference 26

Resolution
unresolved
no resolver link, observed 2026-06-27T13:40:39.598985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T13:40:39.598985Z digest=sha256:5a7fbe96717eaa987af77f8d79dd081596ee8cdb613b351c5cf6a851a482374a

Observation 44674da9-7e50-4e80-9487-9e9220f8ee14 · outbound

This paper cites REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization.

Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization

Reference 27

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T04:47:37.870774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-27T13:40:39.598985Z digest=sha256:71259aa49e17d5321c62511afecc1c2237ad7d2c63997b9bacb1132b4a5156b1

Observation d6f86a76-b645-4da9-a1ec-829abc32a93d · outbound

This paper cites CoRR , volume=.

Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning CoRR , volume=

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-27T13:40:39.598985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T13:40:39.598985Z digest=sha256:d25897799930eb01dd2b8f0b921b0381c57e0577974fcfab860dc9ea376aa880

Observation a61ff2c1-7276-4f66-b0d7-96adf902a2fb · outbound

This paper cites Rethinking the Trust Region in LLM Reinforcement Learning.

Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning Rethinking the Trust Region in LLM Reinforcement Learning

Reference 29

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T04:47:37.873043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-27T13:40:39.598985Z digest=sha256:9c44e2be75bcff59aa68b05ec43e175da24a261951f5a15b7e63e4a25890a8bf

Observation 87dc1642-f0b4-4789-a449-6f0e14205f74 · outbound

This paper cites Multi-Step Likelihood-Ratio Correction for Reinforcement Learning with Verifiable Rewards.

Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning Multi-Step Likelihood-Ratio Correction for Reinforcement Learning with Verifiable Rewards

Reference 30

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T04:47:37.866080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-27T13:40:39.598985Z digest=sha256:792c59920857a87aca24816710145cbc01bc8063778cbc644a3fb895ccf279eb

Observation 2cdaeb9c-f790-426c-9086-d08b0766c84d · outbound

This paper cites Fipo: Eliciting deep reasoning with future-kl influenced policy optimization.

Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning Fipo: Eliciting deep reasoning with future-kl influenced policy optimization

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T04:47:37.865305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-27T13:40:39.598985Z digest=sha256:a482014105bf7a3332a084ef200d99da791a259294fb8c6421e7d876a65676d5

Observation 0dbbe045-60e1-40ef-8f5e-d26700439a33 · outbound

This paper cites Trust Region Masking for Long-Horizon LLM Reinforcement Learning.

Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning Trust Region Masking for Long-Horizon LLM Reinforcement Learning

Reference 32

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T04:47:37.863558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-27T13:40:39.598985Z digest=sha256:748955ba45785823400f2c8414cffd6a6ee9bedf2173e6ee32b27b5425a2355e

Observation 2587ab3f-3995-4846-914c-118d62cd7f08 · outbound

This paper cites The Fourteenth International Conference on Learning Representations , year=.

Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning The Fourteenth International Conference on Learning Representations , year=

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-27T13:40:39.598985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T13:40:39.598985Z digest=sha256:438f0c6a4b657414f5b19c5ca1a15412733c6625c23624bd26073d532c98c150

Pith citing papers

Observation 0eb8cef3-6fd1-42ae-b224-2feaf23564c4 · inbound

Predictive Divergence Masks for LLM RL cites this paper.

Predictive Divergence Masks for LLM RL Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-14T08:49:37.615924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T08:49:37.615924Z digest=sha256:d87a35f215031d718e5db8ee7cad83e10f306eab4864581e0009064ce321500b