Pith. sign in

Paper Citation Record · LEDGER

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2502.06533.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.06533 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:26:01.278862Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T11:41:29.623621Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 7674cb9a-22f7-4bfd-b688-55d4dd1f049a · inbound

Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning cites this paper.

Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-12T12:12:08.960644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T12:12:08.724844Z digest=sha256:522690db162af0c1526e7a2a4698950b84c43588460d389fd28140d739159fd4

Observation 44b0d37e-a9a3-4184-9cb0-6a9915ea55e8 · inbound

Discovering Algorithms with Computational Language Processing cites this paper.

Discovering Algorithms with Computational Language Processing Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T20:26:01.278862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:26:01.278862Z digest=sha256:4b2d090f7ebc6aa7f367c1553a334f82bb1874b3dcadb55b0f37df88a99c7f7c

Observation 6a1c392c-5e2c-4a08-a44c-b3dcb06b6e52 · inbound

CogniSQL-R1-Zero: Lightweight Reinforced Reasoning for Efficient SQL Generation cites this paper.

CogniSQL-R1-Zero: Lightweight Reinforced Reasoning for Efficient SQL Generation Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:38.761833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:18:38.761833Z digest=sha256:6f02f28b098e743c7b85c9e43c00706c44b712daf2b4ef17a35f24d763012618

Observation 1eab7712-b1b2-44ff-98fa-4910f7e8c24c · inbound

Stabilizing Knowledge, Promoting Reasoning: Dual-Token Constraints for RLVR cites this paper.

Stabilizing Knowledge, Promoting Reasoning: Dual-Token Constraints for RLVR Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T23:24:26.241888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-21T23:20:45.685446Z digest=sha256:4baf17ea2a7c455a2d9e47c351fe507c103b8ba426a3c3bff6ecb07ea47e1f56

Observation 376af9d9-ebde-4618-9197-5700c26f2387 · inbound

Self-Reflective Generation at Test Time cites this paper.

Self-Reflective Generation at Test Time Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T12:41:43.057152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:41:43.057152Z digest=sha256:92c0e60faa409bd31e42f0b3b785aedbb7130f7fb696565324172346e5e12d1b

Observation 0b57fbfe-e87d-4eb8-bf7b-53889a8b50a3 · inbound

Attention Illuminates LLM Reasoning: The Preplan-and-Anchor Rhythm Enables Fine-Grained Policy Optimization cites this paper.

Attention Illuminates LLM Reasoning: The Preplan-and-Anchor Rhythm Enables Fine-Grained Policy Optimization Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-04T09:48:09.263325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T09:48:09.263325Z digest=sha256:a7cdbac4543d523bbdd15c57ae69a27229b63c0a56a738e993ad4194c7b5e610

Observation 68e48558-ca96-4c19-b535-78a32c01ac09 · inbound

Training-Trajectory-Aware Token Selection cites this paper.

Training-Trajectory-Aware Token Selection Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-22T11:41:29.626509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T11:41:21.275802Z digest=sha256:5fd878f6aec8368c64b1f7d41f9c853c450a8d8ae8789a71089e9e311cfa1a43

Observation 713e4f03-97af-4015-a2a7-e47620c85e4b · inbound

Embarrassingly Simple Self-Distillation Improves Code Generation cites this paper.

Embarrassingly Simple Self-Distillation Improves Code Generation Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-13T14:33:35.834383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T14:33:35.834383Z digest=sha256:a4666bf481bdb2f5ad7cae234629fe9e4f77f011e8adbbe960f6c444d93c7472

Observation 9ccb437f-a756-4f65-9150-38507645632c · inbound

AtManRL: Towards Faithful Reasoning via Differentiable Attention Saliency cites this paper.

AtManRL: Towards Faithful Reasoning via Differentiable Attention Saliency Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T08:22:37.709068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T08:08:50.330857Z digest=sha256:bfc5f24fb2c48c2825cc5994d9bde36d23728439612cf8d4bbdc7620ffb88191

Observation 3164a8ee-072f-4e86-80d7-3c6c4139884a · inbound

HTPO: Towards Exploration-Exploitation Balanced Policy Optimization via Hierarchical Token-level Objective Control cites this paper.

HTPO: Towards Exploration-Exploitation Balanced Policy Optimization via Hierarchical Token-level Objective Control Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T00:51:14.645501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T00:50:42.836549Z digest=sha256:9cd5f840540686569ebfd1a713db7fb389d000668d87f389b34b2c1682ee4933