Pith. sign in

Paper Citation Record · LEDGER

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning

As of 22 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 10 inbound Pith citation observations for arXiv:2502.06533.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.06533 v1

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T15:11:08.217318Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:26:01.278862Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T11:41:29.623621Z

Reference resolution

38 of 38 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved36
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4335a7de-95b4-4289-beb7-31d9e2830a4a · outbound

This paper cites an unresolved cited work.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-08T15:11:08.598573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-08T15:11:08.108360Z digest=sha256:de9c1867358ae240243832918e51adf3dc16c716920c7e4a18df746e2aaf8616

Observation 2be403d8-dfcc-4cbe-acd2-a0877eae64d9 · outbound

This paper cites an unresolved cited work.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-08T15:11:08.590593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-08T15:11:08.111751Z digest=sha256:759aa99d59207d06b8cbbb9970e1905e66fee9b59b8637291e24d528fcf8f976

Observation cb02443b-5e7c-454d-80a1-cdaf2f15c3f6 · outbound

This paper cites an unresolved cited work.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-08T15:11:08.583473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-08T15:11:08.114846Z digest=sha256:107e5d1509819f6d5a2e1bb2980ca60db8ea07c66d0222b6562b8c7cb752fe30

Observation f5e03d1f-0a1a-4099-98da-8e6dfe7a5b76 · outbound

This paper cites an unresolved cited work.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T15:11:08.119207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:11:08.119207Z digest=sha256:41f4f7df55938867da8832968aaeb58d018cd097380b876b032f5d5b7c5097aa

Observation abae6b32-0469-4597-84ce-2779d1d152b1 · outbound

This paper cites an unresolved cited work.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-08T15:11:08.574958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-08T15:11:08.122210Z digest=sha256:870dca939f63aaa836eca35ce803801f2752bb4a3bd7c33294d854436e23ceea

Observation 918f833d-6546-4426-aabd-16754a3d7bdd · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Evaluating Large Language Models Trained on Code

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T15:11:08.125025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:11:08.125025Z digest=sha256:fb883523dbd2013de046a582811af5ff915e4c4590e46d14c105b10c906b4cd0

Observation a65dbc7e-62ff-433b-8364-619b03875d86 · outbound

This paper cites an unresolved cited work.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-08T15:11:08.567131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-08T15:11:08.128531Z digest=sha256:6787a1c9100ad42fd79ff2c688732002ceaa97f634b516f32729fb3d7be04028

Observation 3db7b570-de64-4666-901b-8d936ac93fe0 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Training Verifiers to Solve Math Word Problems

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T15:11:08.131846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:11:08.131846Z digest=sha256:495f02633217a811be1c3c4eadba034ad54322eade97e1db2637f61c29ae568a

Observation c368a59c-9816-4431-9826-845747dd1018 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T15:11:08.134977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:11:08.134977Z digest=sha256:f804dbae2e007eec7b093015921f47415c1a5d73fb0eb8cf0c61c4ea307d3c0b

Observation 0d0b062e-1036-4c1f-b7fe-403edeb2d6d2 · outbound

This paper cites Teaching Large Language Models to Reason with Reinforcement Learning.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Teaching Large Language Models to Reason with Reinforcement Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T15:11:08.137792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:11:08.137792Z digest=sha256:5c744574d6cceb473139aa4e2b289689f850d5938ebe1ffd0eede3e23bef4f66

Observation ab25b46b-525a-4587-ac65-4057203baa0b · outbound

This paper cites an unresolved cited work.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T15:11:08.140623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:11:08.140623Z digest=sha256:eb1fb7bced96e2901ca04607acda211a5a52d3876ec83a76aeecf99db1539b84

Observation fabb473f-eb47-46c3-a79e-e01d1d4da8f1 · outbound

This paper cites an unresolved cited work.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-08T15:11:08.555008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-08T15:11:08.143465Z digest=sha256:a86d33a482ba5a5422991c0ec646c6e132ed8d3219bb3e3e36cc7386613c91de

Observation ef6d4f98-1704-467a-a0a5-21484918da71 · outbound

This paper cites an unresolved cited work.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-08T15:11:08.547831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-08T15:11:08.145539Z digest=sha256:379b5eaf8dbcbf4aa662d9fef8e215b0b24b3cd505d1de88831749ccf44072e9

Observation f1b2ed8f-87c8-44a9-ade0-4021d5e97acd · outbound

This paper cites an unresolved cited work.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T15:11:08.147904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:11:08.147904Z digest=sha256:3bf276f26ae044e9c19ddaa4433c4a61fd80c5e88e105cdd242cdc06cfd3005c

Observation 6086b97c-1a84-413e-a8a6-fa899a26aca7 · outbound

This paper cites Lee, Kangwook Lee, and Dimitris Papailiopoulos.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Lee, Kangwook Lee, and Dimitris Papailiopoulos

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:11:08.541335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-08T15:11:08.150117Z digest=sha256:50dbac350476b735749fedfeb0c9196450aca29d3ae07f902b0356554b4a9c53

Observation 028bfb75-43f9-46a5-9400-58b7aba0a563 · outbound

This paper cites an unresolved cited work.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T15:11:08.152169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:11:08.152169Z digest=sha256:c529a7d770936b3fe90e2a90ef4f0af67c605d064fc09c3d0f462c4f00c2509f

Observation 4389c21f-a67e-4a21-8939-300e35901201 · outbound

This paper cites Goat: Fine-tuned LLaMA Outperforms GPT-4 on Arithmetic Tasks.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Goat: Fine-tuned LLaMA Outperforms GPT-4 on Arithmetic Tasks

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T15:11:08.154302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:11:08.154302Z digest=sha256:84e9527bd75bb934eca6e4a84108d06b7cf85ed1e66032d3b10ca5786fad2b7f

Observation 92dfa1ca-53fb-45a5-93cf-af728f0d6dee · outbound

This paper cites Bartoldson, Bhavya Kailkhura, Abhinav Bhatele, Jonas Geiping, Avi Schwarzschild, and Tom Goldstein.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Bartoldson, Bhavya Kailkhura, Abhinav Bhatele, Jonas Geiping, Avi Schwarzschild, and Tom Goldstein

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:11:08.534555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-08T15:11:08.157948Z digest=sha256:3092c3f7dc4e6b8d5f2212e9e99658477c4683849677c3138313c7d1efba8051

Observation a0a11701-511c-4602-a25e-807baca51cc0 · outbound

This paper cites an unresolved cited work.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-08T15:11:08.526289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-08T15:11:08.160447Z digest=sha256:af4e3326b54e5129e3942ab1d3816abf0a1fdcf5e9377bbf6ef8035dceef9956

Observation 020d0dcb-a310-4071-aa29-30b33186f5fa · outbound

This paper cites an unresolved cited work.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-08T15:11:08.518595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-08T15:11:08.163134Z digest=sha256:d5637a303b012f26e7dc500bbd119f5852ccc4ea973dd94de6c89ca9557f54dd

Observation 24008427-3b92-4768-b228-d42bd07c16e4 · outbound

This paper cites an unresolved cited work.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T15:11:08.165824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:11:08.165824Z digest=sha256:d679547a0aec394de2c241f4d7ff3947e89baa77930b99906950f59ca1fc63f1

Observation 757c0598-cfca-444c-9323-03d517276ec4 · outbound

This paper cites an unresolved cited work.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-08T15:11:08.506127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-08T15:11:08.168861Z digest=sha256:163d11b0fa37cff8132392f47c4d9e35f0859780cc85e417aec9a650b516e6b0

Observation c9a8232f-e00c-44de-8223-7c30eedb9ef1 · outbound

This paper cites Positional Description Matters for Transformers Arithmetic.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Positional Description Matters for Transformers Arithmetic

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T15:11:08.172446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:11:08.172446Z digest=sha256:e8b64a70bfd674eacf58e7d66395f68378f3584da532a61efed04f5bb6222d9a

Observation cac9a057-a2aa-416f-9395-58ec1ee3f4d2 · outbound

This paper cites an unresolved cited work.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-08T15:11:08.498250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-08T15:11:08.176269Z digest=sha256:9675c541c121447f7dc937b8a7e8b5854868dc4c80122d3c3c582b1cde8a6018

Observation 55a321ac-facf-48ee-b2ac-08c9b019dc82 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T15:11:08.179605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:11:08.179605Z digest=sha256:1f56768895355f8b9ea2ad900afb29462c1926bbe249411066bf84c0f43d53a3

Observation 24c77f65-7e3f-4ac4-acfc-a7da8bfd6b90 · outbound

This paper cites an unresolved cited work.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T15:11:08.182556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:11:08.182556Z digest=sha256:f5bc207f50dd68a59a54248e76dda83b22bdb57a7dd0aacdaac9a39706fe4657

Observation 636c5415-2dd8-4bb2-a643-7afc60724c4b · outbound

This paper cites Chi, Quoc V.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Chi, Quoc V

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T15:11:08.185354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:11:08.185354Z digest=sha256:c88bad6dcd02de3dedcd51d02d4f60a16f412d64bfd24da42c8ea7dd9b52b0f0

Observation 665bf2e8-fd54-4a58-9104-00a2bd0f3aaa · outbound

This paper cites an unresolved cited work.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-08T15:11:08.486457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-08T15:11:08.188614Z digest=sha256:3fa53f949ddc090b45ed8165ececa25083d0adf31e894fc255abc013f1f06e42

Observation 0515158b-359e-4933-9066-0b29407c7474 · outbound

This paper cites an unresolved cited work.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T15:11:08.191550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:11:08.191550Z digest=sha256:02c299b6105ba5ceb5652136f55cc9b70fe78fb5dbbb59372a7d0456aaa80627

Observation cde243c6-e19a-47f4-9383-71f9af953329 · outbound

This paper cites Conditions for Length Generalization in Learning Reasoning Skills.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Conditions for Length Generalization in Learning Reasoning Skills

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T15:11:08.194619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:11:08.194619Z digest=sha256:ed85b4e3b812a8fd9ecc86fae4a8868d18228d1406b373ab939d5dda73396788

Observation d5e1c47b-f1a6-4b0f-9def-eda2b3f42635 · outbound

This paper cites Hausknecht, and Karthik Narasimhan.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Hausknecht, and Karthik Narasimhan

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T15:11:08.197539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:11:08.197539Z digest=sha256:cc9c5d28555678b4f3424c1d9131f613c997e1f4c564f347f88538c18a4d6d18

Observation cdf1b3b7-0ef0-4a28-98f7-d25e3ee8bb00 · outbound

This paper cites How well do Large Language Models perform in Arithmetic tasks?.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning How well do Large Language Models perform in Arithmetic tasks?

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T15:11:08.200325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:11:08.200325Z digest=sha256:a32c2561862bcfbf7e7d251755db6feb98900daf4f73ff7cf01299e11eea354f

Observation 6c344090-0431-4f4e-add7-abf0ab2e5722 · outbound

This paper cites an unresolved cited work.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-08T15:11:08.474833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-08T15:11:08.203740Z digest=sha256:ba83803cb17578c410e6c2e7e6f4427ed1aabafe0ba8cd7e4170a46f2c1dd56f

Observation 18beb684-bb45-4677-b7c7-b82c595574a7 · outbound

This paper cites Chain-of-Thought Reasoning is a Policy Improvement Operator.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Chain-of-Thought Reasoning is a Policy Improvement Operator

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T15:11:08.206499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:11:08.206499Z digest=sha256:dd347aaf52f4c0f6ae84dd344bb2f59dc15aa6da554845225b05ca862103b89a

Observation c16742db-1a7f-40b6-9c3a-95ea4da776e1 · outbound

This paper cites an unresolved cited work.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-08T15:11:08.466799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-08T15:11:08.209904Z digest=sha256:7be1d0655055f6e6999de7ee22d150e4df5c42b1315cbebc223d3106eeb4b893

Observation 9c541d36-6aa4-4a9a-a160-5566293ec96d · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Fine-Tuning Language Models from Human Preferences

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-08T15:11:08.212092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:11:08.212092Z digest=sha256:2f0f5d6183d3f07d29325215b4339eef71dab80a30c598cd040772b279f046a8

Observation dfd56579-ae4d-4fdf-b3ae-cb16ce856c98 · outbound

This paper cites online" 'onlinestring :=.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning online" 'onlinestring :=

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-08T15:11:08.214627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:11:08.214627Z digest=sha256:a1b19b2a26a238feb2be0c402f3b983f4b735feb8adf91ab57ce5d776584c04e

Observation 453d04b3-385b-43cc-b44c-fb4d7f650d30 · outbound

This paper cites write newline.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning write newline

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T15:11:08.217318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:11:08.217318Z digest=sha256:2a949242e28bc33a9a2a28ed124afd9b2c3404e61531a2357cd31c0ce365737f

Pith citing papers

Observation 7674cb9a-22f7-4bfd-b688-55d4dd1f049a · inbound

Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning cites this paper.

Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-12T12:12:08.960644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-12T12:12:08.724844Z digest=sha256:0acca326751651a5cf1e587320aa5105bbec80d8ebc3994f2b71f9440ae63ab6

Observation 44b0d37e-a9a3-4184-9cb0-6a9915ea55e8 · inbound

Discovering Algorithms with Computational Language Processing cites this paper.

Discovering Algorithms with Computational Language Processing Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T20:26:01.278862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:26:01.278862Z digest=sha256:64214b05f767bde276d3dbfb9eecbdcc2f627f66feb0d2b3295bc1ae27628c85

Observation 6a1c392c-5e2c-4a08-a44c-b3dcb06b6e52 · inbound

CogniSQL-R1-Zero: Lightweight Reinforced Reasoning for Efficient SQL Generation cites this paper.

CogniSQL-R1-Zero: Lightweight Reinforced Reasoning for Efficient SQL Generation Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:38.761833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:18:38.761833Z digest=sha256:2a05522e6272488f8b20934dd3f2ce3c6642dd68ffa1be13bbfbb52884220e94

Observation 1eab7712-b1b2-44ff-98fa-4910f7e8c24c · inbound

Stabilizing Knowledge, Promoting Reasoning: Dual-Token Constraints for RLVR cites this paper.

Stabilizing Knowledge, Promoting Reasoning: Dual-Token Constraints for RLVR Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T23:24:26.241888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-21T23:20:45.685446Z digest=sha256:f7584698cb6270c0693caea7a0d22e74bd4811e7d6a56f61a5b3aee301cc0763

Observation 376af9d9-ebde-4618-9197-5700c26f2387 · inbound

Self-Reflective Generation at Test Time cites this paper.

Self-Reflective Generation at Test Time Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T12:41:43.057152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:41:43.057152Z digest=sha256:cb5c5542b163cf8d65c882dfd633dece7ccdfb733cda72abea52d5808937d96b

Observation 0b57fbfe-e87d-4eb8-bf7b-53889a8b50a3 · inbound

Attention Illuminates LLM Reasoning: The Preplan-and-Anchor Rhythm Enables Fine-Grained Policy Optimization cites this paper.

Attention Illuminates LLM Reasoning: The Preplan-and-Anchor Rhythm Enables Fine-Grained Policy Optimization Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-04T09:48:09.263325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T09:48:09.263325Z digest=sha256:0d4a63740d93ae7029577a2cbeec9a41751f33cf508287c90169922aeaa5ef04

Observation 68e48558-ca96-4c19-b535-78a32c01ac09 · inbound

Training-Trajectory-Aware Token Selection cites this paper.

Training-Trajectory-Aware Token Selection Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-22T11:41:29.626509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-22T11:41:21.275802Z digest=sha256:d131ae966935db398810f4af30ac9bbf2f541e3b59b07f2bd069e63dc9bb8324

Observation 713e4f03-97af-4015-a2a7-e47620c85e4b · inbound

Embarrassingly Simple Self-Distillation Improves Code Generation cites this paper.

Embarrassingly Simple Self-Distillation Improves Code Generation Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-13T14:33:35.834383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T14:33:35.834383Z digest=sha256:506c02b053efc3a550e7f19f878c37d460e9ffc1bdf4dd78266af0d68d6b8da1

Observation 9ccb437f-a756-4f65-9150-38507645632c · inbound

AtManRL: Towards Faithful Reasoning via Differentiable Attention Saliency cites this paper.

AtManRL: Towards Faithful Reasoning via Differentiable Attention Saliency Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T08:22:37.709068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T08:08:50.330857Z digest=sha256:1f0dc3f95fcd45b2c938eee8c934e9b044adbb1f62d70ec905bcbc566c2bde11

Observation 3164a8ee-072f-4e86-80d7-3c6c4139884a · inbound

HTPO: Towards Exploration-Exploitation Balanced Policy Optimization via Hierarchical Token-level Objective Control cites this paper.

HTPO: Towards Exploration-Exploitation Balanced Policy Optimization via Hierarchical Token-level Objective Control Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T00:51:14.645501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-12T00:50:42.836549Z digest=sha256:7f8e105bf08026d51df941c8c136827ec32558963d4aaea8d35584c19665a549