Pith. sign in

Paper Citation Record · LEDGER

Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs

As of 20 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 15 inbound Pith citation observations for arXiv:2505.12929.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.12929 v1

Coverage vector

measured 53 of 53 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:29:01.735774Z

measured 68 of 68 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T22:52:02.124869Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T14:27:07.373132Z

Reference resolution

53 of 53 outbound references displayed

  • verified exact0
  • verified fuzzy24
  • unresolved28
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5526fc05-a027-4cc4-87c6-ddaddea30bdd · outbound

This paper cites OpenAI o1 System Card.

Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs OpenAI o1 System Card

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T20:29:01.471084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:29:01.471084Z digest=sha256:08bc9c05d01286c5444b4d7561ef61500783d5224ef5e3bed04c7abc5c105195

Observation 5406fdbb-4131-4219-b078-e436f69483f6 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T20:29:01.477574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:29:01.477574Z digest=sha256:c32d58c5fde3967c5a4551c649beec1eb977792415a571db744d244c88b74597

Observation 6a56aad7-552c-4c90-a4cc-73f0a93fef00 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T20:29:01.483089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:29:01.483089Z digest=sha256:17536c60e040887212cac76a905355a455df7757c443f176748788146ed3f3c0

Observation 849d660e-a703-481c-a394-af141c2fa03e · outbound

This paper cites Monte carlo tree search boosts reasoning via iterative preference learning.

Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs Monte carlo tree search boosts reasoning via iterative preference learning

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:29:02.636536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:29:01.489299Z digest=sha256:93d5edfc13fe2d403233dcc912745ecd77eb795c82af153be6edd40cf5c8ba5f

Observation 8918163f-31d7-4a78-a55c-f9f82be8e9ec · outbound

This paper cites Alphamath almost zero: Process su- pervision without process.

Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs Alphamath almost zero: Process su- pervision without process

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:29:02.620104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:29:01.494519Z digest=sha256:1450089467615e56b13a21caef34f9ae18d1e2b4cbc10a23246367f1dc4c2cfe

Observation 9da1503d-f1d6-4bd1-982f-54a9245f8d6a · outbound

This paper cites Let’s verify step by step.

Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs Let’s verify step by step

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T20:29:01.500140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:29:01.500140Z digest=sha256:89d9b7a2e1212145de9e9dc0d5bed08dc1bc5c28233f764c51aaac97c4de4d33

Observation c1598d75-938b-4f66-afa4-48079d76a431 · outbound

This paper cites Math-shepherd: Verify and reinforce llms step-by-step without human annotations.

Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs Math-shepherd: Verify and reinforce llms step-by-step without human annotations

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:29:02.593102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:29:01.505500Z digest=sha256:1b9515594476e699c32e39af0dd7e99d7ae1c9605a3de1740acb25aeabd0d469

Observation 6bca8f1d-3e62-4948-9d66-ab6019275a4f · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T20:29:01.510372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:29:01.510372Z digest=sha256:57447dfd4b10f6a58f2e4a47691e360aff86681e6a3d2e2c88ba12171d3ca2e4

Observation 05325b6c-6c2f-4ed9-b818-fd78ae52ee50 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T20:29:01.515576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:29:01.515576Z digest=sha256:7b021222c1223ce495c4eafcddc337fa7ac0ac159a6475e054d236ca22b0e1e6

Observation 65b6a274-cf18-4d2e-a71f-f91342c50a10 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs Understanding R1-Zero-Like Training: A Critical Perspective

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T20:29:01.521431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:29:01.521431Z digest=sha256:82860cc48bb1e5531164f7981cde8c9305987b612642dcf38be80d8c4a763acb

Observation ead49eaa-8eb3-46c2-b3bd-ebb996ea348c · outbound

This paper cites A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce.

Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T20:29:01.530228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:29:01.530228Z digest=sha256:7a4e778a52ceb877c83e8986ebb490b0a3d4e42f03e2d6784beb0c572e32e986

Observation 86589244-9023-4876-b9c9-048fc525436b · outbound

This paper cites Deep reinforcement learning from human preferences.

Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs Deep reinforcement learning from human preferences

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:29:02.578463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:29:01.535467Z digest=sha256:ebbad3e1010fb17f1cfb99ea332bc8b58c5be726cec20a28e1ba2b53905ca113

Observation eb4ed842-0e90-48ff-81a3-24067ce60bff · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs Fine-Tuning Language Models from Human Preferences

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T20:29:01.540427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:29:01.540427Z digest=sha256:b84b967a1dc5e22292137cbdef5818f27d962420bbb54468b072e648a247c5ac

Observation fff14750-b579-4fef-8a0f-808dd0f1ef5c · outbound

This paper cites Learning to summarize with human feedback.

Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs Learning to summarize with human feedback

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:29:02.563031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:29:01.545522Z digest=sha256:cc093766fb328dcf1ba3ce5b0660e297a251ba218ddc9270a3ea1539c63b9fef

Observation 19befe5b-abac-4da2-b8f0-3019168b228a · outbound

This paper cites Training language models to follow instructions with human feedback.

Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs Training language models to follow instructions with human feedback

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:29:02.548096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:29:01.550214Z digest=sha256:3254843f02b13dba58145df834e13e569d598190ec5dd5e503046d07875dd463

Observation 6f80dba7-3e47-4faf-9b55-d81fbd7307fb · outbound

This paper cites Proximal Policy Optimization Algorithms.

Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs Proximal Policy Optimization Algorithms

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T20:29:01.555019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:29:01.555019Z digest=sha256:d609b8d78273a75c63659f0874226fdc5effb07459a7d98d57beda20c73b98e7

Observation c0eb6887-8369-4124-bfd6-94ab3d1cf39f · outbound

This paper cites Language models are few-shot learners.

Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs Language models are few-shot learners

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:29:02.532709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:29:01.560022Z digest=sha256:39420f02a5754efe91082046a1b3f8383e03101f0b4cfc127df102f4de8084c8

Observation d9dff6a3-44c1-4cc3-99c9-622bed19f749 · outbound

This paper cites GPT-4 Technical Report.

Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs GPT-4 Technical Report

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T20:29:01.565135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:29:01.565135Z digest=sha256:306dcdb52174409f8544a7615ea08e34722f7b33ac325cc8774386639eea8315

Observation 8721c4e6-4de0-400d-88e9-f525a519c028 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs LLaMA: Open and Efficient Foundation Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T20:29:01.570064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:29:01.570064Z digest=sha256:a84706239879aa2aedb1358830fc087d109fbb570046eefeff2b142e8c24219c

Observation c5494e24-675c-4573-9b1e-238c749465f7 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T20:29:01.574926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:29:01.574926Z digest=sha256:555e75bf056f558fb817c6427a725bb5f53983840981bd24b75783cabd88823d

Observation 3a37f115-ddb2-40c6-8dbd-f4c0800be69e · outbound

This paper cites The Llama 3 Herd of Models.

Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs The Llama 3 Herd of Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T20:29:01.579804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:29:01.579804Z digest=sha256:a5bfa727fa2450c60ddf08372dcf9f50bc3358fa0ca981fca75cffcd76df2afe

Observation 1a5ae833-4b33-4c01-a70a-33b0528dfe75 · outbound

This paper cites Qwen Technical Report.

Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs Qwen Technical Report

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T20:29:01.584511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:29:01.584511Z digest=sha256:c38ee6b4b61eee79a762a874507847a8f47a9a00cc5ff529d65bbac783fddf0b

Observation 592b6e06-1238-4e85-a6d8-5d22647b2ba2 · outbound

This paper cites Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models.

Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T20:29:01.589285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:29:01.589285Z digest=sha256:9e9f8cfd1c41d253f56b63ab35f47a09f822651a3b85d5093ebbdeb1bb377ac5

Observation eb40335f-8ec3-4346-905a-1d47c932203b · outbound

This paper cites Qwen2 Technical Report.

Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs Qwen2 Technical Report

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T20:29:01.594462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:29:01.594462Z digest=sha256:ea803fe740494993365ff35106578984fa733ed6457cd159c394ae13cbffbcef

Observation c936a1ed-f186-4dc2-beef-6ff0e6f89a4d · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs Gemini: A Family of Highly Capable Multimodal Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T20:29:01.599413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:29:01.599413Z digest=sha256:a183c4e2bdacbe924e108ddcd2c5d91296ecc7a3ff214c2b02d91a97f17212e7

Observation 4b64ba0f-edad-4a25-93c6-609c5fca5bf0 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T20:29:01.604503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:29:01.604503Z digest=sha256:e46557fb79b081beb9266d938ef69792ec91aa314ef37684a9f2afccc02ce2f7

Observation e10b5731-abd3-4506-93b6-4454f4999754 · outbound

This paper cites The Claude 3 model family: Opus, sonnet, haiku.

Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs The Claude 3 model family: Opus, sonnet, haiku

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:29:02.516490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:29:01.609712Z digest=sha256:180e7e42df88d4df2bfdc63315e2b0a3ff42108bda428be7b70e9cc296b0301f

Observation 70e37f1d-b0ac-4b72-90ed-508ec8d1f208 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs Direct preference optimization: Your language model is secretly a reward model

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:29:02.500182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:29:01.614552Z digest=sha256:e5b2429e5a7b36d3c275be26c02926ca9241ed6459489d611e8282afb5301a37

Observation f45a84af-8051-43c6-9fe8-9a24b9753eb3 · outbound

This paper cites From r to Q*: Your language model is secretly a Q-function.

Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs From r to Q*: Your language model is secretly a Q-function

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:29:02.484484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:29:01.619470Z digest=sha256:e161f7847658e4c236ceb4e8a605260fe57adf39fe5394b593d889a68f0997b2

Observation 886ec526-8660-4ab8-8745-dd8ab8c30c4d · outbound

This paper cites Is DPO superior to PPO for LLM alignment? a comprehensive study.

Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs Is DPO superior to PPO for LLM alignment? a comprehensive study

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:29:02.469600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:29:01.624457Z digest=sha256:54ed23e10d5af0ed3b9206e7c7f7090b8a959ffa8c5fe36460cf02f2557b45bf

Observation af302d94-5822-44ae-9f6e-1e4bc9e9000c · outbound

This paper cites DPO meets PPO: Reinforced token optimization for RLHF.

Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs DPO meets PPO: Reinforced token optimization for RLHF

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:29:02.453265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:29:01.629387Z digest=sha256:45fe44bafbcbc11f9e4fac34544f8d030dd8b9ff318a0106510d672a9c75dcf4

Observation ca71bb59-21f8-4a8c-a71f-2b463adc45f9 · outbound

This paper cites Sutherland.

Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs Sutherland

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:29:02.436839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:29:01.634236Z digest=sha256:72b417fb9c4bc41bfa8b098733c1afe82bd5046589c470c3c8ba3c54ff0e6380

Observation 358cdf93-5346-44f3-85dc-ec67d831372a · outbound

This paper cites ORPO: Monolithic preference optimization without reference model.

Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs ORPO: Monolithic preference optimization without reference model

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:29:02.418963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:29:01.639159Z digest=sha256:75fd32509616485d8707bec2bf8fffa8a566834cd35ab2626380201c654c6c3f

Observation b03a277f-c07d-4d39-9843-a34a48e52781 · outbound

This paper cites Contrastive preference optimization: Pushing the boundaries of llm performance in machine translation.

Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs Contrastive preference optimization: Pushing the boundaries of llm performance in machine translation

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:29:02.402551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:29:01.644026Z digest=sha256:c9681b4148429876491ecbc70461851ed3925ff49b37fd66d554ed3bf773cd6b

Observation 2598b004-24fb-4140-9093-fafc0e0d9354 · outbound

This paper cites SimPO: Simple preference optimization with a reference-free reward.

Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs SimPO: Simple preference optimization with a reference-free reward

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:29:02.386629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:29:01.648645Z digest=sha256:b70fe7324697ce70f34540ed247f304e08575810b0c7fbc846007bd3f66d177e

Observation c609134d-aa82-41d7-b185-fd9961827394 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs Chain-of-thought prompting elicits reasoning in large language models

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:29:02.370794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:29:01.653279Z digest=sha256:1c5aa849e258c62d236e8b451b8646b78f5e54e137a43255c4d967cc85175a23

Observation 382cad17-3053-43f5-b5a6-843a7bd4733f · outbound

This paper cites Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica.

Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:29:02.354290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:29:01.657932Z digest=sha256:c5a2e6a596fbde936f468726740c69b07c88d2fb7200d1d801e8b986d5893018

Observation c454f95f-1eff-4265-8ec9-d772c025551e · outbound

This paper cites Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning.

Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T20:29:01.662434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:29:01.662434Z digest=sha256:35c48b14ea4d22b83bb45fec6903c8c5177c7764aaab678e773988b117eb6f75

Observation f56f0902-ac21-4a42-b8e3-70542325af25 · outbound

This paper cites Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model.

Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T20:29:01.668030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:29:01.668030Z digest=sha256:b3dc18a391b3897d9f0692948ee335d3701f732ea300d32e3c11ed54678c4a87

Observation be9b3875-e80d-41f1-99ae-44df00a5b9fa · outbound

This paper cites SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild.

Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T20:29:01.673051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:29:01.673051Z digest=sha256:a8b988c3f3186f8426276d7e38e37d10207c6c9b7ef01631f687314b05c4c240

Observation b3a2ab91-fd47-4fbf-be6b-6453c3378dc6 · outbound

This paper cites VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks.

Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T20:29:01.677994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:29:01.677994Z digest=sha256:70c2632565dc356754f18b6b89d5df31e9adc5ffcac13aa7aa6621cd324f996a

Observation 678ca9be-48c7-41f6-9b1e-2c2002903985 · outbound

This paper cites Efficient Reinforcement Finetuning via Adaptive Curriculum Learning.

Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs Efficient Reinforcement Finetuning via Adaptive Curriculum Learning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T20:29:01.683537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:29:01.683537Z digest=sha256:40307a29b6d99baf06a8770e062e00bf834963aa38684e2639b294891b67a9c8

Observation d493bbd5-ac1e-424f-ad86-3f04f77d726a · outbound

This paper cites Attention is all you need.

Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs Attention is all you need

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T20:29:01.688332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:29:01.688332Z digest=sha256:e13f8addb98bab48b2ce90bcd668363affa6516047e53d2d95556bb9eeec2446

Observation d7c8b96d-491a-4a48-9c27-be32b1f6ca01 · outbound

This paper cites Hybridflow: A flexible and efficient RLHF framework.

Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs Hybridflow: A flexible and efficient RLHF framework

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:29:02.327951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:29:01.692734Z digest=sha256:172f61a4d48e7d52fdfe19a243044c08854722939715f71440f98a359116220a

Observation 778f7272-2792-4308-8ac7-820d26be5555 · outbound

This paper cites On Memorization of Large Language Models in Logical Reasoning.

Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs On Memorization of Large Language Models in Logical Reasoning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T20:29:01.697399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:29:01.697399Z digest=sha256:da4d734447567af01f7ed8ca6e73e69763d7c64922475c5c902be46cccec972e

Observation fd308106-049f-4688-bc4d-c98e4417cb56 · outbound

This paper cites REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization.

Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T20:29:01.702314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:29:01.702314Z digest=sha256:56b8869627cf6bd313ed1420ee1cd30fedcf5cc600493c28dba7048398c8104e

Observation dbf63c99-bd86-4c73-b8d2-e1b74458d0fb · outbound

This paper cites What is the name of this book? Touchstone Books Guildford, UK, 1986.

Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs What is the name of this book? Touchstone Books Guildford, UK, 1986

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:29:02.311449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:29:01.707311Z digest=sha256:55f66aa546768a36787d7c61052564cf38552c7ab8906339a7526a59927aeaa2

Observation 166b8592-9003-4ddd-b4ad-2f844a56be24 · outbound

This paper cites Meta-logical problems: Knights, knaves, and rips.

Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs Meta-logical problems: Knights, knaves, and rips

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:29:02.293570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:29:01.712056Z digest=sha256:6ae0cbe8afdd9b633c74657c76a29f0e44d3be6a052e28e173933b159771d09f

Observation dc81e272-2755-4b6f-8cba-3e83b7eeea38 · outbound

This paper cites Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems.

Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:29:02.277686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:29:01.716603Z digest=sha256:f843587600eabce139b24835568e5d53ad4011046a65f9bc45261d4c53fa59fa

Observation 1f3f9c9a-d5b5-41ca-b292-f76012178bd4 · outbound

This paper cites Solving quantitative reasoning problems with language models.

Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs Solving quantitative reasoning problems with language models

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:29:02.260729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:29:01.721440Z digest=sha256:24c748397000592139dced4f0d9164b6f5a05495352efee54ca5246ed0b80dc1

Observation 2c0a9a57-5372-4268-9459-0029ae0f16da · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs Measuring Mathematical Problem Solving With the MATH Dataset

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T20:29:01.726086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:29:01.726086Z digest=sha256:759dc2404bf269ff426f61dca77e132dd8ae452ab1df123aa275f188cd41707b

Observation 1bcb8349-f2d1-4a3a-a853-68f3fa746092 · outbound

This paper cites Approximating KL divergence.

Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs Approximating KL divergence

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:29:02.244387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:29:01.731218Z digest=sha256:56256f8c27078e4031b8bfa1f0aa4f66a4a57e854100bc1812beed6894432e87

Observation 92e56b54-cf11-41e8-9be7-e9594d427974 · outbound

This paper cites Lily is a knave or Lily is a knight.

Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs Lily is a knave or Lily is a knight

Reference 53

Resolution
malformed identifier
raw_fallback, observed 2026-08-15T20:29:02.226983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:29:01.735774Z digest=sha256:e5dd7fa1db9a09ac75d084eeffbfdbe5d78647d10094d6f099d71a939999591b

Pith citing papers

Observation e4be3531-b861-4d71-b485-7ec46ae5d5d7 · inbound

Stabilizing Knowledge, Promoting Reasoning: Dual-Token Constraints for RLVR cites this paper.

Stabilizing Knowledge, Promoting Reasoning: Dual-Token Constraints for RLVR Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T23:24:26.303073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-21T23:20:45.685446Z digest=sha256:5a2212306ad983cfa2112c6aeae7336d6e4a58ad3ceb5ff4ab59909582d9a040

Observation c743253d-9e0d-4fdb-aa86-b63a49ff7c4e · inbound

Sharpness-Guided Group Relative Policy Optimization via Probability Shaping cites this paper.

Sharpness-Guided Group Relative Policy Optimization via Probability Shaping Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-18T03:12:22.271239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T03:10:57.839146Z digest=sha256:03732f5fd8a6d017b22070defd3a15f024b5cacd6fbb99bbeb5cf885ef63fa7a

Observation 02ac3733-2f36-4d5f-8fbe-047402cdfd01 · inbound

STAPO: Stabilizing Reinforcement Learning for LLMs by Silencing Rare Spurious Tokens cites this paper.

STAPO: Stabilizing Reinforcement Learning for LLMs by Silencing Rare Spurious Tokens Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-15T21:46:43.296750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T21:41:48.690125Z digest=sha256:3d590b47b303da6ae6be1261b2a6f0ce5f282c7e895823ba27aa4e132db52301

Observation cad0b166-4da6-4db6-a43a-72de0c98a855 · inbound

STAPO: Stabilizing Reinforcement Learning for LLMs by Silencing Rare Spurious Tokens cites this paper.

STAPO: Stabilizing Reinforcement Learning for LLMs by Silencing Rare Spurious Tokens Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T22:52:02.124869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:52:02.124869Z digest=sha256:4220999be63a4e2d0a3888c9de7d15ec1d06d7cdbafae554b7c2b51cb07f95f0

Observation 648c269a-14bd-455c-a677-c86f9c05996b · inbound

GRPO-VPS: Enhancing Group Relative Policy Optimization with Verifiable Process Supervision for Effective Reasoning cites this paper.

GRPO-VPS: Enhancing Group Relative Policy Optimization with Verifiable Process Supervision for Effective Reasoning Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:14:46.307201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T00:14:29.531510Z digest=sha256:e386863ad25eb955f3c857a6c6879dd2a1063a20ddfae17751ad22fc5438347b

Observation 4c3acee1-4909-4b4f-968c-40287fac6581 · inbound

Holder Policy Optimisation cites this paper.

Holder Policy Optimisation Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:12:22.793408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-13T06:08:28.855671Z digest=sha256:7ba9f532ff51cf168d74f8233bb1e5ef0793adef5f1d19f778820e172523c289

Observation ea8c1185-787b-401f-98a2-5f755ff490cd · inbound

Holder Policy Optimisation cites this paper.

Holder Policy Optimisation Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-22T10:01:22.958441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-22T10:00:58.600743Z digest=sha256:1773ed785b5d57229c9dc185eacd86902ef861bea0e811c52128e2a1cb84fca0

Observation 6a685f5c-6fd3-43b6-8ad2-5f37239a67b5 · inbound

Advantage Collapse in Group Relative Policy Optimization: Diagnosis and Mitigation cites this paper.

Advantage Collapse in Group Relative Policy Optimization: Diagnosis and Mitigation Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T06:09:41.284703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-21T06:08:55.571828Z digest=sha256:5c7f4064a87d6d7e4c047ead039ed487239fbadd224008a830be20a9deb3bc4c

Observation b10868ed-2f85-4319-9de5-65ce2c902b99 · inbound

Advantage Collapse in Group Relative Policy Optimization: Diagnosis and Mitigation cites this paper.

Advantage Collapse in Group Relative Policy Optimization: Diagnosis and Mitigation Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T15:15:47.303183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-30T17:15:53.286925Z digest=sha256:d11bc89e63bd52ffb69c17e806e376e8dfcb2e56a8c4bcf0abbfdb66eb77df78

Observation 131054e5-b69f-4763-ab6b-3bc8dd993145 · inbound

Clipping Bottleneck: Stabilizing RLVR via Stochastic Recovery of Near-Boundary Signals cites this paper.

Clipping Bottleneck: Stabilizing RLVR via Stochastic Recovery of Near-Boundary Signals Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-22T07:51:15.632516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T07:50:44.907952Z digest=sha256:f116dbd7399628e6e289d9dc819185063025cee6415c443e5c72167a00923c87

Observation fd100329-ae05-43c5-ba78-30f6b5c19ba2 · inbound

Not only where, But when: Temporal Scheduling for RLVR cites this paper.

Not only where, But when: Temporal Scheduling for RLVR Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:44:01.829569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-29T22:34:44.791173Z digest=sha256:a9c11d3a7fffa4052b2b58c374c805968c4edf25506d1d46251b812cb60bfca0

Observation 76058b5a-fb62-4cd7-ac8a-b0925351b827 · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs

Reference 239

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:56:13.433277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:baf4d1d40c30b7d670dea8bc3739a66b4b8ee5356254983e69d5efa5a6cc30b7

Observation d296df5d-1a88-43d9-bdd4-bdf8708b5267 · inbound

From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning cites this paper.

From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs

Reference 124

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:38:55.893132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-27T01:13:11.483599Z digest=sha256:8187fa96e4c0a4a3a9fde566303a0233288ad5d6ec52114d27a8128b8c244a42

Observation af7bf6e4-7168-4bd3-8cb3-3bc0f4d082bd · inbound

What are Key Factors for Updates in RL for LLM Reasoning? cites this paper.

What are Key Factors for Updates in RL for LLM Reasoning? Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:09:43.578480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-26T10:24:53.245739Z digest=sha256:df05860b2a79fab6b427eaf4813e92d78654e4c1cdd26b4710026d6a131c3979

Observation de480b0e-a72d-4343-9e30-f44cbd760865 · inbound

When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning cites this paper.

When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-07-10T14:27:07.374977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T14:25:07.006233Z digest=sha256:378862304b4913975ae24378d7c49dc31d107fc894b3c34f0879df66ace49d0b