Pith. sign in

Paper Citation Record · LEDGER

Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning

As of 21 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 0 inbound Pith citation observations for arXiv:2608.00301.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.00301 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T00:53:30.969336Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

51 of 51 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved51
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ab2705ea-ad33-47eb-825f-c2e3b31a29df · outbound

This paper cites and Zhang, Edwin , year =.

Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning and Zhang, Edwin , year =

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T00:53:25.817060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:53:25.817060Z digest=sha256:d3f7003c2e4e05f0873b2dcf3b31a889e5c49315a3446f7d8c6819a2b17166f8

Observation 2abb9b30-a985-4b00-943f-e9301274ae68 · outbound

This paper cites , year =.

Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning , year =

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T00:53:25.888007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:53:25.888007Z digest=sha256:26a4dde38932c5204034aa435779f7531460c2683b8eb56b3f738e70848ee9d1

Observation 57826263-223c-4742-be08-037d3faa7ccf · outbound

This paper cites Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty.

Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T00:53:26.000469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:53:26.000469Z digest=sha256:7d76cfbb4b38ba15098db05c38de3dfd68acba0b0421c69c01cea6e8568c4f07

Observation 676c5148-e63d-4158-bbf2-c2f8b97722f6 · outbound

This paper cites TruthRL: Incentivizing Truthful LLMs via Reinforcement Learning.

Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning TruthRL: Incentivizing Truthful LLMs via Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T00:53:26.113437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:53:26.113437Z digest=sha256:072d1f360837a11dd53175638f75765427b7e96293b71043f10c63f2a756a729

Observation 1d0ec2be-5923-4e84-8d32-c4160d977077 · outbound

This paper cites TIAR: Trajectory-Informed Advantage Reweighting for LLM Abstention Learning.

Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning TIAR: Trajectory-Informed Advantage Reweighting for LLM Abstention Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T00:53:26.206601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:53:26.206601Z digest=sha256:ab59f4e6b0765ed70002bef202849a84e6209bee47e87285c61cd6cca56f9d8c

Observation 54305e50-0fb7-46ae-be6e-c870cb08f6e0 · outbound

This paper cites KARL: Mitigating Hallucinations in LLMs via Knowledge-Boundary-Aware Reinforcement Learning.

Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning KARL: Mitigating Hallucinations in LLMs via Knowledge-Boundary-Aware Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T00:53:26.307965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:53:26.307965Z digest=sha256:949ce1464c29256b983e4c9d73bfc976a541c4f6fa16d5f19cbf31a7a6aaa2ca

Observation 04ea5d40-19af-43f6-9220-e1a19046f1ea · outbound

This paper cites UCPO: Uncertainty-Aware Policy Optimization.

Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning UCPO: Uncertainty-Aware Policy Optimization

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T00:53:26.423440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:53:26.423440Z digest=sha256:6752761e1bda9dbbcf202cf930c949b055ee02e0f8e3f12e5b6a40f12e75885f

Observation fef8b9ba-3f0d-43a2-adfa-7c7b669dd7bb · outbound

This paper cites Abstain-R1: Calibrated Abstention and Post-Refusal Clarification via Verifiable RL.

Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning Abstain-R1: Calibrated Abstention and Post-Refusal Clarification via Verifiable RL

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T00:53:26.556359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:53:26.556359Z digest=sha256:6ab73d70ff51eaea5929c728039679f6e9cc1d5d62841770df0502508342ffd0

Observation 3f7d78d5-7b3a-4d41-9312-00c744b3dfb1 · outbound

This paper cites Enhancing Reliability across Short and Long-Form.

Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning Enhancing Reliability across Short and Long-Form

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T00:53:26.707407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:53:26.707407Z digest=sha256:47b93be7f241ef2385f7d7cd3ae0cff332653ca9567170896d9a7e3d611af47e

Observation 74247269-072c-4a8b-a849-8dbaae796dad · outbound

This paper cites Rejection Improves Reliability: Training LLMs to Refuse Unknown Questions Using RL from Knowledge Feedback.

Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning Rejection Improves Reliability: Training LLMs to Refuse Unknown Questions Using RL from Knowledge Feedback

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T00:53:26.819594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:53:26.819594Z digest=sha256:09efb8e9fccd383fe8401cc12aa2d7a27c0cce6d6f808e22f9f00d89dd55f48b

Observation fdec0010-df81-43b6-9070-93d785f61194 · outbound

This paper cites The Hallucination Tax of Reinforcement Finetuning , eprint =.

Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning The Hallucination Tax of Reinforcement Finetuning , eprint =

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T00:53:26.884202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:53:26.884202Z digest=sha256:cde525131f38d4d3b4e32d34ee793a66d2ff4245c324d0baa52fe913bb5d2d02

Observation d5138077-ffb5-40e6-9fb9-6dc1749a82c3 · outbound

This paper cites 2509.17730 , archivePrefix =.

Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning 2509.17730 , archivePrefix =

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T00:53:26.972189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:53:26.972189Z digest=sha256:6d5bdd84ce5a2f17b8d24705b8b89d9d73197c5e92c17b26163bd67498267c99

Observation ef39490a-612b-4557-a6b6-ad575cf526ed · outbound

This paper cites Knowledge-Level Consistency Reinforcement Learning: Dual-Fact Alignment for Long-Form Factuality , eprint =.

Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning Knowledge-Level Consistency Reinforcement Learning: Dual-Fact Alignment for Long-Form Factuality , eprint =

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T00:53:27.061604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:53:27.061604Z digest=sha256:daf25ea00892347b5c0975b377dc8d4e3f102a980dbbf36e9f5620beb4eb9c98

Observation c3f7dae9-8e02-4c1d-b9b8-aee43e2b6828 · outbound

This paper cites 2505.13529 , archivePrefix =.

Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning 2505.13529 , archivePrefix =

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T00:53:27.195630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:53:27.195630Z digest=sha256:5f71786d0c3fb6c16080acfce2ea8405802502f03f47ea9a5734fa10c1c9c17f

Observation 9ca9606d-b173-4b24-a852-b638efab5d88 · outbound

This paper cites Vanishing Gradients in Reinforcement Finetuning of Language Models , eprint =.

Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning Vanishing Gradients in Reinforcement Finetuning of Language Models , eprint =

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T00:53:27.338083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:53:27.338083Z digest=sha256:58ce61dc6cc4ba5e109c439585755308115be1856606216b10f7a04345e3c74f

Observation 48d7edbf-492b-4f17-ad60-91b93f9125f3 · outbound

This paper cites On the Global Convergence Rates of Softmax Policy Gradient Methods.

Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning On the Global Convergence Rates of Softmax Policy Gradient Methods

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T00:53:27.440871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:53:27.440871Z digest=sha256:c50e2a7ce009b14f853f03a39830e291ea38ee6030a5d74d033c331946ab3b76

Observation c4689bc2-5a92-442e-bde5-03b62c4c55e5 · outbound

This paper cites The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models , eprint =.

Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models , eprint =

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T00:53:27.529142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:53:27.529142Z digest=sha256:966146fc3f53ed298d0bcb8467397912b81a772965ed047e0fbd97cf8945b58a

Observation 05b2a897-3b9e-419b-aadc-8f88a138c232 · outbound

This paper cites Understanding the Effects of RLHF on LLM Generalisation and Diversity.

Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T00:53:27.653093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:53:27.653093Z digest=sha256:ef12450921122d7136c0de8fceafb1e6bdddff82522b133f0d660292dc36a968

Observation 2b1af7e4-cad4-4e02-8915-1aa5cb3720c6 · outbound

This paper cites A Theory of Regularized Markov Decision Processes.

Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning A Theory of Regularized Markov Decision Processes

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T00:53:27.851497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:53:27.851497Z digest=sha256:4dec4266a11328c7fb6f8ba1da5ceb48c433cd4862b6139d8270e0dc5fce1604

Observation 352f964e-166d-4f4f-8e61-4f50a47f5763 · outbound

This paper cites Leverage the Average: an Analysis of KL Regularization in RL.

Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning Leverage the Average: an Analysis of KL Regularization in RL

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T00:53:27.989253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:53:27.989253Z digest=sha256:145c2ce73884d637480007faaca0e413dced3754ad9c97735cce7c511e4bf140

Observation 5cb1418d-f3ba-4e48-abb8-6c1abf01582c · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T00:53:28.131447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:53:28.131447Z digest=sha256:34c2d37b98ace0433f8e7a99f18a4d38dc3fa1b270bbffdeffa765d94dbfecf5

Observation 34d3b5f7-0649-4dcc-af3c-139723f97f59 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning Understanding R1-Zero-Like Training: A Critical Perspective

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T00:53:28.214601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:53:28.214601Z digest=sha256:4d7f671d53aff5c9b8f24e558e5cff8d372e6a56264cbea3b600ddef2968b876

Observation ab733218-cc76-46b6-9b54-775408158617 · outbound

This paper cites Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs.

Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T00:53:28.299798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:53:28.299798Z digest=sha256:7ffad208d923c74c1836aff6a60ba53a981cbde9a6873dc2672ee669eeae38f2

Observation b4773189-261c-489d-8eb2-2ad401d27a4e · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T00:53:28.428593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:53:28.428593Z digest=sha256:272913741b580e318876f873726732553190692c4d0099bd7edeb2465a4cceff

Observation 2f5698f8-f765-4fbe-a238-c6f77a46ffbd · outbound

This paper cites REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization.

Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T00:53:28.530409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:53:28.530409Z digest=sha256:51bf5d206bd3266321ffbc929d577035533eed73ec4d5acc0b9f4bdc089bb361

Observation 4a2fb1da-8723-4123-9238-55217ffb70fc · outbound

This paper cites Scaling Laws for Reward Model Overoptimization , eprint =.

Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning Scaling Laws for Reward Model Overoptimization , eprint =

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T00:53:28.668683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:53:28.668683Z digest=sha256:85e05fc908d74459e53a5dd338206767e51474bd66c9f3e45068ea79b3f7bcfb

Observation f2005d49-85be-4b61-a798-a3125ee9fea5 · outbound

This paper cites an unresolved cited work.

Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T00:53:28.847058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:53:28.847058Z digest=sha256:3e54a379fdbefdf6ff67f747789e97ccb1548c6ed09e6659574b6fb9486db0ae

Observation 5cea912c-8fb8-426c-9511-ada9ed84a4ae · outbound

This paper cites and Martic, Miljan and others , year =.

Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning and Martic, Miljan and others , year =

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T00:53:28.926687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:53:28.926687Z digest=sha256:7a79eb254ffe5495198a2d8f42ac75bcabc5b4346eb450e39298c667c20861fa

Observation da372744-d219-44ed-9d85-ee50eaf7f8d6 · outbound

This paper cites and others , year =.

Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning and others , year =

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T00:53:28.981204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:53:28.981204Z digest=sha256:291ecf423194ac55e6d91d74bc96484238cf589605d7d09edacc18669831de37

Observation 6a66bdae-0c77-466d-93b0-ff8b2785630e · outbound

This paper cites Training Language Models to Follow Instructions with Human Feedback , eprint =.

Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning Training Language Models to Follow Instructions with Human Feedback , eprint =

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T00:53:29.084946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:53:29.084946Z digest=sha256:7d419aef7d8fd90ff84991bde225b3afc5c47bd0e1225cf51f201adda1550638

Observation 94218d1a-a73b-44d1-9a4f-e0c056e33543 · outbound

This paper cites an unresolved cited work.

Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T00:53:29.148434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:53:29.148434Z digest=sha256:5909cfb0f761551c433fc39b095cda2488a7c7f36c72a2c86bdd48120d63d8f8

Observation d9ebbbce-2630-4dc0-83ea-267f841fa17e · outbound

This paper cites On the Foundations of Noise-free Selective Classification , journal =.

Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning On the Foundations of Noise-free Selective Classification , journal =

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T00:53:29.203476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:53:29.203476Z digest=sha256:e79d2452b0023ed6c1b89d9aa9d26d41287bd77554375847b43639d47e96f001

Observation 312536c2-1da6-43f1-b6f6-0a97df50dd84 · outbound

This paper cites Selective Classification for Deep Neural Networks , eprint =.

Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning Selective Classification for Deep Neural Networks , eprint =

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T00:53:29.208690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:53:29.208690Z digest=sha256:9102887ce188e6c84e71d1ba3d9f8697588011de2732461f6c7732e653cf139c

Observation b6d3ec95-38a8-4e46-92a4-1d6507a674c2 · outbound

This paper cites Selective Question Answering under Domain Shift , eprint =.

Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning Selective Question Answering under Domain Shift , eprint =

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T00:53:29.275043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:53:29.275043Z digest=sha256:2cf71c56555aa8f8d8d897a5b25de47471a6138f14c5cd8a8e9e7f9ab7a01a15

Observation 84154b10-a95b-4b4a-b71b-8a7515f75373 · outbound

This paper cites Out-of-Distribution Detection and Selective Generation for Conditional Language Models , eprint =.

Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning Out-of-Distribution Detection and Selective Generation for Conditional Language Models , eprint =

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T00:53:29.487503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:53:29.487503Z digest=sha256:e9f80dcc97717a9da93a2638995594c1eceb21e5533c901058bdbc50011a39c2

Observation 82234b30-2e58-45fa-8300-7f711880624d · outbound

This paper cites Language Models (Mostly) Know What They Know , eprint =.

Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning Language Models (Mostly) Know What They Know , eprint =

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T00:53:29.603289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:53:29.603289Z digest=sha256:3db9d7c9a3e6a976a5166025cebf7c5159ec51ce455217a43a89522c5fa835fd

Observation 5fb6dfa4-3519-424a-b12a-94c35fdf8497 · outbound

This paper cites Teaching Models to Express Their Uncertainty in Words , eprint =.

Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning Teaching Models to Express Their Uncertainty in Words , eprint =

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T00:53:29.696163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:53:29.696163Z digest=sha256:1bee7e0a154e0273adb66340a090ba67aa928a18e6bfeab2bbd88ed665709f2f

Observation 3793a836-01b9-40e1-b150-26d2d3d9b44a · outbound

This paper cites Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback , eprint =.

Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback , eprint =

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T00:53:29.836202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:53:29.836202Z digest=sha256:fb375364f63eca68425cc1f612efb65a3361f28e005197e16ce3b00a524c4704

Observation 6bb6a822-fbde-45a1-920e-d56aa29c8c99 · outbound

This paper cites Reported Confidence in LLMs Tracks Commitment More Than Correctness.

Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning Reported Confidence in LLMs Tracks Commitment More Than Correctness

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T00:53:29.955373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:53:29.955373Z digest=sha256:ee64f94ed0e694afb499d8daf678e2218e7a874e3e4c794b7c956eb529ee1383

Observation 2f7a097b-94e0-402c-8c92-2de6209a4dc2 · outbound

This paper cites , year =.

Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning , year =

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T00:53:30.086006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:53:30.086006Z digest=sha256:0455e353120fdc0e9f429fd160b81a448e604f7903f4c71aa44b813c93c5aa0e

Observation ae44adc7-85fd-42af-afc3-f4f6f889dd0e · outbound

This paper cites Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generation , eprint =.

Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generation , eprint =

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T00:53:30.187212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:53:30.187212Z digest=sha256:5f0ac3336c3774ec3cf276609c07b6935a9c7fd65f8b5600611a9f46f38f86f9

Observation e3e7f2be-5139-4f2a-b3f4-b3c0afd1ee4c · outbound

This paper cites Mitigating LLM Hallucinations via Conformal Abstention.

Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning Mitigating LLM Hallucinations via Conformal Abstention

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T00:53:30.253798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:53:30.253798Z digest=sha256:40dddc1703d03df7a22dbe8ddb8948aec04c8c4e0c84b50b55faec5ad1baf5f4

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T00:53:30.337364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:53:30.337364Z digest=sha256:63ab2e5c417beb4ce42e57bad4173f3118830c6a92de63c4e566c0d775f6fb93

Observation 894f0cf3-5f82-4d9f-8219-47967dfb7db4 · outbound

This paper cites TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension.

Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T00:53:30.434651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:53:30.434651Z digest=sha256:66ebc765e59013494c2f38247c0c82df305a5e909aac55932132347156e356a0

Observation 9a01facf-b339-482a-9fbf-418f418d2f16 · outbound

This paper cites When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories , eprint =.

Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories , eprint =

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T00:53:30.529942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:53:30.529942Z digest=sha256:17397b7a9a83b4eccec7872238aeecc56bf4a1186a7b6018414d67836a916980

Observation 24ce766d-cdd2-4133-991c-c12bf47fb26d · outbound

This paper cites 2511.13029 , archivePrefix =.

Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning 2511.13029 , archivePrefix =

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-04T00:53:30.589609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:53:30.589609Z digest=sha256:4af82bacedef850f785fb48897c53c38bfa6b75a1e0f8b4f43cd07a7bf66e160

Observation c5a2ff5f-f9dc-425f-acd1-3dc3e29d3f8c · outbound

This paper cites , year =.

Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning , year =

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T00:53:30.694160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:53:30.694160Z digest=sha256:2a3efedecec20a36274863d993e6cbf238351641fbfaa6664df0549821cdb746

Observation 264e1248-4b57-4457-b282-25a819bc00a1 · outbound

This paper cites , year =.

Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning , year =

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T00:53:30.769359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:53:30.769359Z digest=sha256:61f7a397e8042be0d717630cde8cf66c107aee369d876d520cd2054fe1e60c21

Observation 8f3000b0-2145-4245-b08d-d9a984197b8a · outbound

This paper cites , year =.

Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning , year =

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-04T00:53:30.816895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:53:30.816895Z digest=sha256:d0edbabe5f433be46df892ab813ae7a558180705ad482368ef833c1caac36d75

Observation d3efdc51-f862-4439-8739-f1cf7d1215df · outbound

This paper cites , year =.

Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning , year =

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T00:53:30.876999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:53:30.876999Z digest=sha256:575760d108bfd3960ae32c5d82d793c6a4f064c7ba4d7b012b37b2827b02dd84

Observation 681405dc-2711-4b67-a25b-d529ac63e90d · outbound

This paper cites Approximating.

Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning Approximating

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T00:53:30.969336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:53:30.969336Z digest=sha256:f30a85fff8455f54dba8328e7fd0d5fe8ac9cb8b7acc44253ab35092e4fd7275

Pith citing papers

No inbound Pith citation observations are available.