Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T00:53:30.969336Z
Paper Citation Record · LEDGER
As of 21 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 0 inbound Pith citation observations for arXiv:2608.00301.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T00:53:30.969336Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
51 of 51 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation ab2705ea-ad33-47eb-825f-c2e3b31a29df · outbound
Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning and Zhang, Edwin , year =
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2abb9b30-a985-4b00-943f-e9301274ae68 · outbound
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57826263-223c-4742-be08-037d3faa7ccf · outbound
Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 676c5148-e63d-4158-bbf2-c2f8b97722f6 · outbound
Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning TruthRL: Incentivizing Truthful LLMs via Reinforcement Learning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d0ec2be-5923-4e84-8d32-c4160d977077 · outbound
Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning TIAR: Trajectory-Informed Advantage Reweighting for LLM Abstention Learning
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54305e50-0fb7-46ae-be6e-c870cb08f6e0 · outbound
Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning KARL: Mitigating Hallucinations in LLMs via Knowledge-Boundary-Aware Reinforcement Learning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04ea5d40-19af-43f6-9220-e1a19046f1ea · outbound
Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning UCPO: Uncertainty-Aware Policy Optimization
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fef8b9ba-3f0d-43a2-adfa-7c7b669dd7bb · outbound
Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning Abstain-R1: Calibrated Abstention and Post-Refusal Clarification via Verifiable RL
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f7d78d5-7b3a-4d41-9312-00c744b3dfb1 · outbound
Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning Enhancing Reliability across Short and Long-Form
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74247269-072c-4a8b-a849-8dbaae796dad · outbound
Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning Rejection Improves Reliability: Training LLMs to Refuse Unknown Questions Using RL from Knowledge Feedback
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdec0010-df81-43b6-9070-93d785f61194 · outbound
Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning The Hallucination Tax of Reinforcement Finetuning , eprint =
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5138077-ffb5-40e6-9fb9-6dc1749a82c3 · outbound
Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning 2509.17730 , archivePrefix =
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef39490a-612b-4557-a6b6-ad575cf526ed · outbound
Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning Knowledge-Level Consistency Reinforcement Learning: Dual-Fact Alignment for Long-Form Factuality , eprint =
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3f7dae9-8e02-4c1d-b9b8-aee43e2b6828 · outbound
Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning 2505.13529 , archivePrefix =
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ca9606d-b173-4b24-a852-b638efab5d88 · outbound
Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning Vanishing Gradients in Reinforcement Finetuning of Language Models , eprint =
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48d7edbf-492b-4f17-ad60-91b93f9125f3 · outbound
Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning On the Global Convergence Rates of Softmax Policy Gradient Methods
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4689bc2-5a92-442e-bde5-03b62c4c55e5 · outbound
Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models , eprint =
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05b2a897-3b9e-419b-aadc-8f88a138c232 · outbound
Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning Understanding the Effects of RLHF on LLM Generalisation and Diversity
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b1af7e4-cad4-4e02-8915-1aa5cb3720c6 · outbound
Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning A Theory of Regularized Markov Decision Processes
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 352f964e-166d-4f4f-8e61-4f50a47f5763 · outbound
Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning Leverage the Average: an Analysis of KL Regularization in RL
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5cb1418d-f3ba-4e48-abb8-6c1abf01582c · outbound
Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34d3b5f7-0649-4dcc-af3c-139723f97f59 · outbound
Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning Understanding R1-Zero-Like Training: A Critical Perspective
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab733218-cc76-46b6-9b54-775408158617 · outbound
Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4773189-261c-489d-8eb2-2ad401d27a4e · outbound
Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f5698f8-f765-4fbe-a238-c6f77a46ffbd · outbound
Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a2fb1da-8723-4123-9238-55217ffb70fc · outbound
Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning Scaling Laws for Reward Model Overoptimization , eprint =
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2005d49-85be-4b61-a798-a3125ee9fea5 · outbound
Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning Unresolved cited work
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5cea912c-8fb8-426c-9511-ada9ed84a4ae · outbound
Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning and Martic, Miljan and others , year =
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da372744-d219-44ed-9d85-ee50eaf7f8d6 · outbound
Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning and others , year =
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a66bdae-0c77-466d-93b0-ff8b2785630e · outbound
Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning Training Language Models to Follow Instructions with Human Feedback , eprint =
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94218d1a-a73b-44d1-9a4f-e0c056e33543 · outbound
Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning Unresolved cited work
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9ebbbce-2630-4dc0-83ea-267f841fa17e · outbound
Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning On the Foundations of Noise-free Selective Classification , journal =
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 312536c2-1da6-43f1-b6f6-0a97df50dd84 · outbound
Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning Selective Classification for Deep Neural Networks , eprint =
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6d3ec95-38a8-4e46-92a4-1d6507a674c2 · outbound
Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning Selective Question Answering under Domain Shift , eprint =
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84154b10-a95b-4b4a-b71b-8a7515f75373 · outbound
Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning Out-of-Distribution Detection and Selective Generation for Conditional Language Models , eprint =
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82234b30-2e58-45fa-8300-7f711880624d · outbound
Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning Language Models (Mostly) Know What They Know , eprint =
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5fb6dfa4-3519-424a-b12a-94c35fdf8497 · outbound
Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning Teaching Models to Express Their Uncertainty in Words , eprint =
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3793a836-01b9-40e1-b150-26d2d3d9b44a · outbound
Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback , eprint =
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6bb6a822-fbde-45a1-920e-d56aa29c8c99 · outbound
Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning Reported Confidence in LLMs Tracks Commitment More Than Correctness
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f7a097b-94e0-402c-8c92-2de6209a4dc2 · outbound
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae44adc7-85fd-42af-afc3-f4f6f889dd0e · outbound
Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generation , eprint =
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3e7f2be-5139-4f2a-b3f4-b3c0afd1ee4c · outbound
Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning Mitigating LLM Hallucinations via Conformal Abstention
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc84d3f3-b641-4ecf-89fd-0dea84786c69 · outbound
Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning Qwen2.5 Technical Report
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 894f0cf3-5f82-4d9f-8219-47967dfb7db4 · outbound
Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a01facf-b339-482a-9fbf-418f418d2f16 · outbound
Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories , eprint =
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24ce766d-cdd2-4133-991c-c12bf47fb26d · outbound
Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning 2511.13029 , archivePrefix =
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5a2ff5f-f9dc-425f-acd1-3dc3e29d3f8c · outbound
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 264e1248-4b57-4457-b282-25a819bc00a1 · outbound
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f3000b0-2145-4245-b08d-d9a984197b8a · outbound
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3efdc51-f862-4439-8739-f1cf7d1215df · outbound
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 681405dc-2711-4b67-a25b-d529ac63e90d · outbound
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.