Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T06:34:30.811523Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 1 inbound Pith citation observation for arXiv:2601.22648.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T06:34:30.811523Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T00:53:26.423440Z
A source-named dated measurement, never combined with another source.
Source: cited_works
23 of 23 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 605cee2b-5d16-4dd8-a92a-4627b74bb8a4 · outbound
UCPO: Uncertainty-Aware Policy Optimization Can AI Assistants Know What They Don't Know?
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f08f8c32-e2bf-4d9c-9b88-cbc0581a328c · outbound
UCPO: Uncertainty-Aware Policy Optimization P., Leang, J
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2bfdbc99-18c0-4f95-a941-a488fd05861a · outbound
UCPO: Uncertainty-Aware Policy Optimization DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a0799e2-24c9-451e-aca2-6aea9e1a51c1 · outbound
UCPO: Uncertainty-Aware Policy Optimization A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f5763b8-3ec8-4ab1-b481-419230e41fb2 · outbound
UCPO: Uncertainty-Aware Policy Optimization AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1433b8d9-02b1-41e6-b1de-b3e4e5bccc9c · outbound
UCPO: Uncertainty-Aware Policy Optimization From System 1 to System 2: A Survey of Reasoning Large Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e745608-7062-4f9c-9a11-36c004a8b6e0 · outbound
UCPO: Uncertainty-Aware Policy Optimization Understanding R1-Zero-Like Training: A Critical Perspective
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e437aee-a72b-487d-82e5-b4de3f4fa9b4 · outbound
UCPO: Uncertainty-Aware Policy Optimization Ngrpo: Negative-enhanced group relative policy optimization
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f512ed6-ca9a-47d8-b001-8db8599267b7 · outbound
UCPO: Uncertainty-Aware Policy Optimization KnowRL: Exploring Knowledgeable Reinforcement Learning for Factuality
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d913b11d-0938-4f15-8e1b-9f467c2fb888 · outbound
UCPO: Uncertainty-Aware Policy Optimization DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64dbf637-8191-4b22-85b7-eb138f78bc3b · outbound
UCPO: Uncertainty-Aware Policy Optimization The Curious Case of Hallucinatory (Un)answerability: Finding Truths in the Hidden States of Over-Confident Large Language Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07c74a56-30da-4241-acee-22357dc2f946 · outbound
UCPO: Uncertainty-Aware Policy Optimization Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e4f6dad-c23b-476e-ba13-b9f248222e08 · outbound
UCPO: Uncertainty-Aware Policy Optimization A Comprehensive Survey of Hallucination Mitigation Techniques in Large Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fcdef762-0f6d-4e7e-9436-2447783ee2f1 · outbound
UCPO: Uncertainty-Aware Policy Optimization TruthRL: Incentivizing Truthful LLMs via Reinforcement Learning
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb98981a-90d9-46d3-9781-ae595a09064b · outbound
UCPO: Uncertainty-Aware Policy Optimization Qwen3 Technical Report
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e50f1400-c5c4-46d5-b9ce-6c0a75ab3b27 · outbound
UCPO: Uncertainty-Aware Policy Optimization Are Reasoning Models More Prone to Hallucination?
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a99d227-4845-400c-9cff-f203d2bd8138 · outbound
UCPO: Uncertainty-Aware Policy Optimization DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd895313-c23f-4a69-9839-681b8117a878 · outbound
UCPO: Uncertainty-Aware Policy Optimization A Survey of Reinforcement Learning for Large Reasoning Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02ba818f-c56b-4a58-8fbd-706e647ed6de · outbound
UCPO: Uncertainty-Aware Policy Optimization The surprising effectiveness of negative reinforcement in llm reasoning.arXiv preprint arXiv:2506.01347,
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c406dff7-26fb-4c27-9210-3e12085cd10e · outbound
UCPO: Uncertainty-Aware Policy Optimization Unresolved cited work
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6adc291c-0df8-4ef8-b116-c208dc6c6f34 · outbound
UCPO: Uncertainty-Aware Policy Optimization Why Language Models Hallucinate
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2709e7c-1776-4bb1-b3f9-9ed4a1d765db · outbound
UCPO: Uncertainty-Aware Policy Optimization Amayuelas, A., Wong, K., Pan, L., Chen, W., and Wang, W
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation adc8f792-31ba-44bb-9ac5-11f2f93e7534 · outbound
UCPO: Uncertainty-Aware Policy Optimization A comprehensive taxonomy of hallucinations in Large Language Models
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04ea5d40-19af-43f6-9220-e1a19046f1ea · inbound
Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning UCPO: Uncertainty-Aware Policy Optimization
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.