Pith. sign in

Paper Citation Record · LEDGER

UCPO: Uncertainty-Aware Policy Optimization

As of 13 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 1 inbound Pith citation observation for arXiv:2601.22648.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2601.22648 v2

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T06:34:30.811523Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T00:53:26.423440Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

23 of 23 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved22
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 605cee2b-5d16-4dd8-a92a-4627b74bb8a4 · outbound

This paper cites Can AI Assistants Know What They Don't Know?.

UCPO: Uncertainty-Aware Policy Optimization Can AI Assistants Know What They Don't Know?

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T06:34:28.294064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:34:28.294064Z digest=sha256:8cc8306416dadcffe6afc392113c20c4de6b3080cc9af1a3b52e8dac7c5d03f7

Observation f08f8c32-e2bf-4d9c-9b88-cbc0581a328c · outbound

This paper cites P., Leang, J.

UCPO: Uncertainty-Aware Policy Optimization P., Leang, J

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T06:34:28.517291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:34:28.517291Z digest=sha256:835855016fc1cfce6b4c1cb91dc81381efa37dbdccaf92f32abeb27436c344fa

Observation 2bfdbc99-18c0-4f95-a941-a488fd05861a · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

UCPO: Uncertainty-Aware Policy Optimization DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T06:34:28.636359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:34:28.636359Z digest=sha256:12439e2d86e2a1396b598e178e050b5abba11dfc1e88abf6b8fe451ff6ffc576

Observation 9a0799e2-24c9-451e-aca2-6aea9e1a51c1 · outbound

This paper cites A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions.

UCPO: Uncertainty-Aware Policy Optimization A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T06:34:28.805851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:34:28.805851Z digest=sha256:666078125e487057cab79e6e8f9fce34a44f187af2ca66231986a2ec41038f3c

Observation 3f5763b8-3ec8-4ab1-b481-419230e41fb2 · outbound

This paper cites AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions.

UCPO: Uncertainty-Aware Policy Optimization AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T06:34:29.025920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:34:29.025920Z digest=sha256:01dd91e4a1ef22c8c36adf5e0f2400c441d76358ad837e963cc1925169a4b170

Observation 1433b8d9-02b1-41e6-b1de-b3e4e5bccc9c · outbound

This paper cites From System 1 to System 2: A Survey of Reasoning Large Language Models.

UCPO: Uncertainty-Aware Policy Optimization From System 1 to System 2: A Survey of Reasoning Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T06:34:29.146626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:34:29.146626Z digest=sha256:643be6c9e9dce32fd8a293070e62ed08cd0dadf7df98a49d45ecd8dead172325

Observation 2e745608-7062-4f9c-9a11-36c004a8b6e0 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

UCPO: Uncertainty-Aware Policy Optimization Understanding R1-Zero-Like Training: A Critical Perspective

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T06:34:29.264212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:34:29.264212Z digest=sha256:e38005a0ad9b91a5b93f94d5de58b9b085fb8a0aeafc456fd1f3823ddaa749ad

Observation 8e437aee-a72b-487d-82e5-b4de3f4fa9b4 · outbound

This paper cites Ngrpo: Negative-enhanced group relative policy optimization.

UCPO: Uncertainty-Aware Policy Optimization Ngrpo: Negative-enhanced group relative policy optimization

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T06:34:29.435633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:34:29.435633Z digest=sha256:93cb75576d897f1160b8e66efdbb9742dac8d16a08ae45f1679ba26bcdea95fa

Observation 7f512ed6-ca9a-47d8-b001-8db8599267b7 · outbound

This paper cites KnowRL: Exploring Knowledgeable Reinforcement Learning for Factuality.

UCPO: Uncertainty-Aware Policy Optimization KnowRL: Exploring Knowledgeable Reinforcement Learning for Factuality

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T06:34:29.551730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:34:29.551730Z digest=sha256:61179405fe098310fe6eabcd12a11c985f6ee983309cfd0326a41f8be9c9789e

Observation d913b11d-0938-4f15-8e1b-9f467c2fb888 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

UCPO: Uncertainty-Aware Policy Optimization DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T06:34:29.722785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:34:29.722785Z digest=sha256:d1b21c185b842ae21256f859ee4155391dca238d4372a44fdb36e826697cde6e

Observation 64dbf637-8191-4b22-85b7-eb138f78bc3b · outbound

This paper cites The Curious Case of Hallucinatory (Un)answerability: Finding Truths in the Hidden States of Over-Confident Large Language Models.

UCPO: Uncertainty-Aware Policy Optimization The Curious Case of Hallucinatory (Un)answerability: Finding Truths in the Hidden States of Over-Confident Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T06:34:29.871108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:34:29.871108Z digest=sha256:a5f495157a4706446d7a0d71f9b40211296d5f58d9c22a800260e9913e0f53f8

Observation 07c74a56-30da-4241-acee-22357dc2f946 · outbound

This paper cites Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback.

UCPO: Uncertainty-Aware Policy Optimization Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T06:34:29.961428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:34:29.961428Z digest=sha256:633dcdf586557924e78ecf87a92c87e2e6b6175a30cd5d4ebcc0ef17237529f0

Observation 1e4f6dad-c23b-476e-ba13-b9f248222e08 · outbound

This paper cites A Comprehensive Survey of Hallucination Mitigation Techniques in Large Language Models.

UCPO: Uncertainty-Aware Policy Optimization A Comprehensive Survey of Hallucination Mitigation Techniques in Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T06:34:30.049121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:34:30.049121Z digest=sha256:670b30f57bc041f9f563f467c25158db438fa79e2b72d0da7915d7654368e0eb

Observation fcdef762-0f6d-4e7e-9436-2447783ee2f1 · outbound

This paper cites TruthRL: Incentivizing Truthful LLMs via Reinforcement Learning.

UCPO: Uncertainty-Aware Policy Optimization TruthRL: Incentivizing Truthful LLMs via Reinforcement Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T06:34:30.106787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:34:30.106787Z digest=sha256:05bfd7cd880b166866cae5bd9a3be05c9d5f3c810423c1dab4dd938f219f1b73

Observation eb98981a-90d9-46d3-9781-ae595a09064b · outbound

This paper cites Qwen3 Technical Report.

UCPO: Uncertainty-Aware Policy Optimization Qwen3 Technical Report

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T06:34:30.194123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:34:30.194123Z digest=sha256:309d230c15696badc41b267530dd22d183879328e87e9659994f96329c4c147c

Observation e50f1400-c5c4-46d5-b9ce-6c0a75ab3b27 · outbound

This paper cites Are Reasoning Models More Prone to Hallucination?.

UCPO: Uncertainty-Aware Policy Optimization Are Reasoning Models More Prone to Hallucination?

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T06:34:30.368552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:34:30.368552Z digest=sha256:9161d95382314345276983e2951e8b9c1cf5a3c25af00eab5e67f6ca14220248

Observation 2a99d227-4845-400c-9cff-f203d2bd8138 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

UCPO: Uncertainty-Aware Policy Optimization DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T06:34:30.480875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:34:30.480875Z digest=sha256:7ab7897c5169199b62c5d20f568ef4b91dcde064e7240bf693ab9cb410b1a0c4

Observation fd895313-c23f-4a69-9839-681b8117a878 · outbound

This paper cites A Survey of Reinforcement Learning for Large Reasoning Models.

UCPO: Uncertainty-Aware Policy Optimization A Survey of Reinforcement Learning for Large Reasoning Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T06:34:30.657970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:34:30.657970Z digest=sha256:d350b640ab92be4b7f467cf4195c581439a04bd4010c908426c0ae04ca5d9139

Observation 02ba818f-c56b-4a58-8fbd-706e647ed6de · outbound

This paper cites The surprising effectiveness of negative reinforcement in llm reasoning.arXiv preprint arXiv:2506.01347,.

UCPO: Uncertainty-Aware Policy Optimization The surprising effectiveness of negative reinforcement in llm reasoning.arXiv preprint arXiv:2506.01347,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T06:34:30.761063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:34:30.761063Z digest=sha256:e6a38b85c4ca984a3bbffbac68445cf92ba0b9d7a496910bc95b717a52330f16

Observation c406dff7-26fb-4c27-9210-3e12085cd10e · outbound

This paper cites an unresolved cited work.

UCPO: Uncertainty-Aware Policy Optimization Unresolved cited work

Reference 23

Resolution
malformed identifier
no resolver link, observed 2026-08-03T06:34:30.811523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:34:30.811523Z digest=sha256:bdb9466d18475aeb1551fbd72e4703fe2a9208b4b26e99593ad207efc3860e7f

Observation 6adc291c-0df8-4ef8-b116-c208dc6c6f34 · outbound

This paper cites Why Language Models Hallucinate.

UCPO: Uncertainty-Aware Policy Optimization Why Language Models Hallucinate

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T06:34:28.912210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:34:28.912210Z digest=sha256:75e561938cb610c5f05eb5f52daf1993813e3b78c0455e0e71c9fd9ae8a8eed9

Observation d2709e7c-1776-4bb1-b3f9-9ed4a1d765db · outbound

This paper cites Amayuelas, A., Wong, K., Pan, L., Chen, W., and Wang, W.

UCPO: Uncertainty-Aware Policy Optimization Amayuelas, A., Wong, K., Pan, L., Chen, W., and Wang, W

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T06:34:28.169956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:34:28.169956Z digest=sha256:71e7c53a488d6194331e374939cdaec04d2a5b2705fb8f22f702e229022261f1

Observation adc8f792-31ba-44bb-9ac5-11f2f93e7534 · outbound

This paper cites A comprehensive taxonomy of hallucinations in Large Language Models.

UCPO: Uncertainty-Aware Policy Optimization A comprehensive taxonomy of hallucinations in Large Language Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T06:34:28.406600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:34:28.406600Z digest=sha256:f251419878563abd7ad65dcb06e1019bc7a1ddb00884367697ed75d3923e0e05

Pith citing papers

Observation 04ea5d40-19af-43f6-9220-e1a19046f1ea · inbound

Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning cites this paper.

Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning UCPO: Uncertainty-Aware Policy Optimization

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T00:53:26.423440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:53:26.423440Z digest=sha256:aa95d4bfd02bd091634957b1e252b08d71abfe7f9895818ff22e56a4b373a06f