Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 28 July 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2401.00243.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-07-28T06:31:03.373048+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-12T07:49:57.204875Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-10T06:15:00.866473Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 048bf3ff-78c4-437b-a3d6-cf4f6004b714 · inbound
Functional-level Uncertainty Quantification for Calibrated Fine-tuning on LLMs Uncertainty-Penalized Reinforcement Learning from Human Feedback with Diverse Reward LoRA Ensembles
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 1d70d49c-04c2-4d1a-9053-0af56be2a8a1 · inbound
Wasserstein Distributionally Robust Regret Optimization for Reinforcement Learning from Human Feedback Uncertainty-Penalized Reinforcement Learning from Human Feedback with Diverse Reward LoRA Ensembles
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 49710199-27b5-495b-918b-c2b02f854056 · inbound
A Unifying Lens on Reward Uncertainty in RLHF Uncertainty-Penalized Reinforcement Learning from Human Feedback with Diverse Reward LoRA Ensembles
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 71b7d724-a887-4a74-8164-0c2720d079cb · inbound
Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization Uncertainty-Penalized Reinforcement Learning from Human Feedback with Diverse Reward LoRA Ensembles
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation dc9bb22a-5858-4d1a-b6c9-656377e7607d · inbound
Reinforcement learning to improve large language model-based automated code compliance systems Uncertainty-Penalized Reinforcement Learning from Human Feedback with Diverse Reward LoRA Ensembles
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 18f85772-110c-474e-a952-ca123ce32896 · inbound
Internal Pluralism and the Limits of Pairwise Comparisons Uncertainty-Penalized Reinforcement Learning from Human Feedback with Diverse Reward LoRA Ensembles
Reference 179
Source-reported events for the cited work
Unavailable: canonical work link unavailable.