Pith. sign in

Paper Citation Record · LEDGER

Uncertainty-Penalized Reinforcement Learning from Human Feedback with Diverse Reward LoRA Ensembles

As of 28 July 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2401.00243.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2401.00243 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-07-28T06:31:03.373048+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-12T07:49:57.204875Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-10T06:15:00.866473Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 048bf3ff-78c4-437b-a3d6-cf4f6004b714 · inbound

Functional-level Uncertainty Quantification for Calibrated Fine-tuning on LLMs cites this paper.

Functional-level Uncertainty Quantification for Calibrated Fine-tuning on LLMs Uncertainty-Penalized Reinforcement Learning from Human Feedback with Diverse Reward LoRA Ensembles

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-23T19:35:47.095212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-23T19:35:29.917096Z digest=sha256:67c9783573665e935af5152e937c9079c17ef5499d8bdd120afe79a6e10d90d7

Observation 1d70d49c-04c2-4d1a-9053-0af56be2a8a1 · inbound

Wasserstein Distributionally Robust Regret Optimization for Reinforcement Learning from Human Feedback cites this paper.

Wasserstein Distributionally Robust Regret Optimization for Reinforcement Learning from Human Feedback Uncertainty-Penalized Reinforcement Learning from Human Feedback with Diverse Reward LoRA Ensembles

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:46:47.024073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-09T21:03:53.045304Z digest=sha256:bdcdd15316d8162d29bb71fba11c2a7bbcaeedcd77cab2e0fd3f418b28dd9959

Observation 49710199-27b5-495b-918b-c2b02f854056 · inbound

A Unifying Lens on Reward Uncertainty in RLHF cites this paper.

A Unifying Lens on Reward Uncertainty in RLHF Uncertainty-Penalized Reinforcement Learning from Human Feedback with Diverse Reward LoRA Ensembles

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T00:27:29.820781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-06-27T17:11:46.549150Z digest=sha256:597af1e248f0f03ea9ced438161126017f80c733df198afba78d5d02ba647fd3

Observation 71b7d724-a887-4a74-8164-0c2720d079cb · inbound

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization cites this paper.

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization Uncertainty-Penalized Reinforcement Learning from Human Feedback with Diverse Reward LoRA Ensembles

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:37:30.424632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=arxiv_source observed=2026-06-27T16:26:34.918099Z digest=sha256:cc36e999c8b95e986bdedf95be4239b4d4acc2e00079018449e670f0c52ecb74

Observation dc9bb22a-5858-4d1a-b6c9-656377e7607d · inbound

Reinforcement learning to improve large language model-based automated code compliance systems cites this paper.

Reinforcement learning to improve large language model-based automated code compliance systems Uncertainty-Penalized Reinforcement Learning from Human Feedback with Diverse Reward LoRA Ensembles

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-06-26T10:19:19.081261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-06-26T10:10:40.301609Z digest=sha256:a3b0be31afc809a43a49172a68306234aef1e190f6ca5dd6417730ea3370d1cc

Observation 18f85772-110c-474e-a952-ca123ce32896 · inbound

Internal Pluralism and the Limits of Pairwise Comparisons cites this paper.

Internal Pluralism and the Limits of Pairwise Comparisons Uncertainty-Penalized Reinforcement Learning from Human Feedback with Diverse Reward LoRA Ensembles

Reference 179

Resolution
unresolved
no resolver link, observed 2026-07-12T07:49:57.204875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T07:49:57.204875Z digest=sha256:361fe214791af900e8e429e3ebab46d451c8ff8d0de7f9ec7d8b27cce96db19b