Pith. sign in

Paper Citation Record · LEDGER

Towards Reliable, Uncertainty-Aware Alignment

As of 13 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 0 inbound Pith citation observations for arXiv:2507.15906.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.15906 v1

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:40:18.315140Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

59 of 59 outbound references displayed

  • verified exact1
  • verified fuzzy12
  • unresolved46
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d809cf6f-ee7c-4ffd-892d-e7588102b78c · outbound

This paper cites The claude 3 model family: Opus, sonnet, haiku.

Towards Reliable, Uncertainty-Aware Alignment The claude 3 model family: Opus, sonnet, haiku

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:21.156461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-06T15:40:13.176395Z digest=sha256:76c156900a2e501cbbe2de98b92b82e2b00dd15c377a1a7780d0b82f10f61a44

Observation 47697553-f704-4bb3-82e3-bfc825ef34c7 · outbound

This paper cites The Llama 3 Herd of Models.

Towards Reliable, Uncertainty-Aware Alignment The Llama 3 Herd of Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:13.266168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:13.266168Z digest=sha256:b0dcbf122750b7b04b4220be32ab84e6dd6ad699f055bc8c2b31424a6835d21a

Observation ee77b394-d87e-49f6-8280-3c14d712c9cb · outbound

This paper cites Concrete Problems in AI Safety.

Towards Reliable, Uncertainty-Aware Alignment Concrete Problems in AI Safety

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:13.333386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:13.333386Z digest=sha256:e83fad71f6d416d131dd389343240ee5ba2ecf496f5cfb37c0545bd913e5a78a

Observation fbb27280-7859-4c32-8d4f-a42d30e3c2f0 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Towards Reliable, Uncertainty-Aware Alignment Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:13.417979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:13.417979Z digest=sha256:bf781b77b82de7cd2a8574aaab4095a203a324a47ab034df9e4fe6cb840b2147

Observation 48ae5dfa-eb9f-465e-aa15-b6cc4542ab59 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Towards Reliable, Uncertainty-Aware Alignment Constitutional AI: Harmlessness from AI Feedback

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:13.494528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:13.494528Z digest=sha256:9ed7d23db408789af5f87fdde176ae95a1d92a5e9386d3c97cfe72637ad189df

Observation 1728c136-ef99-4e4d-a6ff-0be5e299c464 · outbound

This paper cites Deep reinforcement learning from human preferences.

Towards Reliable, Uncertainty-Aware Alignment Deep reinforcement learning from human preferences

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:13.621778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:13.621778Z digest=sha256:63e8f9a6414e9e9ab835818402b2dfa4ec835b40ab43fdb0852b914ae7c55613

Observation d81661bf-1992-4d60-91f9-a40889259d1e · outbound

This paper cites Reward Model Ensembles Help Mitigate Overoptimization.

Towards Reliable, Uncertainty-Aware Alignment Reward Model Ensembles Help Mitigate Overoptimization

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:13.721341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:13.721341Z digest=sha256:2b9c4d3a399f64144e2fe255d923a8c565cdc034b266ced3eb3bb8dd24629f0f

Observation 08da191c-c502-4312-bed6-4226835d001c · outbound

This paper cites UltraFeedback: Boosting Language Models with Scaled AI Feedback.

Towards Reliable, Uncertainty-Aware Alignment UltraFeedback: Boosting Language Models with Scaled AI Feedback

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:13.818475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:13.818475Z digest=sha256:0f5516449c24edb50f1470717da974769978292d66595082ebacea9e57631771

Observation 1d279d30-0dcf-4f19-bfaa-89c952a0a89f · outbound

This paper cites Suphavadeeprasit.

Towards Reliable, Uncertainty-Aware Alignment Suphavadeeprasit

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:20.952919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-06T15:40:13.865665Z digest=sha256:dc2cefcd62102ac6229bfa871a0faf29681e9eb99f51ed8192c5321f46d0206f

Observation 5e10c0f0-fb17-4bfe-9a8f-44419e76630c · outbound

This paper cites Gemini 2.0 flash: Next-generation multimodal ai model.

Towards Reliable, Uncertainty-Aware Alignment Gemini 2.0 flash: Next-generation multimodal ai model

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:20.783795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-06T15:40:13.944596Z digest=sha256:4a08d8eea8d973676f222d8cf25aa68eef128ac612cd94b67e971300c5defa6a

Observation 1361e871-f6d0-4c4d-9c6c-b5e0025049e2 · outbound

This paper cites DeepSeek-V3 Technical Report.

Towards Reliable, Uncertainty-Aware Alignment DeepSeek-V3 Technical Report

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:14.010645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:14.010645Z digest=sha256:4b8f26001603d23b6512e6de4db7631cd649a274e7ced3c2771104a49de1e5b3

Observation 97e0d881-e718-487f-a5bb-1e108dcacce4 · outbound

This paper cites RAFT : Reward ranked finetuning for generative foundation model alignment.

Towards Reliable, Uncertainty-Aware Alignment RAFT : Reward ranked finetuning for generative foundation model alignment

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:14.132426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:14.132426Z digest=sha256:c7a0bbafe9d12d69e0de9a52e97491d3f6e76f21f9f01909a89a5e2891fb0358

Observation e91f2dfc-d7f7-4e9b-8792-7f4fef976c8f · outbound

This paper cites RLHF Workflow: From Reward Modeling to Online RLHF.

Towards Reliable, Uncertainty-Aware Alignment RLHF Workflow: From Reward Modeling to Online RLHF

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:14.208038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:14.208038Z digest=sha256:7bf0bb8415d3d69bfd98d5f6432802755864f5be299a2df9f7319f723ffd9a59

Observation d2b52343-ad36-4d03-85ab-8d46ab45da4f · outbound

This paper cites Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking.

Towards Reliable, Uncertainty-Aware Alignment Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:14.278132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:14.278132Z digest=sha256:c96bf936afceda08961a3de7e92954927b727b2e590f50c00006fa47755c4569

Observation 0e5c7de4-926c-4d9f-8825-509c26dcdb16 · outbound

This paper cites Understanding dataset difficulty with v-usable information.

Towards Reliable, Uncertainty-Aware Alignment Understanding dataset difficulty with v-usable information

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:20.612471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-06T15:40:14.350372Z digest=sha256:7424c2aed06e7e16b3b9faba5b63793ef3caa2d87bf11bece7b59a943942f864

Observation cd467f85-11b8-4976-a8fe-6efab80bbead · outbound

This paper cites Dropout as a bayesian approximation: Representing model uncertainty in deep learning.

Towards Reliable, Uncertainty-Aware Alignment Dropout as a bayesian approximation: Representing model uncertainty in deep learning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:20.401191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-06T15:40:14.467970Z digest=sha256:b692939de2ac4d677a183adc17b229faea571b9ff079f5b85e586d72a2cd2ece

Observation 4115c2db-0a6e-4527-b18e-31c2da76f903 · outbound

This paper cites Scaling laws for reward model overoptimization.

Towards Reliable, Uncertainty-Aware Alignment Scaling laws for reward model overoptimization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:14.648398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:14.648398Z digest=sha256:c2f95d8aabe335cdc9095c5a7b973e13b36f3aa18b9c65fc0632ca6a47121bfe

Observation 8f8c4ae2-e1e8-4d32-b9a1-69738b2f53c3 · outbound

This paper cites https://huggingface.co/google/gemma-2b, 2024.

Towards Reliable, Uncertainty-Aware Alignment https://huggingface.co/google/gemma-2b, 2024

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:20.250009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-06T15:40:14.807037Z digest=sha256:89578679487fc6f85fb49496e5e99c119cbb3adaeec29597f6f5ccdca73c5c17

Observation 1dc032e9-a02f-4dce-b96d-6b068474f4da · outbound

This paper cites Beavertails: Towards improved safety alignment of llm via a human-preference dataset.

Towards Reliable, Uncertainty-Aware Alignment Beavertails: Towards improved safety alignment of llm via a human-preference dataset

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:14.973348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:14.973348Z digest=sha256:a7e129e99056ad5483bf53ac27f62fedd9ccd631b2a7e500ed9b3f53ef9dd8a8

Observation 0130163d-ee01-4d80-8e01-84e6c0938b5b · outbound

This paper cites Mistral 7B.

Towards Reliable, Uncertainty-Aware Alignment Mistral 7B

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:15.092213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:15.092213Z digest=sha256:e5473b64df1108e1ea9a6c300f4bbc9f9cbb3eb2386e91ec94bbbc02825dfcdb

Observation 33e0d3f5-5f7a-4d3d-a8d7-a3ec2544814f · outbound

This paper cites Reward Design with Language Models.

Towards Reliable, Uncertainty-Aware Alignment Reward Design with Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:15.257918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:15.257918Z digest=sha256:b2fc2d0ea642a1c3e56d517cef5e7223cd0d7daa86216c18ddf634f369429fd8

Observation 61587e0a-0251-4e40-b1ba-8cc8fee53523 · outbound

This paper cites RewardBench: Evaluating Reward Models for Language Modeling.

Towards Reliable, Uncertainty-Aware Alignment RewardBench: Evaluating Reward Models for Language Modeling

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:15.368188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:15.368188Z digest=sha256:5731499a412878fc6df8fa003a381d89d3ddb797becaa35bfd2b95a9474b69b5

Observation 24b201b5-75dd-4e28-9733-9bee5051bfbc · outbound

This paper cites Alpacaeval: An automatic evaluator of instruction-following models, 2023.

Towards Reliable, Uncertainty-Aware Alignment Alpacaeval: An automatic evaluator of instruction-following models, 2023

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:15.487856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:15.487856Z digest=sha256:b774cfc3085f5c2f6467508b671ffe3a2ea8f54c050ffb854c18b3249c9a0b74

Observation 6f0b80df-145c-4ef6-af21-d3a6aedfcba2 · outbound

This paper cites Openorca: An open dataset of gpt augmented flan reasoning traces, 2023.

Towards Reliable, Uncertainty-Aware Alignment Openorca: An open dataset of gpt augmented flan reasoning traces, 2023

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:20.024665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-06T15:40:15.623570Z digest=sha256:a060fe573128a05055942f6c5a6b0ea24e05c0631280b00b0be8b22ab2c42eb7

Observation 1544c899-73e4-4a7b-9a41-91ea72ae8d17 · outbound

This paper cites Reward Uncertainty for Exploration in Preference-based Reinforcement Learning.

Towards Reliable, Uncertainty-Aware Alignment Reward Uncertainty for Exploration in Preference-based Reinforcement Learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:15.689847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:15.689847Z digest=sha256:115356b6b86d0e058d7cecd93313723996d777d6de2c97bfda460302ad7f453b

Observation 0babacde-4ada-4f0a-adcc-94e46862e9e8 · outbound

This paper cites Iterative prompting for estimating epistemic uncertainty.

Towards Reliable, Uncertainty-Aware Alignment Iterative prompting for estimating epistemic uncertainty

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:19.849773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-06T15:40:15.762639Z digest=sha256:f225883dfa08c3f1ea16f33be2d724b696ec3b897aba46848bf33b4e2996c7d8

Observation 902799f6-fb92-4983-98a6-59f7a3fff674 · outbound

This paper cites Uncertainty-aware Reward Model: Teaching Reward Models to Know What is Unknown.

Towards Reliable, Uncertainty-Aware Alignment Uncertainty-aware Reward Model: Teaching Reward Models to Know What is Unknown

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:15.880022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:15.880022Z digest=sha256:aeaa27a9d1daa3eb48eb39ea7e320459aaee4d71de6ec421f38499747f4d6d3a

Observation 28c4d9ce-4dcc-4d52-85c6-fc6d793041e9 · outbound

This paper cites Introducing meta llama 3: The most capable openly available llm to date.

Towards Reliable, Uncertainty-Aware Alignment Introducing meta llama 3: The most capable openly available llm to date

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:15.957662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:15.957662Z digest=sha256:5b45c8212b20adabd58b1fb70a2099ff9e1ec1a1405b297079088cf03c35451b

Observation e439b574-48dd-419d-a3bb-64f9c38ba7af · outbound

This paper cites Montgomery and George C.

Towards Reliable, Uncertainty-Aware Alignment Montgomery and George C

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:19.638664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-06T15:40:16.020519Z digest=sha256:5399537b7baa9e79814377c72b160bbf980058e50d73436a856bfa3a3d3ee38d

Observation 757574e9-7c31-43a4-abd1-614f6e78f258 · outbound

This paper cites GPT-4 Technical Report.

Towards Reliable, Uncertainty-Aware Alignment GPT-4 Technical Report

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:16.098615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:16.098615Z digest=sha256:0a3e61f1389f313b1f70757743b0788f518e87ece8d837137343fa3b8fa41222

Observation 91246783-23e2-48a4-8de0-3e8e6fa149ca · outbound

This paper cites Training language models to follow instructions with human feedback.

Towards Reliable, Uncertainty-Aware Alignment Training language models to follow instructions with human feedback

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:16.193548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:16.193548Z digest=sha256:c288a31672e7723d07713f6c72f2cedd6512df9d6379aad6777c9470781f0b8c

Observation 18c522b9-122d-4a4a-adbd-92b17f02a131 · outbound

This paper cites Language models are unsupervised multitask learners.

Towards Reliable, Uncertainty-Aware Alignment Language models are unsupervised multitask learners

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:16.285137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:16.285137Z digest=sha256:e09eeee6593b0fa2f1e63828f8cd352b18ab04798b3a247297a4af15a9dedf7e

Observation 57c3ab16-89df-48eb-a27d-8adab523beee · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Towards Reliable, Uncertainty-Aware Alignment Direct preference optimization: Your language model is secretly a reward model

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:16.368788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:16.368788Z digest=sha256:eb48e8e050a2e1a8e68c2cd5c513397b7e6e9e4fff8e89c312a1405ad9090835

Observation 5650bcbd-ea5c-4a08-9b87-388f4fe98063 · outbound

This paper cites WARM: On the Benefits of Weight Averaged Reward Models.

Towards Reliable, Uncertainty-Aware Alignment WARM: On the Benefits of Weight Averaged Reward Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:16.470462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:16.470462Z digest=sha256:1c31aa0ed7e46c2ca5d09adc3a6dee63633eba05190b938f03f056ad1cf43ab1

Observation 9b210cb0-f49a-48d1-aea4-6c6df8c36850 · outbound

This paper cites Trust Region Policy Optimization.

Towards Reliable, Uncertainty-Aware Alignment Trust Region Policy Optimization

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:16.538614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:16.538614Z digest=sha256:26e8c40ad10a45f396d0e8809de14a3d6e9b20989bc4751dc13a2f4a7ca5a8e4

Observation 58d43223-87e6-472e-b055-18d7a48a7cdb · outbound

This paper cites Proximal Policy Optimization Algorithms.

Towards Reliable, Uncertainty-Aware Alignment Proximal Policy Optimization Algorithms

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:16.600627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:16.600627Z digest=sha256:9fc28cd86a6eb0acc12582fcda427d5963ea8018d6ce408a0cb92be70377a23b

Observation 98871c98-ad6d-461e-90d4-40083a555d8e · outbound

This paper cites Mutual fund performance.

Towards Reliable, Uncertainty-Aware Alignment Mutual fund performance

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:19.401125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-06T15:40:16.680687Z digest=sha256:f90bf2f9015f9ef42ada3e732c6cacc3a97fd1e1da3fd807612ab05b53c01efd

Observation a399d956-22d0-4734-b28b-5c7ad4a851cb · outbound

This paper cites an unresolved cited work.

Towards Reliable, Uncertainty-Aware Alignment Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:16.751394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:16.751394Z digest=sha256:3a0fdcbb638cd3599b04236cdaf426bb46a52096519d7b1823f388fcf3418415

Observation 9aeebd62-11a7-44fa-94b3-9d09552a9746 · outbound

This paper cites LLM-as-a-Judge & Reward Model: What They Can and Cannot Do.

Towards Reliable, Uncertainty-Aware Alignment LLM-as-a-Judge & Reward Model: What They Can and Cannot Do

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:16.846785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:16.846785Z digest=sha256:dda0f52c5a2b9675b3a68980eefbc75f7903f56ffc46dc11024e912d2db667fa

Observation cc34c0a9-eb6a-439b-851c-fc50b32b7883 · outbound

This paper cites Learning to summarize with human feedback.

Towards Reliable, Uncertainty-Aware Alignment Learning to summarize with human feedback

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:16.908542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:16.908542Z digest=sha256:e4c44fcf0175d0bbcbef45608f7984c2caca63d2738b5fa93739461981dcfc86

Observation 362f149e-6940-42e8-8316-2a92bb0a8b50 · outbound

This paper cites Policy gradient methods for reinforcement learning with function approximation.

Towards Reliable, Uncertainty-Aware Alignment Policy gradient methods for reinforcement learning with function approximation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:16.999910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:16.999910Z digest=sha256:705f8a0a8cf846dd7fb729ed64b2cfe455caa688918ad4da5fd0a7f45e4605bc

Observation b75730e9-da5d-469a-ad94-d67b210d359e · outbound

This paper cites Quantifying Uncertainty in Natural Language Explanations of Large Language Models.

Towards Reliable, Uncertainty-Aware Alignment Quantifying Uncertainty in Natural Language Explanations of Large Language Models

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:40:18.689752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-06T15:40:17.037690Z digest=sha256:5fe27aa6ac8d05aa7fc6acc0363a5145ec439ca142d081a065f61671613792f5

Observation 99d1eca3-0d57-44c2-89f2-f0e50ca4e355 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Towards Reliable, Uncertainty-Aware Alignment Gemini: A Family of Highly Capable Multimodal Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:17.125727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:17.125727Z digest=sha256:74b178b2f2f05ba7860a4636937b8225bf2ea4b780da8ca818442bd07abbc3b2

Observation 6b251672-a983-4cec-8bd6-84eb6017d14b · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Towards Reliable, Uncertainty-Aware Alignment Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:17.194495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:17.194495Z digest=sha256:54abc3e094a867548906e9b0168ba11f885f58280f392b7ae12bfb0f5d5b0cd8

Observation 8d6d2705-1af5-4aef-a24c-771fdf608aa3 · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

Towards Reliable, Uncertainty-Aware Alignment Gemma: Open Models Based on Gemini Research and Technology

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:17.315138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:17.315138Z digest=sha256:1d98b51783026b1504ae3a15bbadfe2bbf1267fe3ebe2a68a168d566acaa230c

Observation 4b7ec90c-cf1c-40a0-b378-fd8e214e9552 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Towards Reliable, Uncertainty-Aware Alignment Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:17.350383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:17.350383Z digest=sha256:05cb4aedc1c07f301b56b4d4c7327f08409e09e07b7c38c54bb0fdf150615f07

Observation 98e4ed2e-64d5-47b2-b840-901b19b509cf · outbound

This paper cites Trl: Transformer reinforcement learning.

Towards Reliable, Uncertainty-Aware Alignment Trl: Transformer reinforcement learning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:17.443886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:17.443886Z digest=sha256:78aa22661839c958125b09d1d36fc8bc0bbefa7910e80a57cc806033fa245514

Observation c8a35bd8-9eeb-4310-b713-fbe3f304dbc7 · outbound

This paper cites HelpSteer: Multi-attribute Helpfulness Dataset for SteerLM.

Towards Reliable, Uncertainty-Aware Alignment HelpSteer: Multi-attribute Helpfulness Dataset for SteerLM

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:17.507422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:17.507422Z digest=sha256:3d629212c4bea07017a07c5247f05102275a3762eea69c5337ddb53e38f686f0

Observation 91d61d80-9a94-43c0-bbbd-640ab06fd754 · outbound

This paper cites an unresolved cited work.

Towards Reliable, Uncertainty-Aware Alignment Unresolved cited work

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:17.562762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:17.562762Z digest=sha256:1670b436a713c9e4689fe94000a480a4116113b451e8cbb84c97ebcad4703afa

Observation d0847c30-0dc7-4bb2-b24e-876bd2a7af18 · outbound

This paper cites an unresolved cited work.

Towards Reliable, Uncertainty-Aware Alignment Unresolved cited work

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:17.663383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:17.663383Z digest=sha256:be6f423334165ea794ae86ea368749575c64fe295d2fa4ad1c604548bd72b657

Observation 86902715-0dcf-4e6c-ac7e-5e21f4300593 · outbound

This paper cites Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs.

Towards Reliable, Uncertainty-Aware Alignment Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:17.718260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:17.718260Z digest=sha256:6c895244fb619d183c95b0f34aa71c2df035fe849ba8702c0f5dc4e2537e8f76

Observation b9e8f8b0-26ef-4a1d-ab22-ede3d8ba02d1 · outbound

This paper cites Iterative preference learning from human feedback: Bridging theory and practice for rlhf under kl-constraint.

Towards Reliable, Uncertainty-Aware Alignment Iterative preference learning from human feedback: Bridging theory and practice for rlhf under kl-constraint

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:19.197653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-06T15:40:17.830385Z digest=sha256:42338a3b43ee9e3b9b73931f2486c6f134710b04cfed0e6f2a86a5bff2e2040d

Observation 9624b5bf-a1d1-4fa5-8d09-f4b118e216e4 · outbound

This paper cites Qwen2.5 Technical Report.

Towards Reliable, Uncertainty-Aware Alignment Qwen2.5 Technical Report

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:17.864241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:17.864241Z digest=sha256:7be285ed43931a3614da5857b7eddb5ac8500cd00f059bb22e78b55ef4a1efb1

Observation a0780e28-bddc-4838-9500-e3b1140c06a7 · outbound

This paper cites Advancing LLM Reasoning Generalists with Preference Trees.

Towards Reliable, Uncertainty-Aware Alignment Advancing LLM Reasoning Generalists with Preference Trees

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:17.967384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:17.967384Z digest=sha256:4f1167790875f2f2da9cf30a7cc27e6aaa856b382bfaf18fe3abad140f659062

Observation 07acfe67-c226-46b8-9be1-2d768943a270 · outbound

This paper cites Uncertainty-Penalized Reinforcement Learning from Human Feedback with Diverse Reward LoRA Ensembles.

Towards Reliable, Uncertainty-Aware Alignment Uncertainty-Penalized Reinforcement Learning from Human Feedback with Diverse Reward LoRA Ensembles

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:18.065331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:18.065331Z digest=sha256:ffce0b6bdaeed58f7d31b44ec3ba0893d5381e73d55913082c6ab99e5a82049d

Observation 1cc00e7e-3488-4566-9cd4-acba6e84c90a · outbound

This paper cites Bayesian prompt ensembles: Model uncertainty estimation for black-box large language models.

Towards Reliable, Uncertainty-Aware Alignment Bayesian prompt ensembles: Model uncertainty estimation for black-box large language models

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:18.999138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-06T15:40:18.152809Z digest=sha256:8ad4f0bec1d418423b2ca16cffba5af53e20f9397d7d21d1b62a31610bff06b5

Observation ac20c0ed-0aa3-4bb5-bf93-3e9d82819382 · outbound

This paper cites Improving Reinforcement Learning from Human Feedback with Efficient Reward Model Ensemble.

Towards Reliable, Uncertainty-Aware Alignment Improving Reinforcement Learning from Human Feedback with Efficient Reward Model Ensemble

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:18.178863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:18.178863Z digest=sha256:ac922a4e06b4f28c7e2ca24918af9974ec36faadcdce478607baaa6e142f6a20

Observation 5790b6ad-18f7-4d2b-96cb-4127ecc14bb7 · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.

Towards Reliable, Uncertainty-Aware Alignment Judging llm-as-a-judge with mt-bench and chatbot arena

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:18.225082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:18.225082Z digest=sha256:3aab8ffaa4da153173efb517ceaddd25692b8ee39120534e18f0897c7d46f229

Observation 6d9b15d1-25f5-428b-b2c2-6e41fc92c2a5 · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Towards Reliable, Uncertainty-Aware Alignment Fine-Tuning Language Models from Human Preferences

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:18.315140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:18.315140Z digest=sha256:dacf618d63c4e67f8a9a177e095a8f56d00eb21ca88ace798d2fae2daf0a1afe

Pith citing papers

No inbound Pith citation observations are available.