Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T15:40:18.315140Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 0 inbound Pith citation observations for arXiv:2507.15906.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T15:40:18.315140Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
59 of 59 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation d809cf6f-ee7c-4ffd-892d-e7588102b78c · outbound
Towards Reliable, Uncertainty-Aware Alignment The claude 3 model family: Opus, sonnet, haiku
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 47697553-f704-4bb3-82e3-bfc825ef34c7 · outbound
Towards Reliable, Uncertainty-Aware Alignment The Llama 3 Herd of Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee77b394-d87e-49f6-8280-3c14d712c9cb · outbound
Towards Reliable, Uncertainty-Aware Alignment Concrete Problems in AI Safety
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fbb27280-7859-4c32-8d4f-a42d30e3c2f0 · outbound
Towards Reliable, Uncertainty-Aware Alignment Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48ae5dfa-eb9f-465e-aa15-b6cc4542ab59 · outbound
Towards Reliable, Uncertainty-Aware Alignment Constitutional AI: Harmlessness from AI Feedback
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1728c136-ef99-4e4d-a6ff-0be5e299c464 · outbound
Towards Reliable, Uncertainty-Aware Alignment Deep reinforcement learning from human preferences
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d81661bf-1992-4d60-91f9-a40889259d1e · outbound
Towards Reliable, Uncertainty-Aware Alignment Reward Model Ensembles Help Mitigate Overoptimization
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08da191c-c502-4312-bed6-4226835d001c · outbound
Towards Reliable, Uncertainty-Aware Alignment UltraFeedback: Boosting Language Models with Scaled AI Feedback
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d279d30-0dcf-4f19-bfaa-89c952a0a89f · outbound
Towards Reliable, Uncertainty-Aware Alignment Suphavadeeprasit
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5e10c0f0-fb17-4bfe-9a8f-44419e76630c · outbound
Towards Reliable, Uncertainty-Aware Alignment Gemini 2.0 flash: Next-generation multimodal ai model
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 1361e871-f6d0-4c4d-9c6c-b5e0025049e2 · outbound
Towards Reliable, Uncertainty-Aware Alignment DeepSeek-V3 Technical Report
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97e0d881-e718-487f-a5bb-1e108dcacce4 · outbound
Towards Reliable, Uncertainty-Aware Alignment RAFT : Reward ranked finetuning for generative foundation model alignment
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e91f2dfc-d7f7-4e9b-8792-7f4fef976c8f · outbound
Towards Reliable, Uncertainty-Aware Alignment RLHF Workflow: From Reward Modeling to Online RLHF
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2b52343-ad36-4d03-85ab-8d46ab45da4f · outbound
Towards Reliable, Uncertainty-Aware Alignment Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e5c7de4-926c-4d9f-8825-509c26dcdb16 · outbound
Towards Reliable, Uncertainty-Aware Alignment Understanding dataset difficulty with v-usable information
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation cd467f85-11b8-4976-a8fe-6efab80bbead · outbound
Towards Reliable, Uncertainty-Aware Alignment Dropout as a bayesian approximation: Representing model uncertainty in deep learning
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4115c2db-0a6e-4527-b18e-31c2da76f903 · outbound
Towards Reliable, Uncertainty-Aware Alignment Scaling laws for reward model overoptimization
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f8c4ae2-e1e8-4d32-b9a1-69738b2f53c3 · outbound
Towards Reliable, Uncertainty-Aware Alignment https://huggingface.co/google/gemma-2b, 2024
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 1dc032e9-a02f-4dce-b96d-6b068474f4da · outbound
Towards Reliable, Uncertainty-Aware Alignment Beavertails: Towards improved safety alignment of llm via a human-preference dataset
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0130163d-ee01-4d80-8e01-84e6c0938b5b · outbound
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33e0d3f5-5f7a-4d3d-a8d7-a3ec2544814f · outbound
Towards Reliable, Uncertainty-Aware Alignment Reward Design with Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61587e0a-0251-4e40-b1ba-8cc8fee53523 · outbound
Towards Reliable, Uncertainty-Aware Alignment RewardBench: Evaluating Reward Models for Language Modeling
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24b201b5-75dd-4e28-9733-9bee5051bfbc · outbound
Towards Reliable, Uncertainty-Aware Alignment Alpacaeval: An automatic evaluator of instruction-following models, 2023
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f0b80df-145c-4ef6-af21-d3a6aedfcba2 · outbound
Towards Reliable, Uncertainty-Aware Alignment Openorca: An open dataset of gpt augmented flan reasoning traces, 2023
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 1544c899-73e4-4a7b-9a41-91ea72ae8d17 · outbound
Towards Reliable, Uncertainty-Aware Alignment Reward Uncertainty for Exploration in Preference-based Reinforcement Learning
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0babacde-4ada-4f0a-adcc-94e46862e9e8 · outbound
Towards Reliable, Uncertainty-Aware Alignment Iterative prompting for estimating epistemic uncertainty
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 902799f6-fb92-4983-98a6-59f7a3fff674 · outbound
Towards Reliable, Uncertainty-Aware Alignment Uncertainty-aware Reward Model: Teaching Reward Models to Know What is Unknown
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28c4d9ce-4dcc-4d52-85c6-fc6d793041e9 · outbound
Towards Reliable, Uncertainty-Aware Alignment Introducing meta llama 3: The most capable openly available llm to date
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e439b574-48dd-419d-a3bb-64f9c38ba7af · outbound
Towards Reliable, Uncertainty-Aware Alignment Montgomery and George C
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 757574e9-7c31-43a4-abd1-614f6e78f258 · outbound
Towards Reliable, Uncertainty-Aware Alignment GPT-4 Technical Report
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91246783-23e2-48a4-8de0-3e8e6fa149ca · outbound
Towards Reliable, Uncertainty-Aware Alignment Training language models to follow instructions with human feedback
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18c522b9-122d-4a4a-adbd-92b17f02a131 · outbound
Towards Reliable, Uncertainty-Aware Alignment Language models are unsupervised multitask learners
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57c3ab16-89df-48eb-a27d-8adab523beee · outbound
Towards Reliable, Uncertainty-Aware Alignment Direct preference optimization: Your language model is secretly a reward model
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5650bcbd-ea5c-4a08-9b87-388f4fe98063 · outbound
Towards Reliable, Uncertainty-Aware Alignment WARM: On the Benefits of Weight Averaged Reward Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b210cb0-f49a-48d1-aea4-6c6df8c36850 · outbound
Towards Reliable, Uncertainty-Aware Alignment Trust Region Policy Optimization
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58d43223-87e6-472e-b055-18d7a48a7cdb · outbound
Towards Reliable, Uncertainty-Aware Alignment Proximal Policy Optimization Algorithms
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98871c98-ad6d-461e-90d4-40083a555d8e · outbound
Towards Reliable, Uncertainty-Aware Alignment Mutual fund performance
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a399d956-22d0-4734-b28b-5c7ad4a851cb · outbound
Towards Reliable, Uncertainty-Aware Alignment Unresolved cited work
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9aeebd62-11a7-44fa-94b3-9d09552a9746 · outbound
Towards Reliable, Uncertainty-Aware Alignment LLM-as-a-Judge & Reward Model: What They Can and Cannot Do
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc34c0a9-eb6a-439b-851c-fc50b32b7883 · outbound
Towards Reliable, Uncertainty-Aware Alignment Learning to summarize with human feedback
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 362f149e-6940-42e8-8316-2a92bb0a8b50 · outbound
Towards Reliable, Uncertainty-Aware Alignment Policy gradient methods for reinforcement learning with function approximation
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b75730e9-da5d-469a-ad94-d67b210d359e · outbound
Towards Reliable, Uncertainty-Aware Alignment Quantifying Uncertainty in Natural Language Explanations of Large Language Models
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 99d1eca3-0d57-44c2-89f2-f0e50ca4e355 · outbound
Towards Reliable, Uncertainty-Aware Alignment Gemini: A Family of Highly Capable Multimodal Models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b251672-a983-4cec-8bd6-84eb6017d14b · outbound
Towards Reliable, Uncertainty-Aware Alignment Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d6d2705-1af5-4aef-a24c-771fdf608aa3 · outbound
Towards Reliable, Uncertainty-Aware Alignment Gemma: Open Models Based on Gemini Research and Technology
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b7ec90c-cf1c-40a0-b378-fd8e214e9552 · outbound
Towards Reliable, Uncertainty-Aware Alignment Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98e4ed2e-64d5-47b2-b840-901b19b509cf · outbound
Towards Reliable, Uncertainty-Aware Alignment Trl: Transformer reinforcement learning
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8a35bd8-9eeb-4310-b713-fbe3f304dbc7 · outbound
Towards Reliable, Uncertainty-Aware Alignment HelpSteer: Multi-attribute Helpfulness Dataset for SteerLM
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91d61d80-9a94-43c0-bbbd-640ab06fd754 · outbound
Towards Reliable, Uncertainty-Aware Alignment Unresolved cited work
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0847c30-0dc7-4bb2-b24e-876bd2a7af18 · outbound
Towards Reliable, Uncertainty-Aware Alignment Unresolved cited work
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86902715-0dcf-4e6c-ac7e-5e21f4300593 · outbound
Towards Reliable, Uncertainty-Aware Alignment Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9e8f8b0-26ef-4a1d-ab22-ede3d8ba02d1 · outbound
Towards Reliable, Uncertainty-Aware Alignment Iterative preference learning from human feedback: Bridging theory and practice for rlhf under kl-constraint
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9624b5bf-a1d1-4fa5-8d09-f4b118e216e4 · outbound
Towards Reliable, Uncertainty-Aware Alignment Qwen2.5 Technical Report
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0780e28-bddc-4838-9500-e3b1140c06a7 · outbound
Towards Reliable, Uncertainty-Aware Alignment Advancing LLM Reasoning Generalists with Preference Trees
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07acfe67-c226-46b8-9be1-2d768943a270 · outbound
Towards Reliable, Uncertainty-Aware Alignment Uncertainty-Penalized Reinforcement Learning from Human Feedback with Diverse Reward LoRA Ensembles
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1cc00e7e-3488-4566-9cd4-acba6e84c90a · outbound
Towards Reliable, Uncertainty-Aware Alignment Bayesian prompt ensembles: Model uncertainty estimation for black-box large language models
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ac20c0ed-0aa3-4bb5-bf93-3e9d82819382 · outbound
Towards Reliable, Uncertainty-Aware Alignment Improving Reinforcement Learning from Human Feedback with Efficient Reward Model Ensemble
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5790b6ad-18f7-4d2b-96cb-4127ecc14bb7 · outbound
Towards Reliable, Uncertainty-Aware Alignment Judging llm-as-a-judge with mt-bench and chatbot arena
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d9b15d1-25f5-428b-b2c2-6e41fc92c2a5 · outbound
Towards Reliable, Uncertainty-Aware Alignment Fine-Tuning Language Models from Human Preferences
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.