Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T11:17:36.764467Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 0 inbound Pith citation observations for arXiv:2510.06096.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T11:17:36.764467Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
16 of 16 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 1a4c187e-5c6c-4a98-9c82-14f8146df2af · outbound
The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives Uncertainty quantification in fine-tuned LLMs using LoRA ensembles
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2174ce8-2c66-4898-a84a-1f6dcd1579e8 · outbound
The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e0b2168-f3dc-4494-9489-80b906e6c9ba · outbound
The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives Toy Models of Superposition
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9c2b000-f501-40e1-85e4-21a600ba3d39 · outbound
The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives Alignment of Language Agents
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8c570dd-d43a-4150-b0d1-544d1c28b749 · outbound
The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives Simple Yet Effective: An Information-Theoretic Approach to Multi-LLM Uncertainty Quantification
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7206c980-9ad0-4ebd-b63c-4191a50e06af · outbound
The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives Tulu 3: Pushing Frontiers in Open Language Model Post-Training
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8960377-10e3-4407-a63f-80af2ab5a6f5 · outbound
The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives Bayesian prompt ensembles: Model uncertainty estimation for black-box large language models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4440dea5-165d-4523-8693-b37bfd7be1bc · outbound
The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives Bayesian Reward Models for LLM Alignment
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6feb97e1-b97a-4e3f-a7d9-b00823355b0e · outbound
The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives toget back at fuckboys
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86eb37f9-d4fa-4adb-8efd-97e1592fb523 · outbound
The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives Textual Bayes: Quantifying Prompt Uncertainty in LLM-Based Systems
Reference 2007
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69cb4738-ee36-4eff-815f-889f1f09edca · outbound
The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58ce233a-7689-4a88-977a-8ab27b2358d0 · outbound
The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e9a5758-72b3-4456-9766-bff79fea411a · outbound
The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives Emergent misalignment: Narrow finetuning can produce broadly misaligned llms.arXiv preprint arXiv:2502.17424,
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7416fa0-5e01-4096-8f7b-9953471000e4 · outbound
The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd63256c-b9d3-431d-8e54-c3cdc8957097 · outbound
The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives Ethical and social risks of harm from Language Models
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce3d0d09-b88d-4b46-a180-347335720062 · outbound
The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives On the Opportunities and Risks of Foundation Models
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.