Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T19:28:59.743545Z
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 2 inbound Pith citation observations for arXiv:2507.05913.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T19:28:59.743545Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T12:27:28.637485Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-14T20:39:29.243714Z
64 of 64 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 8ce6e034-1836-4f15-b008-c71c9bb35f0e · outbound
Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d04662fd-b050-4620-bb54-e25119632517 · outbound
Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis GPT-4 Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ebd86b7-d44a-4208-906b-59d51d5aecf8 · outbound
Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Variational best-of-n alignment
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 601540f5-577a-44bd-969b-b44d2261e935 · outbound
Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Theoretical analysis of kl-regularized rlhf with multiple reference models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1af68a8-0e1c-43bf-a1fd-43ded7d1787d · outbound
Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Concrete Problems in AI Safety
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19bfc809-1ae5-4de0-bbfc-a668ae52244c · outbound
Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Infalign: Inference-aware language model alignment
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2472b580-9c99-416e-b729-29f25acd089f · outbound
Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Theoretical guarantees on the best-of-n alignment policy
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 161a9e20-845a-48ee-839d-021f3162647f · outbound
Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Q-learning for risk-sensitive control
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 88478503-8ef4-4b14-9a3e-694e37266884 · outbound
Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3bad5a91-22c5-419d-8c94-a3902c77982d · outbound
Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis A short note on an inequality between KL and TV
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5aa90d68-8dd2-4b55-8f6e-92ca8bc28ee2 · outbound
Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis The master equation and the convergence problem in mean field games:(ams-201)
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ac2747d2-67f7-42a0-9154-a4e87dbab96e · outbound
Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb73da34-49fc-4bb7-908b-50f43ba162e5 · outbound
Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Deep reinforcement learning from human preferences
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation daf9ad92-3588-4f28-b6ce-f0ab9bda33f9 · outbound
Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Soft best-of-n sampling for model alignment
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3500e86d-7259-4f5b-85ee-db9d1d77baf2 · outbound
Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Reward model ensembles help mitigate overoptimization
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8d3a0bcf-b5a8-421e-8c57-287a77f28537 · outbound
Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Safe RLHF: Safe Reinforcement Learning from Human Feedback
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f2a30f0-326f-48cf-a273-7622b5ece621 · outbound
Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Helping or herding? reward model ensembles mitigate but do not eliminate reward hacking
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3366544c-778c-497b-a6c4-f82a7b6c2df9 · outbound
Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 186afac0-7b57-4d07-a848-7ac07c1fadee · outbound
Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis A framework for few-shot language model evaluation
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6fdeb58c-e912-4005-876e-3cb8d2b7e924 · outbound
Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Scaling laws for reward model overoptimization
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d9d32290-4ce1-49f2-a2fe-9067a5ef89d8 · outbound
Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Guided Speculative Inference for Efficient Test-Time Alignment of LLMs
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab799296-c1a3-44d0-b8c4-81e837b32367 · outbound
Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis BoNBoN Alignment for Large Language Models and the Sweetness of Best-of-n Sampling
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 539fb220-9f5e-44d2-8be8-54fe107213b2 · outbound
Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Statistical theory of extreme values and some practical applications
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 64402ae6-45dc-4991-90ef-4421c7277eb4 · outbound
Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Hilton, P
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b1f6fedc-0060-4073-84ae-6bea47ec2d06 · outbound
Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Risk-sensitive markov decision processes
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c363c589-ea24-4b53-b71c-b72722679d1d · outbound
Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Is Best-of-N the Best of Them? Coverage, Scaling, and Optimality in Inference-Time Alignment
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71faf483-00ac-4258-909e-e4c8a15e3999 · outbound
Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Best-of-N Jailbreaking
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9389ae74-a7e5-4806-a549-42a4c80e1ee8 · outbound
Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Evaluation of best-of-n sampling strategies for language model alignment
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5e1823eb-b27b-49b4-aef6-df26a2505cc7 · outbound
Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Regularized best-of-n sampling to mitigate reward hacking for language model alignment
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6ed88f36-d430-445e-b00c-1c10a059d193 · outbound
Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Inference-time reward hacking in large language models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3223f1b-673c-43ab-ae62-693610b0990b · outbound
Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis ARGS: Alignment as Reward-Guided Search
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05b1f509-175e-4ba2-8111-9c46923220e3 · outbound
Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis On reinforcement learning and distribution matching for fine-tuning language models with no catastrophic forgetting
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 93333f7d-8616-41bb-b553-deae042a3f7d · outbound
Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis RL with KL penalties is better viewed as Bayesian inference
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4880f3a2-9ccf-4207-a706-833ace878a9d · outbound
Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis A new penalty function method for constrained minimization
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 310532ee-2c97-4b8b-af17-a07f09c7015f · outbound
Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Unveiling Safety Vulnerabilities of Large Language Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53e330cf-3069-4bea-8c0c-a36497ded7ec · outbound
Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis On tilted losses in machine learning: Theory and applications
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6dd95ac2-a8a0-4e63-bd5d-56f1633c3e6c · outbound
Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Deep Learning Theory Review: An Optimal Control and Dynamical Systems Perspective
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0642541b-daa0-41ce-8282-c91540f41148 · outbound
Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Information Theoretic Guarantees For Policy Alignment In Large Language Models
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2be4de91-739a-4700-bbfc-75b639fe3726 · outbound
Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Controlled decoding from language models
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 76b71c17-22b7-47fa-a71a-192e9cdd5cb0 · outbound
Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis WebGPT: Browser-assisted question-answering with human feedback
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b786f6be-7028-450b-a487-f725761866db · outbound
Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis 2 OLMo 2 Furious
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 727ea2c9-e0d9-49a4-a3b4-07f05e4e6421 · outbound
Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Training language models to follow instructions with human feedback
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5753a7e-95fc-4b31-ba63-0362f90f5ff4 · outbound
Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis On solving large-scale finite minimax problems using exponential smoothing
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4d635e82-6652-41e2-8114-782abf0743bf · outbound
Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Information theory: From coding to learning, 2022
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 380b06d1-6599-4930-a23b-5248a0d52bd6 · outbound
Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis TreeBoN: Enhancing Inference-Time Alignment with Speculative Tree-Search and Best-of-N Sampling
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1f2b118-520e-4cbd-9198-e42a0ac0638e · outbound
Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1082371d-7dc0-4884-9d76-215b731ec668 · outbound
Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis BOND: Aligning LLMs with Best-of-N Distillation
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da09c6cb-4cdc-4fc2-b991-3d8e071a0571 · outbound
Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3eff8743-70ea-4993-a1df-8e259b7fba33 · outbound
Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis The importance of online data: Understanding preference fine-tuning via coverage
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 43b2d2e0-e4b5-44d8-8694-8a8679a424a0 · outbound
Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Learning to summarize with human feedback
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0da765a-9292-43ce-8a78-0d46ac347b0e · outbound
Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Inference scaling f-laws: The limits of llm resampling with imperfect verifiers
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c54f448-de15-4244-b20a-a49a0b647a27 · outbound
Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Fast Best-of-N Decoding via Speculative Rejection
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a49e034-e132-4766-87c8-61fa95b9cdb6 · outbound
Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Gemini: A Family of Highly Capable Multimodal Models
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e3eca5e-52af-4aa6-970f-3d4ee506d061 · outbound
Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55bbc992-aee4-4daa-b6ee-a5cf6912e674 · outbound
Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Interpretable preferences via multi-objective reward modeling and mixture-of-experts
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation db710494-7c25-499f-9fda-597c279565c2 · outbound
Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Robust variable selection with exponential squared loss
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 12b2e296-184b-436e-b81b-7bece514809b · outbound
Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Iterative preference learning from human feedback: Bridging theory and practice for rlhf under kl-constraint
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c9134f7e-168b-4b3b-8458-985d2c47c02d · outbound
Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Asymptotics of language model alignment
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 093679e0-110d-4480-907f-5811acb13c66 · outbound
Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Convergence of the inexact langevin algorithm and score-based generative models in kl divergence
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 683f5707-3ded-4ea2-9359-4f792525dc08 · outbound
Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Online Iterative Reinforcement Learning from Human Feedback with General Preference Model
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation faa4c54d-4083-4620-95f4-cb0d669efcad · outbound
Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Provable Offline Preference-Based Reinforcement Learning
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c246fe7b-c98b-49c4-ab0f-333531e32621 · outbound
Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Sharp Analysis for KL-Regularized Contextual Bandits and RLHF
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d617809-b2c2-46f8-acdb-1da6b3ccd165 · outbound
Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Calibrating sequence likelihood improves conditional language generation
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c3eea33b-c332-4ce8-a9da-7bb7bec78699 · outbound
Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Judging llm-as-a-judge with mt-bench and chatbot arena
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3545a9e9-d152-4a23-b9c4-7580a44a94cb · inbound
Certifiable Safe RLHF: Semantic Grounding and Fixed Penalty Constraint Optimization for Safer LLM Alignment Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8584fc7a-97be-49fe-a67a-187801539ab3 · inbound
Temper and Tilt Lead to SLOP: Reward Hacking Mitigation with Inference-Time Alignment Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.