Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T23:55:19.365804Z
Paper Citation Record · LEDGER
As of 15 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 8 inbound Pith citation observations for arXiv:2506.15651.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T23:55:19.365804Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T16:07:42.225595Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T08:09:40.705912Z
28 of 28 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 5d924e4a-6e1c-4a66-a136-cfe0b695ed8a · outbound
AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning The claude 3 model family: Opus, sonnet, haiku, 2024
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 6c02a401-4af8-4245-ac7f-1578f504b921 · outbound
AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning Training a helpful and harmless assistant with reinforcement learning from human feedback,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation d409543e-9789-4a80-a823-91960e13630e · outbound
AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning Constitutional AI: Harmlessness from AI Feedback
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf3ebc87-f99e-47ec-bca2-45f2ea52e711 · outbound
AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning Odin: disentangled reward mitigates hacking in rlhf
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 8e933ef1-de09-41c8-b638-2b2aa1613da6 · outbound
AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning Ultrafeedback: Boosting language models with high-quality feedback, 2024
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation dc5c5bfc-8105-49c2-a982-d9c7d62efdea · outbound
AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c38fd67-dc42-46cd-b62f-8d1ad1879e81 · outbound
AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning Length-controlled alpacaeval: A simple debiasing of automatic evaluators
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 1632e14a-4c0e-4f8b-a419-19fc043d74dc · outbound
AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning Reward shaping to mitigate reward hacking in rlhf, 2025
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a17db659-2111-4ca2-a188-08f38718066a · outbound
AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning Scaling laws for reward model overoptimization
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 3b3496ff-6536-4ee3-8b78-27e842cb7a96 · outbound
AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning S., Green, R., Mokrá, S., Fernando, N., Wu, B., Foley, R., Young, S., Gabriel, I., Isaac, W., Mellor, J., Hassabis, D., Kavukcuoglu, K., Hendricks, L
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 33486952-d28f-4202-89ff-01085059ccad · outbound
AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning Gemini: A Family of Highly Capable Multimodal Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e46063c1-7856-4808-89d6-2be2185331d0 · outbound
AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning Openrlhf: An easy-to-use, scalable and high-performance rlhf framework, 2024
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 2c941af4-2b1b-46cc-81cb-7852c8327d7c · outbound
AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning The Llama 3 Herd of Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 168800b7-45b3-482e-883a-7fb8975cd15d · outbound
AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning Inform: Mitigating reward hacking in rlhf via information-theoretic reward modeling
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 96037b73-3393-4f8f-9b99-617ffb2e20f2 · outbound
AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning Rule based rewards for language model safety
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 3be88102-f3ab-457f-a736-06037b856f0e · outbound
AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning GPT-4 Technical Report
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6baabdaf-39d8-49c0-9a0a-198d02f2d057 · outbound
AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning F., Leike, J., and Lowe, R
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation f4682ebe-c9bd-4415-aa2b-018b7cde02f4 · outbound
AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning D., Ermon, S., and Finn, C
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation aef2970b-9987-4889-a699-3eeb389b82eb · outbound
AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning Warm: on the benefits of weight averaged reward models
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 91244e0a-5e96-480d-92cc-fd7a69542366 · outbound
AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning Proximal Policy Optimization Algorithms
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a06f961-8d08-4459-8a29-ff3f0708672a · outbound
AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38a3a96d-3812-48de-9d11-7143f4d24650 · outbound
AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning A long way to go: Investigating length correlations in RLHF, 2024
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation f1bce08e-2765-4650-ad99-e26236de101d · outbound
AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning Learning to summarize from human feedback
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe93116b-d111-474a-a295-271f3d6f21d5 · outbound
AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning Interpretable preferences via multi- objective reward modeling and mixture-of-experts
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f3038dd-c30c-4500-b988-cc7392c35913 · outbound
AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning Transforming and combining rewards for aligning large language models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 2bfa2f02-2fad-454c-b3c1-0db4f0351f45 · outbound
AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning E., and Stoica, I
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 3e40a4df-025b-4e51-bfef-1be957b8cf7f · outbound
AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29deb25a-cbbc-4ee4-ba28-98eb74ad27e1 · outbound
AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning confidence
Reference 2023
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation eb9dc29b-af21-4a33-837f-297378dc94bc · inbound
A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning
Reference 262
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation ada27720-584f-40ab-a0a2-a0c9f39baa49 · inbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning
Reference 170
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ca00ec7-4d2a-4e55-9fcd-e12388adae42 · inbound
Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation bfb0290e-40fb-4bfc-a48a-58a088745556 · inbound
Evaluation-driven Scaling for Scientific Discovery AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning
Reference 151
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation abdb9b9b-5c93-48c9-94f5-4a5ab4fc5248 · inbound
AutoRubric-T2I: Robust Rule-Based Reward Model for Text-to-Image Alignment AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation cee7903f-9cbe-4bc3-b9c4-dbedb0d01671 · inbound
AutoRubric-T2I: Robust Rule-Based Reward Model for Text-to-Image Alignment AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation e2dc734c-b2b6-479d-9a4c-e9bd28f1bad1 · inbound
Generating and Refining Dynamic Evaluation Rubrics for LLM-as-a-Judge AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 2107dac5-c1d6-48f2-8801-6c2dd0e50e20 · inbound
Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning
Reference 219
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.