Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 87 inbound Pith citation observations for arXiv:1811.07871.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-12T18:28:47.669686Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
125
pith, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation b94fdaa7-3843-4324-8855-114572c29861 · inbound
Risks from Learned Optimization in Advanced Machine Learning Systems Scalable agent alignment via reward modeling: a research direction
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d1e9dfd4-2010-4583-89d1-ed849309bbce · inbound
Modeling AGI Safety Frameworks with Causal Influence Diagrams Scalable agent alignment via reward modeling: a research direction
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d203dae5-4607-4bfa-86b9-7b75b60e10cb · inbound
Towards Empathic Deep Q-Learning Scalable agent alignment via reward modeling: a research direction
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f26a7609-fa77-4a48-9253-ea8d790485c0 · inbound
Requisite Variety in Ethical Utility Functions for AI Value Alignment Scalable agent alignment via reward modeling: a research direction
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 52c0de86-d26c-4b11-9896-78bc6a622517 · inbound
Fine-Tuning Language Models from Human Preferences Scalable agent alignment via reward modeling: a research direction
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d87632bb-493b-4517-bc0f-7ec226ed6527 · inbound
Learning to summarize from human feedback Scalable agent alignment via reward modeling: a research direction
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation bb487044-62be-4736-a970-e6ac7f7712f4 · inbound
Scaling Laws for Reward Model Overoptimization Scalable agent alignment via reward modeling: a research direction
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation bc09f87c-8d8a-4b89-85ed-780447feea9e · inbound
Measuring Progress on Scalable Oversight for Large Language Models Scalable agent alignment via reward modeling: a research direction
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 838a7ba4-b93f-4f1f-9f8e-aa2f772f759b · inbound
Discovering Latent Knowledge in Language Models Without Supervision Scalable agent alignment via reward modeling: a research direction
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 376e6a77-db9e-484a-a804-6d0eca1b892b · inbound
RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment Scalable agent alignment via reward modeling: a research direction
Reference 114
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 321cb681-049f-486e-96dc-72cd871218ab · inbound
Universal and Transferable Adversarial Attacks on Aligned Language Models Scalable agent alignment via reward modeling: a research direction
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 086e11e0-497c-40dd-850b-e057234545be · inbound
A Roadmap to Pluralistic Alignment Scalable agent alignment via reward modeling: a research direction
Reference 271
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation cb019c6b-8076-4151-a9e9-5fb79ce59bcc · inbound
LLM Evaluators Recognize and Favor Their Own Generations Scalable agent alignment via reward modeling: a research direction
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 63a74004-5198-45c2-a7ab-c65270c68565 · inbound
Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models Scalable agent alignment via reward modeling: a research direction
Reference 278
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation eb7444a7-eb29-4c93-9394-6d45ee2f02de · inbound
Search, Verify and Feedback: Towards Next Generation Post-training Paradigm of Foundation Models via Verifier Engineering Scalable agent alignment via reward modeling: a research direction
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69fad9f2-d5b7-4d57-b5ce-393613f57313 · inbound
DLBacktrace: A Model Agnostic Explainability for any Deep Learning Models Scalable agent alignment via reward modeling: a research direction
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b76db881-6c47-4909-abd7-0cb3edd0fab4 · inbound
Effective Reward Specification in Deep Reinforcement Learning Scalable agent alignment via reward modeling: a research direction
Reference 188
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a28e9c60-0409-4a0e-ade6-96397527604a · inbound
The Superalignment of Superhuman Intelligence with Large Language Models Scalable agent alignment via reward modeling: a research direction
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 146ee1b6-60df-4fb6-a233-fbaab0d3e5fd · inbound
The Road to Artificial SuperIntelligence: A Comprehensive Survey of Superalignment Scalable agent alignment via reward modeling: a research direction
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2855cfac-44af-460d-9951-9d95e4373c8c · inbound
Aligning LLMs with Domain Invariant Reward Models Scalable agent alignment via reward modeling: a research direction
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18412e5e-0a41-4f6d-a3dc-2769650d39ec · inbound
Debate Helps Weak-to-Strong Generalization Scalable agent alignment via reward modeling: a research direction
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 067cdd38-c458-4409-b732-4577e4929598 · inbound
Leveraging Sparsity for Sample-Efficient Preference Learning: A Theoretical Perspective Scalable agent alignment via reward modeling: a research direction
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e076956e-6c58-435c-8b9f-f2c327833128 · inbound
Process Reinforcement through Implicit Rewards Scalable agent alignment via reward modeling: a research direction
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ab5c7713-fe20-4ef9-a9f2-ad7da4b506d9 · inbound
Learning from Active Human Involvement through Proxy Value Propagation Scalable agent alignment via reward modeling: a research direction
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f76cd35-3b8d-44d2-9919-f1446876d4a9 · inbound
Active Inference through Incentive Design in Markov Decision Processes Scalable agent alignment via reward modeling: a research direction
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8dccded-8c51-48f6-b699-731479c7493b · inbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization Scalable agent alignment via reward modeling: a research direction
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbb23fa1-8626-40e0-8e4b-0195ccf6c77d · inbound
Reinforcement Learning from Human Feedback Scalable agent alignment via reward modeling: a research direction
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2667f632-86dc-4ae1-b8b1-24eca8b177bb · inbound
Exploring Societal Concerns and Perceptions of AI: A Thematic Analysis through the Lens of Problem-Seeking Scalable agent alignment via reward modeling: a research direction
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45a0b585-35d0-4e79-b213-a2c76e17bc74 · inbound
Adversarial Attacks on Robotic Vision Language Action Models Scalable agent alignment via reward modeling: a research direction
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d2bb129-4ecb-4225-92d3-85b487ae0178 · inbound
Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Scalable agent alignment via reward modeling: a research direction
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 853a6945-caa1-4574-bb72-90775fe66610 · inbound
Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models Scalable agent alignment via reward modeling: a research direction
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4fc3d327-5031-45a9-bc38-86ab82cea280 · inbound
Towards Efficient and Effective Alignment of Large Language Models Scalable agent alignment via reward modeling: a research direction
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a75905cd-e4ab-40db-969b-80a1713f02c9 · inbound
Out of Control -- Why Alignment Needs Formal Control Theory (and an Alignment Control Stack) Scalable agent alignment via reward modeling: a research direction
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0891b52-ab44-40d8-b5bb-2d89b1cb5696 · inbound
Optimising Language Models for Downstream Tasks: A Post-Training Perspective Scalable agent alignment via reward modeling: a research direction
Reference 120
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca30404d-f858-4923-bbef-4c3b8c65c32b · inbound
Data Diversification Methods In Alignment Enhance Math Performance In LLMs Scalable agent alignment via reward modeling: a research direction
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de72642e-059b-4375-af75-3174116c1109 · inbound
5C Prompt Contracts: A Minimalist, Creative-Friendly, Token-Efficient Design Framework for Individual and SME LLM Usage Scalable agent alignment via reward modeling: a research direction
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00f4f874-3bec-43ef-a11b-88b37a366202 · inbound
One Token to Fool LLM-as-a-Judge Scalable agent alignment via reward modeling: a research direction
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f13d4424-1aed-4394-9070-21d3e2c96fd9 · inbound
The AI Ethical Resonance Hypothesis: The Possibility of Discovering Moral Meta-Patterns in AI Systems Scalable agent alignment via reward modeling: a research direction
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6893840f-8913-456c-8bfe-03992331d013 · inbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Scalable agent alignment via reward modeling: a research direction
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 968bd471-84e3-41b6-900d-994e907fdce0 · inbound
Learning the Value Systems of Societies from Preferences Scalable agent alignment via reward modeling: a research direction
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ea747f1-0089-4256-a17b-1ce8ffe9fec6 · inbound
Sample-efficient LLM Optimization with Reset Replay Scalable agent alignment via reward modeling: a research direction
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 71bf551d-680e-43a1-a83b-17835cde46b4 · inbound
SSRL: Self-Search Reinforcement Learning Scalable agent alignment via reward modeling: a research direction
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5251e13c-1174-463d-a131-8903d4b20d22 · inbound
Governance-as-a-Service: A Multi-Agent Framework for AI System Compliance and Policy Enforcement Scalable agent alignment via reward modeling: a research direction
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a40b2d6-0391-4062-8c0f-1bc4c973d466 · inbound
Democracy-in-Silico: Institutional Design as Alignment in AI-Governed Polities Scalable agent alignment via reward modeling: a research direction
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1f490a5-886d-44a7-ae8a-6e568a64fc81 · inbound
Safety Alignment Should Be Made More Than Just A Few Attention Heads Scalable agent alignment via reward modeling: a research direction
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6500af1-ac20-43ba-9e3a-d7c9f5d72e2d · inbound
What Fundamental Structure in Reward Functions Enables Efficient Sparse-Reward Learning? Scalable agent alignment via reward modeling: a research direction
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57594bd1-0c0c-4f0a-9891-1937d44e3e1c · inbound
Contrastive Weak-to-strong Generalization Scalable agent alignment via reward modeling: a research direction
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a71921f-3217-4d3b-ba42-5238a305e796 · inbound
Human-AI Complementarity: A Goal for Amplified Oversight Scalable agent alignment via reward modeling: a research direction
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6814fecb-f832-43ca-a1e4-42b35894e8e7 · inbound
An Onto-Relational-Sophic Framework for Governing Synthetic Minds Scalable agent alignment via reward modeling: a research direction
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6f694e5f-4080-4fd4-a56e-e8d684ac58a4 · inbound
Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems Scalable agent alignment via reward modeling: a research direction
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 346dd22f-f42d-4ff7-83ba-9109cb697879 · inbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Scalable agent alignment via reward modeling: a research direction
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 290b8545-bbe7-4356-8202-21e39b3928a5 · inbound
The Alignment Target Problem: Divergent Moral Judgments of Humans, AI Systems, and Their Designers Scalable agent alignment via reward modeling: a research direction
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7eb3076e-cc94-4147-bf1f-713ed36275b1 · inbound
The Alignment Target Problem: Divergent Moral Judgments of Humans, AI Systems, and Their Designers Scalable agent alignment via reward modeling: a research direction
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8f5f280a-5801-49d8-b7b9-e24a601ce69b · inbound
The Alignment Target Problem: Divergent Moral Judgments of Humans, AI Systems, and Their Designers Scalable agent alignment via reward modeling: a research direction
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ecec322c-54ad-4663-8969-26e9f2fa5233 · inbound
The Alignment Target Problem: Divergent Moral Judgments of Humans, AI Systems, and Their Designers Scalable agent alignment via reward modeling: a research direction
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba094ca4-f60c-4c44-9d2b-3237657e92f1 · inbound
Uncertainty-Aware Reward Discounting for Mitigating Reward Hacking Scalable agent alignment via reward modeling: a research direction
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 15c88cbe-4c01-4dec-86a6-9f0b373e65f4 · inbound
Uncertainty-Aware Reward Discounting for Mitigating Reward Hacking Scalable agent alignment via reward modeling: a research direction
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation bd65d651-7eeb-464e-a2e7-6eecf7809936 · inbound
AI Safety as Control of Irreversibility: A Systems Framework for Decision-Energy and Sovereignty Boundaries Scalable agent alignment via reward modeling: a research direction
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 91f3e312-3992-454c-9b1d-5a4be522007a · inbound
AI Alignment via Incentives and Correction Scalable agent alignment via reward modeling: a research direction
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation dd840523-7196-4847-812e-c7d2c448744b · inbound
AI Alignment via Incentives and Correction Scalable agent alignment via reward modeling: a research direction
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 11287020-73ed-4169-9207-669247e1193f · inbound
Brainrot: Deskilling and Addiction are Overlooked AI Risks Scalable agent alignment via reward modeling: a research direction
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ee9e7ea6-dfa6-4125-bf1c-c3544f0324dc · inbound
Automated alignment is harder than you think Scalable agent alignment via reward modeling: a research direction
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 0a830ea8-b730-42b0-b3fd-fec8c5087a4d · inbound
Correcting Influence: Unboxing LLM Outputs with Orthogonal Latent Spaces Scalable agent alignment via reward modeling: a research direction
Reference 222
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 302cc156-bda3-4183-ae8a-9a7e927eb4c5 · inbound
Silent Collapse in Recursive Learning Systems Scalable agent alignment via reward modeling: a research direction
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 1ef68a19-2b98-47dc-ab8a-36b11d113ca7 · inbound
Deep Pre-Alignment for VLMs Scalable agent alignment via reward modeling: a research direction
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b979e1aa-d578-4595-8afc-95e9cbceae06 · inbound
When In-Distribution Gains Fail: Evaluating Weak-to-Strong Reward Models under Preference Shift Scalable agent alignment via reward modeling: a research direction
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2858b5f3-6215-4f10-9aa0-61270355617d · inbound
The Behavioral Credibility Trilemma: When Calibrated Autonomy Becomes Impossible Scalable agent alignment via reward modeling: a research direction
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2096c4f9-dca9-4785-8631-9ae77de9396e · inbound
The Behavioral Credibility Trilemma: When Calibrated Autonomy Becomes Impossible Scalable agent alignment via reward modeling: a research direction
Reference 2007
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2e71a42-012c-47ec-892c-040bd312a826 · inbound
When Helping Hurts and How to Fix It: Multi-Agent Debate for Data Cleaning Scalable agent alignment via reward modeling: a research direction
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6d0bac49-ca62-46cc-a8b3-0edc1f23e5ce · inbound
Self-Commitment Latency: A Reward-Free Probe for Prompted Implicit Hacking Scalable agent alignment via reward modeling: a research direction
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 0a850ee1-20d6-4505-89d1-7c224499ee90 · inbound
Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization Scalable agent alignment via reward modeling: a research direction
Reference 243
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 604eccae-230a-4312-8723-db246ae5877e · inbound
AI Researchers Must Help Lead Arms Control to Mitigate Military AI Risks Scalable agent alignment via reward modeling: a research direction
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 34fd7f76-8c20-48b0-a40f-eac988dbbaff · inbound
The Tao of Agency: Autotelic AI, Embedded Agency and Dissolution of the Self Scalable agent alignment via reward modeling: a research direction
Reference 112
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 45b653c3-268d-494b-967d-96ed487c5700 · inbound
Objective-Behavior Alignment: Diagnostics for MORL Policy Selection Scalable agent alignment via reward modeling: a research direction
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8bd9b28e-0363-4cd5-b607-d46f485e04d0 · inbound
AI Alignment From Social Choice Perspectives Scalable agent alignment via reward modeling: a research direction
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2deda88b-774f-4af2-985e-a446b8904f59 · inbound
The Unverifiability of Artificial General Intelligence (AGI) Alignment, Static and Dynamic: From Trakhtenbrot's Wall to the Safety-Generality Tension Scalable agent alignment via reward modeling: a research direction
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f76ab2fe-e0e2-4b9f-873b-193d816330ba · inbound
Constructive Alignment: Governing Preference Dynamics in Human-AI Interaction Scalable agent alignment via reward modeling: a research direction
Reference 138
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d133d6e5-7d41-4d4a-a333-c5e7d559ee95 · inbound
MentalThink: Shaping Thoughts in Mental SVG World Scalable agent alignment via reward modeling: a research direction
Reference 170
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fcaf0638-9baf-4b6f-ba54-6d5a5ed255e1 · inbound
How to Avoid Debate: Scalable AI Safety via Doubly-Efficient Interactive Proofs Scalable agent alignment via reward modeling: a research direction
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 514601cb-6a7d-4996-94de-f388951ae611 · inbound
Attention Limited Reward Learning Scalable agent alignment via reward modeling: a research direction
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0ef56ba-a906-45f4-8759-0fcb74e602fd · inbound
Weak-to-Strong Generalization via Direct On-Policy Distillation Scalable agent alignment via reward modeling: a research direction
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 0fbe11a8-925b-4d49-8365-804d4389dcb3 · inbound
Weak-to-Strong Generalization via Direct On-Policy Distillation Scalable agent alignment via reward modeling: a research direction
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee0177bf-6ff1-4f8b-9caa-d4173551a443 · inbound
Learning Safe Agent Behaviour from Human Preferences and Justifications via World Models Scalable agent alignment via reward modeling: a research direction
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5349a75b-0367-4be7-99c1-a8f420d5ca75 · inbound
Relative Value Learning Scalable agent alignment via reward modeling: a research direction
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8aba963f-f187-4d10-82b5-07710c08c0d8 · inbound
When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops Scalable agent alignment via reward modeling: a research direction
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4a00365-adb6-467d-b5f8-9001742877bb · inbound
Calibrating Artificial Guilt: Neurally Grounded Reward Shaping for Prosocial Multi-Agent Reinforcement Learning Scalable agent alignment via reward modeling: a research direction
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3cc8b0a5-815d-48cf-9357-d58ee0e680b0 · inbound
Why Study Emergent Behavior When You Can Regulate It? Aligning Multi-Agent Systems with Reward Prediction Scalable agent alignment via reward modeling: a research direction
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.