Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 100 inbound Pith citation observations for arXiv:2307.15217.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-11T22:06:54.024699Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
92
pith, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation d8a95642-eca0-4579-8bb7-7b6ef46e5aa6 · inbound
RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 126
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a07e6ed0-a364-4ab9-97be-515c3ad34b6b · inbound
Active teacher selection for reward learning Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 828da731-537b-4454-aea0-1201131bcf23 · inbound
A Roadmap to Pluralistic Alignment Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 154
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 25cf5175-e1be-4c62-b356-a41a5f68d7b9 · inbound
LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 140
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 31597a6d-0474-4ec5-96f5-788952ad2efb · inbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 264
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3c04b55c-8786-430e-bb16-57578f39a27c · inbound
AI Safety Landscape for Large Language Models: Taxonomy, State-of-the-art, and Future Directions Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 102
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b586926b-b09e-49cc-83be-fd9e878bacb0 · inbound
Training Language Models to Self-Correct via Reinforcement Learning Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 141
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d1131658-53a2-4c1b-8da7-02f27aafd9fb · inbound
Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 00bf9566-314e-43d9-b38e-03e081c3ab0b · inbound
Towards Data Governance of Frontier AI Models Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a718bade-719b-4609-8a77-e6cf523f3bdb · inbound
ReFF: Reinforcing Format Faithfulness in Language Models across Varied Tasks Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98458dd8-32be-44da-89a8-380a9eab4476 · inbound
Active Inference and Human--Computer Interaction Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97938775-c86a-4530-a403-76c05dd628f5 · inbound
MPPO: Multi Pair-wise Preference Optimization for LLMs with Arbitrary Negative Samples Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 453f6f5d-cbc5-4d96-a03a-fd10d120fdb0 · inbound
Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71e97d38-5174-4b94-ba1f-49e4796217cf · inbound
AI Agent for Education: von Neumann Multi-Agent System Framework Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e75027bd-2ea8-4057-830a-9d5dc3d72fa9 · inbound
Many of Your DPOs are Secretly One: Attempting Unification Through Mutual Information Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c945e325-1d29-4ab8-b70e-07f7936b83cd · inbound
AlphaPO: Reward Shape Matters for LLM Alignment Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e054cc71-f857-4a29-84a1-efc37b332d2f · inbound
Stream Aligner: Efficient Sentence-Level Alignment via Distribution Induction Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 769e38ec-1831-4bfe-a1e9-89e1706b2c4f · inbound
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2973a96-1e78-4e54-8240-458dcbcedce9 · inbound
Debate Helps Weak-to-Strong Generalization Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8f884b5-31a6-40bb-89b3-0c29f5c8b440 · inbound
Trustworthiness in Stochastic Systems: Towards Opening the Black Box Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db8e2882-4d5e-4719-b39d-617a0ab8bac5 · inbound
Offline Learning for Combinatorial Multi-armed Bandits Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1bf8590-3090-406b-a708-657b7bd30e25 · inbound
The Energy Loss Phenomenon in RLHF: A New Perspective on Mitigating Reward Hacking Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81b336dd-1a86-4be9-b198-ad2b53e51433 · inbound
Process-Supervised Reinforcement Learning for Code Generation Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19ba90b6-35f8-4d14-996f-9b0f03fd7e62 · inbound
Efficiently Generating Expressive Quadruped Behaviors via Language-Guided Preference Learning Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f12cd1d-928d-4dad-8eb7-ce13dc188a92 · inbound
TruthFlow: Truthful LLM Generation via Representation Flow Correction Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a55315f-54a4-494a-955d-74c839955e2f · inbound
Use of Winsome Robots for Understanding Human Feedback (UWU) Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ceb3ae33-a214-48ea-a9b7-9747da5c0bc8 · inbound
Probabilistic Artificial Intelligence Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 2013
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1163456b-9200-41ef-9695-08a2fd6e6d56 · inbound
Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7dc7c782-b221-45b8-b989-3fd527b53881 · inbound
Thinking beyond the anthropomorphic paradigm benefits LLM research Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c96f6ae-cab5-45d5-bba6-5fbf314e9818 · inbound
Kaleidoscope Gallery: Exploring Ethics and Generative AI Through Art Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9bf3d758-8c50-40ea-88e2-3edf4d488e02 · inbound
A Unified Theoretical Analysis of Private and Robust Offline Alignment: from RLHF to DPO Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c234d425-3000-4e85-8a67-7c03b38df923 · inbound
Reality Check: A New Evaluation Ecosystem Is Necessary to Understand AI's Real World Effects Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a629e6ad-7be2-4c03-81fd-d1473e5d3b6d · inbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b351382a-f6ee-491e-9b37-d8518700b8c6 · inbound
Modeling and Optimizing User Preferences in AI Copilots: A Comprehensive Survey and Taxonomy Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 109
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c0b73b2-aa5c-49b1-990e-9906ac6bb7ed · inbound
Thompson Sampling in Online RLHF with General Function Approximation Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aaebfc88-f326-44ee-abbd-07aab08942f9 · inbound
Risks of AI-driven product development and strategies for their mitigation Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b0746f2-8130-4e13-81be-37aa6ee289b5 · inbound
Multi-objective Aligned Bidword Generation Model for E-commerce Search Advertising Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef7b58ca-d551-4fcd-bb97-aa7a6aee3c69 · inbound
Crowd-SFT: Crowdsourcing for LLM Alignment Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22e083f8-5051-4f3b-aa9e-db05d3809bee · inbound
AssertBench: A Benchmark for Evaluating Self-Assertion in Large Language Models Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a262c9d3-2641-4e1a-af74-b2c73504f47d · inbound
Exploring the Secondary Risks of Large Language Models Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation cab35d18-a447-4352-ac03-fd37e5153ac7 · inbound
Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cddfd517-9ac4-4910-a2d7-25882bebf5e6 · inbound
Collaborative Editable Model Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df5fe3e3-889a-428e-a270-c6de61da79d2 · inbound
Affective-CARA: A Knowledge Graph Driven Framework for Culturally Adaptive Emotional Intelligence in HCI Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1dd5e96f-9931-4ee4-84e8-729017ed3d0f · inbound
PrefPaint: Enhancing Medical Image Inpainting through Expert Human Feedback Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 77d01e76-e017-48ab-9720-a3eca423d311 · inbound
Lessons from a Chimp: AI "Scheming" and the Quest for Ape Language Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1cd622ca-9fb5-4544-a414-a41cc0fa6ecd · inbound
Exploring a Gamified Personality Assessment Method through Interaction with LLM Agents Embodying Different Personalities Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ac2747d2-67f7-42a0-9154-a4e87dbab96e · inbound
Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79a2f52d-fd27-4fbb-9f10-74f12a55cf3c · inbound
Granular feedback merits sophisticated aggregation Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81e681ba-1912-40da-ba7b-1d2f75ac1ce6 · inbound
PrefPalette: Personalized Preference Modeling with Latent Attributes Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 1952
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a5a0256-0ed2-4a11-a324-0a6ae5dfe732 · inbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b05867a-b2bb-41e9-bdae-b1cfd9d5b789 · inbound
Towards Mitigation of Hallucination for LLM-empowered Agents: Progressive Generalization Bound Exploration and Watchdog Monitor Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8b2e8f2-8d63-476c-91d6-e9393a28425a · inbound
RecGPT Technical Report Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 2009
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d235c00-dbf4-4475-9b75-70eb77925cd5 · inbound
Difficulty-Based Preference Data Selection by DPO Implicit Reward Gap Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 05948fb1-cc9e-4c62-833a-44cbbf013ac9 · inbound
Beyond Prediction: Reinforcement Learning as the Defining Leap in Healthcare AI Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd4f5822-fdd6-4f1e-b279-7c9a34e7a1ff · inbound
FocusDPO: Dynamic Preference Optimization for Multi-Subject Personalized Image Generation via Adaptive Focus Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97bb6238-1d2c-4042-b3e8-2ea3e4dddebe · inbound
SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 204f11e3-f7e5-475c-9fe1-17de2c24ec84 · inbound
The Anti-Ouroboros Effect: Emergent Resilience in Large Language Models from Recursive Selective Feedback Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95956f12-4d1f-464e-a2b7-b46706ebc560 · inbound
Large Foundation Models for Trajectory Prediction in Autonomous Driving: A Comprehensive Survey Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2f5499a-556a-4cb9-b502-68a2b81a64ec · inbound
EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2174ce8-2c66-4898-a84a-1f6dcd1579e8 · inbound
The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4bec7b9b-952a-4327-a2fe-21ec7b80fcf2 · inbound
A Multifaceted Analysis of Social Biases in Large Language Models Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e93959a-afe3-4f8a-977d-a8e32867b405 · inbound
Beyond Context: Large Language Models' Failure to Grasp Users' Intent Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ee760c3f-a2e5-4896-b32f-a326426bfb41 · inbound
ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 114285a7-d4c2-4d75-b1a5-181c265bfa08 · inbound
ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0e13fa8f-ac8b-46b0-a53d-0f724e22b70f · inbound
Acquiring Human-Like Data-Efficient Mechanics Prediction from Deep Reinforcement Learning Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14962ea5-6d68-42a0-b81d-f9a56741eaab · inbound
After Talking with 1,000 Personas: Learning Preference-Aligned Proactive Assistants From Large-Scale Persona Interactions Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c3f6a3f-121e-4ebc-bd54-8368f0cd73ba · inbound
Meta-Learning and Meta-Reinforcement Learning -- Tracing the Path towards DeepMind's Adaptive Agent Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 42c1d9f3-25d4-4821-a6a5-414e6fc24da4 · inbound
Beyond Semantic Manipulation: Token-Space Attacks on Reward Models Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8b647c86-fe95-4a56-842f-3bf8c19fe435 · inbound
The Theorems of Dr. David Blackwell and Their Contributions to Artificial Intelligence Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2e66137a-7514-48f5-b823-6564a678bd7a · inbound
The Art of (Mis)alignment: How Fine-Tuning Methods Effectively Misalign and Realign LLMs in Post-Training Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f2b2cbde-4eca-4447-9353-aeacf297aa06 · inbound
PlanGuard: Defending Agents against Indirect Prompt Injection via Planning-based Consistency Verification Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 13c48fb7-dd71-4edc-8adc-4923ed9218ac · inbound
Gaslight, Gatekeep, V1-V3: Early Visual Cortex Alignment Shields Vision-Language Models from Sycophantic Manipulation Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 36751cd0-2cc4-4a3e-8169-3fa3be83a811 · inbound
AIRA: AI-Induced Risk Audit: A Structured Inspection Framework for AI-Generated Code Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f7d9e837-e024-45cd-b794-36bf595eac1b · inbound
Policy Gradient Primal-Dual Method for Safe Reinforcement Learning from Human Feedback Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 21963a03-975b-4c12-8210-0430445ba567 · inbound
Post-AGI Economies: Autonomy and the First Fundamental Theorem of Welfare Economics Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d85e55c0-fb36-4e7e-8db8-167d471bb613 · inbound
Three Models of RLHF Annotation: Extension, Evidence, and Authority Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f4fbdf6f-ea2f-4df3-9781-2c01b6941981 · inbound
Learning from Disagreement: Clinician Overrides as Implicit Preference Signals for Clinical AI in Value-Based Care Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 594bfc7a-b5ee-4499-9129-3b49da21f9ec · inbound
Learning from Disagreement: Clinician Overrides as Implicit Preference Signals for Clinical AI in Value-Based Care Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 92b09ebb-7d35-401e-935b-ef0397d63fbc · inbound
Moira: Language-driven Hierarchical Reinforcement Learning for Pair Trading Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 10cbd289-f5b3-4176-b54d-eef634840de4 · inbound
Efficient Preference Poisoning Attack on Offline RLHF Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7bd7c218-ff88-40ea-9006-d14104726a17 · inbound
Dr. Post-Training: A Data Regularization Perspective on LLM Post-Training Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 124
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a7882be5-e14b-4d10-9a28-b824a2c66a79 · inbound
Curated Synthetic Data Doesn't Have to Collapse: A Theoretical Study of Generative Retraining with Pluralistic Preferences Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a228f44e-35ad-4a0f-88c3-af473d5152d3 · inbound
Reinforcement Learning for Scalable and Trustworthy Intelligent Systems Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 97b9a3c1-143e-47c4-be16-64e2753cb0cd · inbound
Can Revealed Preferences Clarify LLM Alignment and Steering? Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 28c1bddf-928e-457b-96f3-c92bf7e3840d · inbound
Power Reinforcement Post-Training of Text-to-Image Models with Super-Linear Advantage Shaping Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 9d69cc25-5418-41e5-97c0-3b07659707fb · inbound
Spurious Correlation Learning in Preference Optimization: Mechanisms, Consequences, and Mitigation via Tie Training Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 91045b5c-6b56-470a-80bc-2d0ff7c0f63f · inbound
Correcting Influence: Unboxing LLM Outputs with Orthogonal Latent Spaces Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 215
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 57819ffd-b01a-45c3-895e-8b7debcb0fda · inbound
Common-agency Games for Multi-Objective Test-Time Alignment Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 184
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ba3bf406-2a29-4790-b9ba-99a65925cc69 · inbound
TeamTR: Trust-Region Fine-Tuning for Multi-Agent LLM Coordination Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 9e08c13b-3780-4ef4-b8d5-341dd7ec9d55 · inbound
Formal Methods Meet LLMs: Auditing, Monitoring, and Intervention for Compliance of Advanced AI Systems Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d47a6bbf-8f01-48ce-ab89-8f3b3b24f7a1 · inbound
Preference Instability in Reward Models: Detection and Mitigation via Sparse Autoencoders Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 542eb250-0ca5-4b23-b1cc-734e8f95d18e · inbound
Some[Body] Must Receive That Pain for Agent Accountability Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 134094df-a128-4114-b62e-dc477c9eadb9 · inbound
ClaHF: A Human Feedback-inspired Reinforcement Learning Framework for Improving Classification Tasks Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation da158d90-1513-4cff-b6ca-bf2fc83d9b8f · inbound
Base Models Look Human To AI Detectors Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ae5723ab-12d5-4757-9601-5cb5fb29229b · inbound
REFLECTOR: Internalizing Step-wise Reflection against Indirect Jailbreak Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 85dbb226-d4be-4a22-8545-5ed867158fa0 · inbound
Echo: Learning from Experience Data via User-Driven Refinement Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b4c8e3d3-9f7a-4a62-8a41-4f51d03fa8c7 · inbound
Emotional intelligence in large language models is fragmented across perception, cognition, and interaction Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 17887a60-a748-4acf-a1e5-5c02444d73f5 · inbound
Cultivating Machine Intelligence: The OMEGA Shift from Top-Down Optimization to Autopoietic Cognitive Ecologies Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 988ba260-04d3-4ee6-a0e9-134ab4555582 · inbound
Does Capability Transfer to Subjective Behavior -- and Would Our Instruments Tell Us? A Self-Evolving, Trust-by-Construction Evaluation Paradigm Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4b911027-c79e-418a-a753-723082778d06 · inbound
In-Context Reward Adaptation for Robust Preference Modeling Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.