Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 47 inbound Pith citation observations for arXiv:2304.05302.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:48:04.370319Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-06-30T23:15:07.886084Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 72976909-ea44-46d9-88ba-833ff05566db · inbound
RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment RRHF: Rank Responses to Align Language Models with Human Feedback without tears
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9ec7fa5b-d481-4847-b890-79898516f46a · inbound
WizardLM: Empowering large pre-trained language models to follow complex instructions RRHF: Rank Responses to Align Language Models with Human Feedback without tears
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a59b01a7-cea2-4b29-b9d2-895df4f38fb0 · inbound
A Comprehensive Overview of Large Language Models RRHF: Rank Responses to Align Language Models with Human Feedback without tears
Reference 170
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e61a961f-06fa-44a7-8530-96ed0368ac11 · inbound
Trustworthy LLMs: a Survey and Guideline for Evaluating Large Language Models' Alignment RRHF: Rank Responses to Align Language Models with Human Feedback without tears
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 64edbcdf-4674-4375-b44a-2eb0f48d4618 · inbound
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models RRHF: Rank Responses to Align Language Models with Human Feedback without tears
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9d66a7fd-3736-466f-abf7-3f142b5d1427 · inbound
A Survey on Knowledge Distillation of Large Language Models RRHF: Rank Responses to Align Language Models with Human Feedback without tears
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8fa9f457-f4b5-46d5-b3bf-c65a6b2866e0 · inbound
UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types RRHF: Rank Responses to Align Language Models with Human Feedback without tears
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b510fc61-533d-4dc4-b764-8299bbf4e1bc · inbound
Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey RRHF: Rank Responses to Align Language Models with Human Feedback without tears
Reference 175
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation baab4df5-0a6d-4668-9373-c572083d1de9 · inbound
HybridFlow: A Flexible and Efficient RLHF Framework RRHF: Rank Responses to Align Language Models with Human Feedback without tears
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1944b7f2-59b0-450d-b4d3-a13dc4e8c1eb · inbound
A Survey on LLM-as-a-Judge RRHF: Rank Responses to Align Language Models with Human Feedback without tears
Reference 200
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation eb4a2363-412a-44da-98fa-a8a3cae7000a · inbound
AnnoDPO: Protein Functional Annotation Learning with Direct Preference Optimization RRHF: Rank Responses to Align Language Models with Human Feedback without tears
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6c9871d-7ba5-4d27-ba1d-24fa4ab96c53 · inbound
AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models RRHF: Rank Responses to Align Language Models with Human Feedback without tears
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d3ab8c7-82c6-472a-85b1-493e848b096b · inbound
Value-Free Policy Optimization via Reward Partitioning RRHF: Rank Responses to Align Language Models with Human Feedback without tears
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a45831ae-3018-49f0-b5f9-0c5dc3b00652 · inbound
From General to Targeted Rewards: Surpassing GPT-4 in Open-Ended Long-Context Generation RRHF: Rank Responses to Align Language Models with Human Feedback without tears
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8f5b9bb-35ec-4f6f-84c6-d5d8950363da · inbound
Activation Reward Models for Few-Shot Model Alignment RRHF: Rank Responses to Align Language Models with Human Feedback without tears
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de9c2c0e-ad7c-44cd-93c5-03e4f816e9f0 · inbound
Flipping Knowledge Distillation: Leveraging Small Models' Expertise to Enhance LLMs in Text Matching RRHF: Rank Responses to Align Language Models with Human Feedback without tears
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5e67739-421f-4012-a12f-e1cddf0cac87 · inbound
Enhancing Safe and Controllable Protein Generation via Knowledge Preference Optimization RRHF: Rank Responses to Align Language Models with Human Feedback without tears
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26a849c7-7e62-4a1a-9237-5698fd209cb9 · inbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities RRHF: Rank Responses to Align Language Models with Human Feedback without tears
Reference 110
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48f014e5-7d23-4d3f-bc37-cd9703e6a585 · inbound
Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges RRHF: Rank Responses to Align Language Models with Human Feedback without tears
Reference 260
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8cc6708-05ee-4ece-a79a-1f80bff6b554 · inbound
A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence RRHF: Rank Responses to Align Language Models with Human Feedback without tears
Reference 196
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b5d488ce-ff01-4ef7-b4b2-c88ba28494c0 · inbound
Secure Tug-of-War (SecTOW): Iterative Defense-Attack Training with Reinforcement Learning for Multimodal Model Security RRHF: Rank Responses to Align Language Models with Human Feedback without tears
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6735acff-92f5-4e47-a1e8-a09b988f2502 · inbound
Beyond Surface-Level Detection: Towards Cognitive-Driven Defense Against Jailbreak Attacks via Meta-Operations Reasoning RRHF: Rank Responses to Align Language Models with Human Feedback without tears
Reference 122
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e311480-4017-468b-932f-7df4c5757651 · inbound
Flexible Agent Alignment with Goal Inference from Open-Ended Dialog RRHF: Rank Responses to Align Language Models with Human Feedback without tears
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 768289db-0364-4ac0-93cd-24ed24bee1a4 · inbound
PKG-DPO: Optimizing Domain-Specific AI systems with Physics Knowledge Graphs and Direct Preference Optimization RRHF: Rank Responses to Align Language Models with Human Feedback without tears
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79211895-3090-4cd1-92ae-5e0fbb5630a0 · inbound
Failure Modes of Maximum Entropy RLHF RRHF: Rank Responses to Align Language Models with Human Feedback without tears
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 875c2ac0-b05f-4cb7-9448-571525ed98bb · inbound
Simultaneous Multi-objective Alignment Across Verifiable and Non-verifiable Rewards RRHF: Rank Responses to Align Language Models with Human Feedback without tears
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 722fc942-83e3-4dc7-bde2-9b6156d27149 · inbound
POPI: Personalizing LLMs via Optimized Natural Language Preference Inference RRHF: Rank Responses to Align Language Models with Human Feedback without tears
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 30d48663-d659-4a54-bc46-2fd92d99faf2 · inbound
OOWM: Structuring Embodied Reasoning and Planning via Object-Oriented Programmatic World Modeling RRHF: Rank Responses to Align Language Models with Human Feedback without tears
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 31089a63-6ed0-4a01-8500-42cf20d59c9d · inbound
GroupDPO: Memory efficient Group-wise Direct Preference Optimization RRHF: Rank Responses to Align Language Models with Human Feedback without tears
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9cfa3750-ac93-497a-965b-151d7f1f4984 · inbound
RLearner-LLM: Balancing Logical Grounding and Fluency in Large Language Models via Hybrid Direct Preference Optimization RRHF: Rank Responses to Align Language Models with Human Feedback without tears
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 54258816-4d2a-4120-b1c3-b6ea9f080ad3 · inbound
RLearner-LLM: Balancing Logical Grounding and Fluency in Large Language Models via Hybrid Direct Preference Optimization RRHF: Rank Responses to Align Language Models with Human Feedback without tears
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7bb9b7cd-b001-41b4-9ec4-06c1ae9352fe · inbound
RLearner-LLM: Balancing Logical Grounding and Fluency in Large Language Models via Hybrid Direct Preference Optimization RRHF: Rank Responses to Align Language Models with Human Feedback without tears
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 985730b5-b7d3-4c31-88ed-4b0322f4ae3e · inbound
RLearner-LLM: Balancing Logical Grounding and Fluency in Large Language Models via Hybrid Direct Preference Optimization RRHF: Rank Responses to Align Language Models with Human Feedback without tears
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9bc0e629-2ba4-4e2f-bfeb-16dbda53b79a · inbound
PairAlign: A Framework for Sequence Tokenization via Self-Alignment with Applications to Audio Tokenization RRHF: Rank Responses to Align Language Models with Human Feedback without tears
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 04d0c94e-1d3a-44d2-8183-29975cecc6ec · inbound
PairAlign: A Framework for Sequence Tokenization via Self-Alignment with Applications to Audio Tokenization RRHF: Rank Responses to Align Language Models with Human Feedback without tears
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 96dbc87c-d15d-4ab1-a18a-ff08132f1f16 · inbound
Response Time Enhances Alignment with Heterogeneous Preferences RRHF: Rank Responses to Align Language Models with Human Feedback without tears
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation da909d21-cd01-46ed-9ba2-568e47c4b775 · inbound
Offline Preference Optimization for Rectified Flow with Noise-Tracked Pairs RRHF: Rank Responses to Align Language Models with Human Feedback without tears
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 34a8d2c4-5ccb-461c-abae-f3fc672c673d · inbound
TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching RRHF: Rank Responses to Align Language Models with Human Feedback without tears
Reference 108
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 987a504b-d28b-4e5f-b1fc-d8f7db65f81c · inbound
TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching RRHF: Rank Responses to Align Language Models with Human Feedback without tears
Reference 108
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 70559131-adfd-4148-8834-7d0046e98dde · inbound
Convex Optimization for Alignment and Preference Learning on a Single GPU RRHF: Rank Responses to Align Language Models with Human Feedback without tears
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 79fba758-3eae-47f0-aff4-3efa0a651e42 · inbound
BV-Blend: Uncertainty-Weighted Historical Baselines for Stable Critic-Free RL with Verifiable Rewards RRHF: Rank Responses to Align Language Models with Human Feedback without tears
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 70a8db04-0fd0-4fd6-be39-2f830736006c · inbound
Multi-Turn On-Policy Distillation with Prefix Replay RRHF: Rank Responses to Align Language Models with Human Feedback without tears
Reference 157
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc17eec0-cf70-4c10-bc40-2a677e905421 · inbound
Multi-Turn On-Policy Distillation with Prefix Replay RRHF: Rank Responses to Align Language Models with Human Feedback without tears
Reference 158
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98eb984c-4f54-4d6e-aa35-87962f7db8cd · inbound
Every Sample Counts: Supervised Fine-Tuning of Language Models with Pointwise Constraints RRHF: Rank Responses to Align Language Models with Human Feedback without tears
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3585521-9000-4a22-9543-274e88dc66bb · inbound
Test-Time Scaling via Error Localization RRHF: Rank Responses to Align Language Models with Human Feedback without tears
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da752a54-d624-4528-859d-44f132c7ca38 · inbound
SCOPE and SCION: A Benchmark and an Auditable Reference Pipeline for Schema Induction and Fusion from Text RRHF: Rank Responses to Align Language Models with Human Feedback without tears
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e36ee06a-daf0-477f-a23d-7923d915f246 · inbound
Quo Vadis, World Modeling? RRHF: Rank Responses to Align Language Models with Human Feedback without tears
Reference 211
Source-reported events for the cited work
Unavailable: canonical work link unavailable.