Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T18:21:37.480991Z
Paper Citation Record · LEDGER
As of 22 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 4 inbound Pith citation observations for arXiv:2501.11463.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T18:21:37.480991Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T11:50:28.097413Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T07:59:40.158119Z
55 of 55 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 75aaf1dd-d8b1-4051-8a58-6cb50b6fb916 · outbound
Curiosity-Driven Reinforcement Learning from Human Feedback Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b448799-681c-4306-a1f4-1edc01ff7d79 · outbound
Curiosity-Driven Reinforcement Learning from Human Feedback A General Language Assistant as a Laboratory for Alignment
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23a26cf7-86d9-44bf-b0bf-abff340a0304 · outbound
Curiosity-Driven Reinforcement Learning from Human Feedback Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48ae12b7-7a43-412d-897b-15367a132529 · outbound
Curiosity-Driven Reinforcement Learning from Human Feedback Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation f2faac84-9df4-4bec-aa82-43928cf6a9f8 · outbound
Curiosity-Driven Reinforcement Learning from Human Feedback Quality-Diversity through AI Feedback
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e50ef19-dd91-4319-a18b-3e5a0390f9a4 · outbound
Curiosity-Driven Reinforcement Learning from Human Feedback Unresolved cited work
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 666cd16a-1fb7-4351-a0a0-c78a4cfa3bed · outbound
Curiosity-Driven Reinforcement Learning from Human Feedback Unresolved cited work
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation c9bdcca6-4cd8-4dbb-99c0-759070e88f8a · outbound
Curiosity-Driven Reinforcement Learning from Human Feedback Robust Preference Learning for Storytelling via Contrastive Reinforcement Learning
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation f1dda98a-c9c6-40e0-9824-1d24df36789d · outbound
Curiosity-Driven Reinforcement Learning from Human Feedback MA-RLHF: Reinforcement Learning from Human Feedback with Macro Actions
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ebe5560-1554-4741-9855-523088a1fb29 · outbound
Curiosity-Driven Reinforcement Learning from Human Feedback Unresolved cited work
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ea5b238-dd95-4fb4-8867-54af34ec5c3f · outbound
Curiosity-Driven Reinforcement Learning from Human Feedback Unresolved cited work
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 842d757c-ead6-4c2d-8ce0-a471e3e31bc0 · outbound
Curiosity-Driven Reinforcement Learning from Human Feedback Christiano, Jan Leike, Tom B
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fbeb767d-7f03-4b23-a49c-a90dbf7d9824 · outbound
Curiosity-Driven Reinforcement Learning from Human Feedback UltraFeedback: Boosting Language Models with Scaled AI Feedback
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6585c1d5-7de4-4f01-91df-cf180f67d474 · outbound
Curiosity-Driven Reinforcement Learning from Human Feedback Unresolved cited work
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 5e36b825-d1cf-4589-9f7b-27b6dede4d57 · outbound
Curiosity-Driven Reinforcement Learning from Human Feedback The Llama 3 Herd of Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39dea6f4-620d-4b49-a5b4-3a966585eaa4 · outbound
Curiosity-Driven Reinforcement Learning from Human Feedback Unresolved cited work
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 18c92a1c-b7ba-4152-a81c-b2cac710e41e · outbound
Curiosity-Driven Reinforcement Learning from Human Feedback Unresolved cited work
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 1ecc97d1-0f34-4502-80d3-321822372c27 · outbound
Curiosity-Driven Reinforcement Learning from Human Feedback Glass, Akash Srivastava, and Pulkit Agrawal
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b51389eb-5a07-410e-899b-9d179eee7150 · outbound
Curiosity-Driven Reinforcement Learning from Human Feedback Unresolved cited work
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6413a6cc-e154-43e1-8836-14f1a8c47f90 · outbound
Curiosity-Driven Reinforcement Learning from Human Feedback Unresolved cited work
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 125e65f3-5493-4cb3-96fd-4659e6c2d048 · outbound
Curiosity-Driven Reinforcement Learning from Human Feedback Unresolved cited work
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22d86c07-4060-4598-a1d1-99bcc18ad473 · outbound
Curiosity-Driven Reinforcement Learning from Human Feedback Unresolved cited work
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d0f8288-3628-4560-af52-5966ab4753d4 · outbound
Curiosity-Driven Reinforcement Learning from Human Feedback Unresolved cited work
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 11c83419-a2d3-47d8-a4c9-2ad2b1337591 · outbound
Curiosity-Driven Reinforcement Learning from Human Feedback I-SHEEP: Self-Alignment of LLM from Scratch through an Iterative Self-Enhancement Paradigm
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 159aa98e-5493-4cc1-b04e-153dc25cde4b · outbound
Curiosity-Driven Reinforcement Learning from Human Feedback Unresolved cited work
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 73826fd4-a384-47c3-9445-292b58ca9ffd · outbound
Curiosity-Driven Reinforcement Learning from Human Feedback StarCoder 2 and The Stack v2: The Next Generation
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1848864-b63e-4121-9bf7-a0041bd85d02 · outbound
Curiosity-Driven Reinforcement Learning from Human Feedback Unresolved cited work
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation f3e900cd-80be-470a-9a06-5f795c75c074 · outbound
Curiosity-Driven Reinforcement Learning from Human Feedback Bellemare, A \"a ron van den Oord, and R \'e mi Munos
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 24ad0f3b-fb08-4d31-be67-10791b1ad772 · outbound
Curiosity-Driven Reinforcement Learning from Human Feedback Unresolved cited work
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9e33f2d-a512-4e51-90a8-fd0197310855 · outbound
Curiosity-Driven Reinforcement Learning from Human Feedback Unresolved cited work
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 1f42b683-9b06-4713-9877-755898c86271 · outbound
Curiosity-Driven Reinforcement Learning from Human Feedback Unresolved cited work
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b04304f-65e7-487a-9893-aab167ddb2c4 · outbound
Curiosity-Driven Reinforcement Learning from Human Feedback Unresolved cited work
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a42c2fe3-91bd-4bf5-abf6-c0c534bbc3a1 · outbound
Curiosity-Driven Reinforcement Learning from Human Feedback Unresolved cited work
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c38027f-6ba5-4ec3-9133-85a1c3cb501d · outbound
Curiosity-Driven Reinforcement Learning from Human Feedback Code Llama: Open Foundation Models for Code
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1da51b33-7f97-4def-a11e-51c4077e6036 · outbound
Curiosity-Driven Reinforcement Learning from Human Feedback Unresolved cited work
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 14f9ed15-7a7e-458f-9738-a57fd626bb02 · outbound
Curiosity-Driven Reinforcement Learning from Human Feedback Unresolved cited work
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ac394264-efe0-4f54-ab63-be50791917c2 · outbound
Curiosity-Driven Reinforcement Learning from Human Feedback Incentivizing Exploration In Reinforcement Learning With Deep Predictive Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a5e5681-7f98-48b0-8489-2c90bc12f057 · outbound
Curiosity-Driven Reinforcement Learning from Human Feedback Learning to summarize from human feedback
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 853d3ae0-f135-4c9e-bd84-3094b2799e6b · outbound
Curiosity-Driven Reinforcement Learning from Human Feedback Gemma: Open Models Based on Gemini Research and Technology
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73c0118b-78fe-46a0-ba7c-35517fe81d22 · outbound
Curiosity-Driven Reinforcement Learning from Human Feedback Unresolved cited work
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e63fdbf0-1375-47de-9bdf-14c58977c34a · outbound
Curiosity-Driven Reinforcement Learning from Human Feedback Solving math word problems with process- and outcome-based feedback
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6992dfdd-56f4-45eb-92eb-93f79a76edfd · outbound
Curiosity-Driven Reinforcement Learning from Human Feedback Unresolved cited work
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 374518da-571b-4f5c-8973-6f09962a0e53 · outbound
Curiosity-Driven Reinforcement Learning from Human Feedback Unresolved cited work
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 899a1e39-071f-4e75-a52c-c0f63297ae94 · outbound
Curiosity-Driven Reinforcement Learning from Human Feedback Generative Monoculture in Large Language Models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b99d810a-0b1b-4ab1-bb12-87fcabc06165 · outbound
Curiosity-Driven Reinforcement Learning from Human Feedback Unresolved cited work
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 86be2fbe-e6bc-4699-9001-2d4920c47f3a · outbound
Curiosity-Driven Reinforcement Learning from Human Feedback WizardLM: Empowering large pre-trained language models to follow complex instructions
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0272511-cb13-4617-85da-c6a49d36e289 · outbound
Curiosity-Driven Reinforcement Learning from Human Feedback Unresolved cited work
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 2f8f7139-3aad-4d1e-916c-170fa4030d46 · outbound
Curiosity-Driven Reinforcement Learning from Human Feedback SEED-Story: Multimodal Long Story Generation with Large Language Model
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9fb83aab-dfe0-4eaa-83ea-22db674c60c9 · outbound
Curiosity-Driven Reinforcement Learning from Human Feedback DeepSpeed-Chat: Easy, Fast and Affordable RLHF Training of ChatGPT-like Models at All Scales
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2eb9e11-3434-471e-95ae-9fe5d863cd18 · outbound
Curiosity-Driven Reinforcement Learning from Human Feedback Unresolved cited work
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 631f5bbc-199c-44e1-b99f-bda9d3859013 · outbound
Curiosity-Driven Reinforcement Learning from Human Feedback Unresolved cited work
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a8021b5-4082-424a-ba4e-fc28d6241e91 · outbound
Curiosity-Driven Reinforcement Learning from Human Feedback Unresolved cited work
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation c206f32a-14d1-4478-a8e1-af62c62500df · outbound
Curiosity-Driven Reinforcement Learning from Human Feedback Fine-Tuning Language Models from Human Preferences
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f1d644f-fc87-4860-8618-e5a40ceaf127 · outbound
Curiosity-Driven Reinforcement Learning from Human Feedback online" 'onlinestring :=
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe0568f0-7da8-45a5-b224-a8057dea6954 · outbound
Curiosity-Driven Reinforcement Learning from Human Feedback write newline
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1f33125-130b-4524-aa1c-2a2cadc76e4a · inbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Curiosity-Driven Reinforcement Learning from Human Feedback
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67b78455-0897-42a2-acf6-43fce24cef22 · inbound
Avoidance Decoding for Diverse Multi-Branch Story Generation Curiosity-Driven Reinforcement Learning from Human Feedback
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ba30e36-23e4-44aa-bc17-7e45b0c1344d · inbound
Quality-constrained Entropy Maximization Policy Optimization for LLM Diversity Curiosity-Driven Reinforcement Learning from Human Feedback
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3383f364-3f79-4012-bfb2-90f49dbc3e79 · inbound
Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning Curiosity-Driven Reinforcement Learning from Human Feedback
Reference 187
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.