Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T15:02:45.665708Z
Paper Citation Record · LEDGER
As of 16 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2608.02951.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T15:02:45.665708Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
27 of 27 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation ee422279-496a-4840-9e7d-bfb307e0ec00 · outbound
SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d290f816-01cf-4562-8b67-462bbce1b352 · outbound
SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling Contrastive Preference Learning: Learning from Human Feedback without RL
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98b5c4aa-31eb-42e3-a4ba-65e39627fcad · outbound
SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a480b73-c493-4684-92e3-dccca44d5ce0 · outbound
SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f6a9377-259a-45e1-9173-21492ddfe58f · outbound
SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling Reinforcement Learning with Segment Feedback
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 605b043e-97f6-4fc4-89d8-ad25e4b7c7df · outbound
SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling Adam: A Method for Stochastic Optimization
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08c4f306-f481-4970-b9c9-d8c977c836e5 · outbound
SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling Decoupled Weight Decay Regularization
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e28fc222-6012-45fa-b84a-03663fb3c60f · outbound
SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling Approximating kl divergence, 2020.URL http://joschu
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 4fb484bb-1eb7-4c64-a6d9-305513f72586 · outbound
SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling 15 A.2 Comparison of Preference Models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation ba2a8584-11b2-422a-9c6f-84a761d8d5d0 · outbound
SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling This makes them inapplicable to many modern RL problems
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 13ac6ae2-d619-4207-9ba0-d10702023097 · outbound
SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling Specifically, even if the partial returns of two segments are the same, 15 humans could still prefer one trajectory segment based on the value of the last state and action
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 156a5eb1-7f77-42e1-b4bb-c7055b8afd38 · outbound
SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling During training, policies were Gaussian with a parameter controlling the log standard deviation for each dimension in the action space
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation ff72a96f-415a-4d8e-9695-9a1ab3b7e032 · outbound
SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling The global gradient norm is clipped to1.0
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation fe301495-77e8-4b26-9e33-1856d6f1db27 · outbound
SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling We observe that in Ant-v5 usingπθt´1 as the second policy out performs setting the second policy toπθref
Reference 1000
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 34a6412f-ff93-4201-886c-6617e19fa7e0 · outbound
SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling Models of human preference for learning reward functions
Reference 1952
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14ed6ea2-811c-45d0-84b0-51b5553374de · outbound
SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 2012
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ff964b8-3fa5-467a-8dfa-88ee91dd45af · outbound
SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling Proximal Policy Optimization Algorithms
Reference 2013
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52bf253f-7a60-4fdd-acb2-db84ff3810f2 · outbound
SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling Making Reinforcement Learning Work on Swimmer
Reference 2014
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 526c189c-caec-415d-a723-96ba1e20aceb · outbound
SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-training
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 179e76b8-8da8-4c1f-8fe0-de3fa777a8c2 · outbound
SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling Aligning Text-to-Image Models using Human Feedback
Reference 2018
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0d7fa27-dac7-474f-9532-0b7e02d9daa7 · outbound
SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling Playing Atari with Deep Reinforcement Learning
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38abb7ee-237b-4218-993a-661c3d4087de · outbound
SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling Preference Transformer: Modeling Human Preferences using Transformers for RL
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b85b8801-d6b7-4c14-a2b0-d595bc8b65a0 · outbound
SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling Gymnasium: A Standard Interface for Reinforcement Learning Environments
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea2ce2c8-9538-4056-8515-e1ea757041a9 · outbound
SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 227b8e35-4610-4640-bc74-c6d145f81fd4 · outbound
SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7f7b12a-67c5-4df6-b868-14f03998d776 · outbound
SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d978a87-4205-43fd-bd9a-0b116899fbf5 · outbound
SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling Direct Language Model Alignment from Online AI Feedback
Reference 2026
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.