Pith. sign in

Paper Citation Record · LEDGER

SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling

As of 16 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2608.02951.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.02951 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:02:45.665708Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact1
  • verified fuzzy7
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ee422279-496a-4840-9e7d-bfb307e0ec00 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T15:02:45.558196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:02:45.558196Z digest=sha256:0f8e60f533331a41f821c59804dbd1c45768481cf3e432481a877cc75a1fcbd3

Observation d290f816-01cf-4562-8b67-462bbce1b352 · outbound

This paper cites Contrastive Preference Learning: Learning from Human Feedback without RL.

SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling Contrastive Preference Learning: Learning from Human Feedback without RL

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T15:02:45.577850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:02:45.577850Z digest=sha256:0f8a558c32f45567b5b6085746a463a26d78f1bc3de73071b70f33addb9136bb

Observation 98b5c4aa-31eb-42e3-a4ba-65e39627fcad · outbound

This paper cites Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback.

SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T15:02:45.582408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:02:45.582408Z digest=sha256:6c366f227b83209b1c177d89fd181fd3e9e616e44e9ae06a407e409a48dbf3d1

Observation 3a480b73-c493-4684-92e3-dccca44d5ce0 · outbound

This paper cites RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models.

SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T15:02:45.609451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:02:45.609451Z digest=sha256:520f06d2dbc7d5e9bc59a2e2152247f168749f30c46d399d8232ffcb0e38fcb3

Observation 2f6a9377-259a-45e1-9173-21492ddfe58f · outbound

This paper cites Reinforcement Learning with Segment Feedback.

SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling Reinforcement Learning with Segment Feedback

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-15T15:02:45.752291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:02:45.617996Z digest=sha256:2d4a3de13d8499d84f5408aa64bf6809da25cfce8ebfeff1270b9c2566a42b46

Observation 605b043e-97f6-4fc4-89d8-ad25e4b7c7df · outbound

This paper cites Adam: A Method for Stochastic Optimization.

SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling Adam: A Method for Stochastic Optimization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T15:02:45.622442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:02:45.622442Z digest=sha256:571f9120146ac36322fe3725ed4fe7bb65a516275d57d986d0bb88e02f593b87

Observation 08c4f306-f481-4970-b9c9-d8c977c836e5 · outbound

This paper cites Decoupled Weight Decay Regularization.

SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling Decoupled Weight Decay Regularization

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T15:02:45.631715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:02:45.631715Z digest=sha256:d915d13cbc67a1be50cae93f892c3c2ef41a882e9d0c1ce4a3352a702e145d26

Observation e28fc222-6012-45fa-b84a-03663fb3c60f · outbound

This paper cites Approximating kl divergence, 2020.URL http://joschu.

SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling Approximating kl divergence, 2020.URL http://joschu

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:02:46.087144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:02:45.636173Z digest=sha256:26bb21de1c2908637bfd80180c8b9d3172468f2744601a8a6946271ad9614eda

Observation 4fb484bb-1eb7-4c64-a6d9-305513f72586 · outbound

This paper cites 15 A.2 Comparison of Preference Models.

SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling 15 A.2 Comparison of Preference Models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:02:46.071148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:02:45.641390Z digest=sha256:d7f063a29ff5405a054ae1c1572fdf324f5f8ac4ab5df7b0cac60672858941fe

Observation ba2a8584-11b2-422a-9c6f-84a761d8d5d0 · outbound

This paper cites This makes them inapplicable to many modern RL problems.

SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling This makes them inapplicable to many modern RL problems

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:02:46.055188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:02:45.646471Z digest=sha256:e306ea5a928b16c74fed280b7c751b96c128b7aab32d670f10ca486571dc62f1

Observation 13ac6ae2-d619-4207-9ba0-d10702023097 · outbound

This paper cites Specifically, even if the partial returns of two segments are the same, 15 humans could still prefer one trajectory segment based on the value of the last state and action.

SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling Specifically, even if the partial returns of two segments are the same, 15 humans could still prefer one trajectory segment based on the value of the last state and action

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:02:46.040674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:02:45.651001Z digest=sha256:8537953073cafe1680b6498b581f82f2534225d9bbad3dcdc07a9fa717767b8c

Observation 156a5eb1-7f77-42e1-b4bb-c7055b8afd38 · outbound

This paper cites During training, policies were Gaussian with a parameter controlling the log standard deviation for each dimension in the action space.

SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling During training, policies were Gaussian with a parameter controlling the log standard deviation for each dimension in the action space

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:02:46.010880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:02:45.660952Z digest=sha256:27db6100fa2650fee292aaef62249f767b201a0a3ce4ff204c17479511ba029b

Observation ff72a96f-415a-4d8e-9695-9a1ab3b7e032 · outbound

This paper cites The global gradient norm is clipped to1.0.

SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling The global gradient norm is clipped to1.0

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:02:45.996416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:02:45.665708Z digest=sha256:4e9a224fb7b736c45f962f15d0cfecc34d0ef733230c3c39b0c319f6ce38b0dd

Observation fe301495-77e8-4b26-9e33-1856d6f1db27 · outbound

This paper cites We observe that in Ant-v5 usingπθt´1 as the second policy out performs setting the second policy toπθref.

SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling We observe that in Ant-v5 usingπθt´1 as the second policy out performs setting the second policy toπθref

Reference 1000

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:02:46.025703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:02:45.655771Z digest=sha256:aa183ad06022b0c5e07a0331ec986ae2051c75eeae7a9a31ef267c60590952a0

Observation 34a6412f-ff93-4201-886c-6617e19fa7e0 · outbound

This paper cites Models of human preference for learning reward functions.

SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling Models of human preference for learning reward functions

Reference 1952

Resolution
unresolved
no resolver link, observed 2026-08-15T15:02:45.573168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:02:45.573168Z digest=sha256:e9491d44d25e45b73b15a48a03a61a12a3f8c8acab71e295c43633b7f2c42912

Observation 14ed6ea2-811c-45d0-84b0-51b5553374de · outbound

This paper cites Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs.

SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-15T15:02:45.591812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:02:45.591812Z digest=sha256:39e0627696ac5d493dbec72064633e57203daad63699f171a4ea9588c3da45cc

Observation 1ff964b8-3fa5-467a-8dfa-88ee91dd45af · outbound

This paper cites Proximal Policy Optimization Algorithms.

SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling Proximal Policy Optimization Algorithms

Reference 2013

Resolution
unresolved
no resolver link, observed 2026-08-15T15:02:45.548076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:02:45.548076Z digest=sha256:b4268581266a0b96cc688be4e08ff18785a38b522593860ded51c7458ea01909

Observation 52bf253f-7a60-4fdd-acb2-db84ff3810f2 · outbound

This paper cites Making Reinforcement Learning Work on Swimmer.

SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling Making Reinforcement Learning Work on Swimmer

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-15T15:02:45.626921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:02:45.626921Z digest=sha256:01b55bb9dc84392ce877ccf0e2e6b12a02a4eb44b73c60773221950cc85a454d

Observation 526c189c-caec-415d-a723-96ba1e20aceb · outbound

This paper cites PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-training.

SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-training

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-15T15:02:45.601173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:02:45.601173Z digest=sha256:7e5af111217c18a4052ee065f196d8de7b721dc056cac5e3b38989bcdc287f56

Observation 179e76b8-8da8-4c1f-8fe0-de3fa777a8c2 · outbound

This paper cites Aligning Text-to-Image Models using Human Feedback.

SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling Aligning Text-to-Image Models using Human Feedback

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-15T15:02:45.567940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:02:45.567940Z digest=sha256:b135b30bdda3e30d222b0242f2b7de894c4db603b76f84e06716d265809a7b1d

Observation b0d7fa27-dac7-474f-9532-0b7e02d9daa7 · outbound

This paper cites Playing Atari with Deep Reinforcement Learning.

SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling Playing Atari with Deep Reinforcement Learning

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-15T15:02:45.542769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:02:45.542769Z digest=sha256:7997090ad575b6bc6faa8b8acd519a815ee378e849152fa3c19ca88dc77dede9

Observation 38abb7ee-237b-4218-993a-661c3d4087de · outbound

This paper cites Preference Transformer: Modeling Human Preferences using Transformers for RL.

SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling Preference Transformer: Modeling Human Preferences using Transformers for RL

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-15T15:02:45.613860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:02:45.613860Z digest=sha256:ebf295ce9966e9fac998c44d374b789d77bdd84261b414f0acdb117aad072973

Observation b85b8801-d6b7-4c14-a2b0-d595bc8b65a0 · outbound

This paper cites Gymnasium: A Standard Interface for Reinforcement Learning Environments.

SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-15T15:02:45.605223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:02:45.605223Z digest=sha256:c77c3c5f5b62253cdbf2c796b77e0603b0b9101ded3192b67e023b427e229594

Observation ea2ce2c8-9538-4056-8515-e1ea757041a9 · outbound

This paper cites RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback.

SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-15T15:02:45.552796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:02:45.552796Z digest=sha256:6ccaba8c773b6e36967388bc9d6369a7e6997441f61923c0fe5d882b6e5ad877

Observation 227b8e35-4610-4640-bc74-c6d145f81fd4 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-15T15:02:45.563041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:02:45.563041Z digest=sha256:433968a16ddb4cc6b664405e5b860fcea067768419c45e322f9bc5232bef7dae

Observation f7f7b12a-67c5-4df6-b868-14f03998d776 · outbound

This paper cites Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function.

SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-15T15:02:45.587036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:02:45.587036Z digest=sha256:89eab526ee9f2fcd135e459be55921e3d051bc2a6cfcb97a53c59a401ef24fe6

Observation 5d978a87-4205-43fd-bd9a-0b116899fbf5 · outbound

This paper cites Direct Language Model Alignment from Online AI Feedback.

SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling Direct Language Model Alignment from Online AI Feedback

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-15T15:02:45.596461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:02:45.596461Z digest=sha256:ef05718b2841bc4397a7c2372b96937880991ad0d0c75497e810b826214ccfac

Pith citing papers

No inbound Pith citation observations are available.