Pith. sign in

Paper Citation Record · LEDGER

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function

As of 7 August 2026, this Paper Citation Record lists 87 of 87 outbound references and 1 inbound Pith citation observation for arXiv:2506.03066.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.03066 v2

Coverage vector

measured 87 of 87 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:20:15.714867Z

measured 88 of 88 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T04:48:15.394329Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-10T11:35:19.307998Z

Reference resolution

87 of 87 outbound references displayed

  • verified exact1
  • verified fuzzy44
  • unresolved42
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e9703832-2747-4b73-8c47-4467da086c5e · outbound

This paper cites Reinforcement learning: An introduction, volume 1.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Reinforcement learning: An introduction, volume 1

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:10.114854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:10.114854Z digest=sha256:1cb07115b97499a98a2be2e28b0e97308acc93e432f35352aae15aed2e514a5b

Observation 70fe8ede-870e-4e44-ae03-67b1798c8153 · outbound

This paper cites Controlled experiments on the web: survey and practical guide.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Controlled experiments on the web: survey and practical guide

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:10.203695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:10.203695Z digest=sha256:07c56fd26958c106ecd6a954b94c499a18e46185be18ca14c1c2d8d693bf4f4a

Observation 2de9bc8e-27eb-4749-95be-5ba5d2ac8ec5 · outbound

This paper cites Deep reinforcement learning from human preferences.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Deep reinforcement learning from human preferences

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:10.296587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:10.296587Z digest=sha256:8f33e93f5397250fe16c4975dd4af8d72e03a49f05151391516fd41b86f447eb

Observation 2916cde9-d8b5-465c-97bc-8968f3927110 · outbound

This paper cites Al Sallab, Senthil Yogamani, and Patrick Pérez.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Al Sallab, Senthil Yogamani, and Patrick Pérez

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:10.405514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:10.405514Z digest=sha256:61808870b64a85d6ec1ed0286166825aa23155b11647d210c8c91c9dab2d6f32

Observation 739bf0c4-c7c9-4cd1-bc45-3be231cf17ea · outbound

This paper cites Training language models to follow instructions with human feedback.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Training language models to follow instructions with human feedback

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:10.525911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:10.525911Z digest=sha256:7ff78368d10b724fa60769957f5acb7333f313630d696cd1bb25d458cf04b6c4

Observation ca01d7dd-fa85-480d-9975-8a0b310cb5f4 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:10.608779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:10.608779Z digest=sha256:0d9c54f1aa8a05afe3efe155c49b55ec48baf1fca8d9d86230a3c4ceeb3e0d0e

Observation 99c7fe35-9628-4b78-b4ab-abb6f85799f3 · outbound

This paper cites Dynamic programming and stochastic control processes.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Dynamic programming and stochastic control processes

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:10.716180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:10.716180Z digest=sha256:3da3f3fe8a6c9976e98f9de76ed27a645f79f882864ec8e5d55b0a744bffaccd

Observation 192aba53-7d83-4907-ad1b-7861438a4a85 · outbound

This paper cites Markov decision processes: discrete stochastic dynamic programming.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Markov decision processes: discrete stochastic dynamic programming

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:10.827799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:10.827799Z digest=sha256:73193db92a922b8694bd9f6c2fa2a648778b9a482951d8fe4d0317dd0db39e87

Observation 5d6406f9-6fc8-452a-874a-8d9f581f79e9 · outbound

This paper cites Inverse reward design.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Inverse reward design

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:10.891470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:10.891470Z digest=sha256:b551cf8cc08fc89ee43da3057bc58d4a7f6a0449e2b1463c71d944a332ef57dd

Observation de54f59e-6712-4905-8a83-7fe6714d6b9d · outbound

This paper cites Reward Design with Language Models.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Reward Design with Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:10.973403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:10.973403Z digest=sha256:b210f08aef1d4820b94d15986234956d004ab9352d12d15335590ac516d02c92

Observation 6f2ccb55-f03d-4c47-9196-aa7b47ef13c6 · outbound

This paper cites Defining and characterizing reward gaming.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Defining and characterizing reward gaming

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:11.056676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:11.056676Z digest=sha256:0081d700c87f71f964e7c78c021bbac2ff8b4c21927ccd39a18c9b8583dcbe23

Observation c50e7a09-cd8c-4f95-834e-d184db24cd46 · outbound

This paper cites Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:11.242520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:11.242520Z digest=sha256:dbfbf40a61f4e00a8816103cfca96516a421759112de35bcbf726be9c6feb877

Observation 31d800c5-7bad-4ed5-8efc-6e2d80ecf082 · outbound

This paper cites GPT-4o System Card.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function GPT-4o System Card

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:11.353435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:11.353435Z digest=sha256:dd3badb7cc757ed0e541ddaba719c7637ee84057b5c173e570bb9c3ab2efb8a7

Observation 1534fe5b-74b5-4f5f-a2bb-9c1b5495e071 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Direct preference optimization: Your language model is secretly a reward model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:11.439190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:11.439190Z digest=sha256:4cf4ef99209725c165968d697e968f87516bf08b5ae7c825102327f858977ef0

Observation e5c0ff52-d75e-47a2-bb45-e07577dc1094 · outbound

This paper cites Zeroth-order policy gradient for reinforcement learning from human feedback without reward inference.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Zeroth-order policy gradient for reinforcement learning from human feedback without reward inference

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:24.434223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:20:11.554482Z digest=sha256:8ea0f449ee3e16283b34a3fc354c74dde56c48c43194a1fda204c61416a2bc37

Observation 5def0db4-fbcf-475d-b4d3-3d100ea593d7 · outbound

This paper cites Proximal policy optimization algorithms, 2017.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Proximal policy optimization algorithms, 2017

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:11.665749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:11.665749Z digest=sha256:7e1474e0556c54e61f2a5c2357b95dfda49ee1c891f60f6ef324b1f64550afc5

Observation a80d8297-66c6-4289-8a23-988584021b24 · outbound

This paper cites Open problems and fundamental limitations of reinforcement learning from human feedback.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Open problems and fundamental limitations of reinforcement learning from human feedback

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:24.322900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:20:11.755552Z digest=sha256:5e1b530bf84574d0f57609210efdefba376cd7871959d9b14964f859af9c0356

Observation 2b6d06d7-162f-4354-abc4-664102270b38 · outbound

This paper cites Lee, Wotao Yin, Mingyi Hong, Zhangyang Wang, Sijia Liu, and Tianlong Chen.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Lee, Wotao Yin, Mingyi Hong, Zhangyang Wang, Sijia Liu, and Tianlong Chen

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:24.164237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:20:11.793743Z digest=sha256:212d61507a5a4c2fce5f96b363147aac6cf762c79f0d701bceacec8a29275c79

Observation 4656b496-a7a3-473d-9590-2e947e847154 · outbound

This paper cites From r to q^ * : Your language model is secretly a q-function.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function From r to q^ * : Your language model is secretly a q-function

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:23.963362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:20:11.799620Z digest=sha256:357971e0cac873fd12f7915de7044b2f37ece1f9d6d1f9ddf47d7a4543956a37

Observation b76e2247-cc93-4d0b-841a-15c87dfc1c31 · outbound

This paper cites an unresolved cited work.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:20:23.787862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:20:11.839237Z digest=sha256:f9d2280d14f4a1a776a243b0cd7690e5bcfb53b19893b45c0031cee207ef1985

Observation 2105833b-782a-4fa4-b5d9-68dbaa0f8459 · outbound

This paper cites Random utility theory for social choice.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Random utility theory for social choice

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:11.906112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:11.906112Z digest=sha256:26a1aea46af80388cfab1be3991c7bc212cae1aa090ab2877f4b095e22e7dfc1

Observation 7519c8fe-e391-4f82-b614-f4e0aa2e2cd9 · outbound

This paper cites Modeling ordered choices: A primer, 2010.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Modeling ordered choices: A primer, 2010

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:23.658611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:20:11.971612Z digest=sha256:d1b0c80608560a4ac9e2d8f9e4b567f32789804dccf13680d48349572d3253b4

Observation a1d649d2-8446-4572-87bf-486af6d3ef41 · outbound

This paper cites Econometric analysis 4th edition.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Econometric analysis 4th edition

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:23.504855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:20:12.030921Z digest=sha256:c777ca87708c86a22c8c64b54dcec8fe945be4a50fb7a36cd56b4c965aa3a77c

Observation 571facf2-89e0-497e-8702-1fd5b464a369 · outbound

This paper cites Sensory evaluation of food: principles and practices.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Sensory evaluation of food: principles and practices

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:23.282474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:20:12.101142Z digest=sha256:c31a2dda5f12b0dce02e88075bb04b0e40d4629ac618f8ac91ae7316c4ef9d46

Observation 097214b3-cd0f-4089-a3e8-c302ce0d7ae6 · outbound

This paper cites Sensory evaluation techniques.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Sensory evaluation techniques

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:23.069453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:20:12.166328Z digest=sha256:bc97625b46cc150e033ff80e822686f6d6245bcf2368b294e121736828178a10

Observation 0b3b9f97-ef92-479c-bb26-deb82ab8b5de · outbound

This paper cites Nash learning from human feedback.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Nash learning from human feedback

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:22.880574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:20:12.215572Z digest=sha256:470651ae7d8104a48ad7b0dd9f526ff034d3670ae80a3a3c239081a5289b8a79

Observation 9cca57ca-66e0-4d12-bbd7-2c97d1726096 · outbound

This paper cites A general theoretical paradigm to understand learning from human preferences.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function A general theoretical paradigm to understand learning from human preferences

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:22.715947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:20:12.318773Z digest=sha256:764c6d879f872a8d4b58f33ace1f7cec47bdd54dad38365d2d9b1a64480d557a

Observation f44cfb7f-350f-4cc4-ae86-c53b6055920e · outbound

This paper cites Models of human preference for learning reward functions.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Models of human preference for learning reward functions

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:22.557457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:20:12.395294Z digest=sha256:789a49655c825657a3f673bb320c0ca730cdf7658d94993ff26cf4e85bc43f8f

Observation 1120e8b4-c18f-4faa-8d3b-7be8f969d57a · outbound

This paper cites Reinforcement learning from human feedback without reward inference: Model-free algorithm and instance-dependent analysis.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Reinforcement learning from human feedback without reward inference: Model-free algorithm and instance-dependent analysis

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:22.399902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:20:12.569486Z digest=sha256:b4dc16dbfe1e40c03846f28b64b567c733aff54b7d5c0e05813ccc9c991965a9

Observation 16fc529c-6624-49cd-89d7-04165b8a6dc4 · outbound

This paper cites Preference-based online learning with dueling bandits: A survey.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Preference-based online learning with dueling bandits: A survey

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:22.244253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:20:12.737974Z digest=sha256:a3c68429d303eeb2af3736a76b37e166e61dfbb1ebd17444f8d42b0a2ae49bb9

Observation 68c03123-55da-4bc6-a574-bbf61d4e86f5 · outbound

This paper cites Is RLHF More Difficult than Standard RL?.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Is RLHF More Difficult than Standard RL?

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:12.854603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:12.854603Z digest=sha256:2422373b9fd1dd4f22fad7471126c40269c9957975f19cf0b767a56ab8377cc2

Observation 155b393a-2365-4814-8dc9-7b9c21daf1bc · outbound

This paper cites A survey of reinforcement learning from human feedback.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function A survey of reinforcement learning from human feedback

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:12.936227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:12.936227Z digest=sha256:9d228061ce21ec64521aa5966798c9edc22b71299e26808078c7a008a2eadfef

Observation ce96e202-23b6-4bcb-aa22-30374648414c · outbound

This paper cites Scaling laws for reward model overoptimization.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Scaling laws for reward model overoptimization

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:22.106083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:20:12.957795Z digest=sha256:66ddeaa860ca7ab4a6326d9cffef68b40e8e429f974ef37c6f035447fa2c77c4

Observation eb2419d6-f622-4eaa-b050-50fb1789f50c · outbound

This paper cites Model-free preference-based reinforcement learning.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Model-free preference-based reinforcement learning

Reference 34

Resolution
verified exact
doi, observed 2026-08-07T11:20:15.864764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:20:13.050288Z digest=sha256:fb8e40bcbae1f97d7d7c3182eaa94c5edd7df7c87fee48b08edc69ff97405f6f

Observation 852239dc-55ea-4a60-bd60-6d5db5a9e88b · outbound

This paper cites Scaling Language Models: Methods, Analysis & Insights from Training Gopher.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Scaling Language Models: Methods, Analysis & Insights from Training Gopher

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:13.111892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:13.111892Z digest=sha256:17a5d31c32b33eee49288a35d8826c96b01cbedbdb9fb86a077ed345d5d827a3

Observation 47499b1b-0fdc-41be-b7ed-be6849a3b848 · outbound

This paper cites Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:13.164246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:13.164246Z digest=sha256:c468cd78697b1ac4f2ee04055a3991a7045fa1c36a0761499177a300c81b8612

Observation eb008823-5efe-47a6-9bb5-5b36aaab89fe · outbound

This paper cites Iterative data smoothing: Mitigating reward overfitting and overoptimization in RLHF.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Iterative data smoothing: Mitigating reward overfitting and overoptimization in RLHF

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:21.959204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:20:13.204264Z digest=sha256:33a699fc6e27f2ff8b6114a9b244b1e4316c8127309bdc8e3d3aeb67697bd78b

Observation c53d914b-6a9f-429c-97e9-e5b27852a45f · outbound

This paper cites RLHF Workflow: From Reward Modeling to Online RLHF.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function RLHF Workflow: From Reward Modeling to Online RLHF

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:13.243909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:13.243909Z digest=sha256:fe6b174a84b93fe32561ba1d63e0f073429610ecca3a1608cc37bc38da4dee5a

Observation 72383b3d-0aa4-47c7-952e-3f082282bcbf · outbound

This paper cites Iterative preference learning from human feedback: Bridging theory and practice for RLHF under KL -constraint.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Iterative preference learning from human feedback: Bridging theory and practice for RLHF under KL -constraint

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:21.818904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:20:13.281397Z digest=sha256:8bc0c654fb10f5599a026172b107ed3b04d2cfacbf9a06c5c793d09c27d0b824

Observation faf8bccb-9cd4-44ef-92c4-992109603475 · outbound

This paper cites SLiC-HF: Sequence Likelihood Calibration with Human Feedback.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:13.316647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:13.316647Z digest=sha256:4dfc6f2f215d3dc46f82a528a41dc31b8dc8de88ff1b0b0a505431f775c2c508

Observation 4e900513-ea97-4973-99b3-e7ed056381b2 · outbound

This paper cites an unresolved cited work.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:20:21.642253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:20:13.349331Z digest=sha256:5d19edb989e068e0b3cc4762f0ffaa539eda86039b6404dbb74f4c88113e7cb6

Observation fde34b94-7148-45ca-9177-fabbbd68ac2d · outbound

This paper cites Dueling rl: Reinforcement learning with trajectory preferences.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Dueling rl: Reinforcement learning with trajectory preferences

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:21.447430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:20:13.386029Z digest=sha256:3a81ba93e4182c72b083f54c5d671f5f6f78bff2cda9341f6df7fd7cb7ed4d32

Observation d3376d72-f218-4f16-979d-2b5ac0b7f007 · outbound

This paper cites Lee, and Wen Sun.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Lee, and Wen Sun

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:21.240433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:20:13.420179Z digest=sha256:2ed150023e3de8277fde9744eea3b6c3374fdf2256572b76d0059dfbad03a2a7

Observation c27ff57d-d834-4f70-9e04-45b415b4032f · outbound

This paper cites an unresolved cited work.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:20:21.078126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:20:13.440411Z digest=sha256:c5b80e89d04b2a2194f1f2322a219a687f83407597bc988ef7a3fd95b0d61f79

Observation cb357bfe-a7f0-42b9-8cf9-09477257b24f · outbound

This paper cites Principled reinforcement learning with human feedback from pairwise or k-wise comparisons.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Principled reinforcement learning with human feedback from pairwise or k-wise comparisons

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:20.936211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:20:13.473666Z digest=sha256:2ead79e980bd7083f86ad4f28ea3a1cb75c917d431332c7502205fd73786d7aa

Observation c998c1d4-4d99-46c4-bd9a-40bf5731e69c · outbound

This paper cites Provably feedback-efficient reinforcement learning via active reward learning.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Provably feedback-efficient reinforcement learning via active reward learning

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:20.811129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:20:13.509008Z digest=sha256:e9c42b1a598f699370527711166474dce7225c008022830af19703ccfc7735fe

Observation 0233317e-db04-4053-aae1-f918664243a0 · outbound

This paper cites Making RL with preference-based feedback efficient via randomization.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Making RL with preference-based feedback efficient via randomization

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:13.514997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:13.514997Z digest=sha256:839f9ddfb962a44524e2aa4d58e805143d4dff5d796cc69a69ba2f6e461d76bf

Observation c477f12d-8642-45e5-9f7c-6d8ee20d0c30 · outbound

This paper cites Reinforcement Learning with Human Feedback: Learning Dynamic Choices via Pessimism.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Reinforcement Learning with Human Feedback: Learning Dynamic Choices via Pessimism

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:13.572546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:13.572546Z digest=sha256:b3b9227dfeb7ec985b9952b830416c0830fae8b154ebce52c81bcdf05c929bf2

Observation dd8c175d-60e3-4c22-a873-2d522bf430d3 · outbound

This paper cites PARL : A unified framework for policy alignment in reinforcement learning from human feedback.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function PARL : A unified framework for policy alignment in reinforcement learning from human feedback

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:20.698009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:20:13.622186Z digest=sha256:a36988751f1396e61db0d4164fef2448687a9c4182856654edc4e6315e638c43

Observation 5e79b9c9-f0f6-461b-8c8b-9946d45997bc · outbound

This paper cites A Theoretical Framework for Partially Observed Reward-States in RLHF.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function A Theoretical Framework for Partially Observed Reward-States in RLHF

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:13.676261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:13.676261Z digest=sha256:c11e5562c9e5cb1b5283878540b2baee2a1e60023f23d6634cb37a7b0fb234f9

Observation ed9d7dcf-c55c-4cba-b980-2322526f4b1e · outbound

This paper cites Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:13.714921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:13.714921Z digest=sha256:00ccc32f8539afa7e8df469ccf24a7482a02602fed7961b95eb8eecf2453e641

Observation 6690ffa4-8c86-4f9f-92e0-78b78d683504 · outbound

This paper cites Preference-based reinforcement learning with finite-time guarantees.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Preference-based reinforcement learning with finite-time guarantees

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:13.768418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:13.768418Z digest=sha256:0253fa0c1bac0e73f788289915c8fbae4a505bfeed8e3a7716a3943ecdbb6754

Observation a625c456-97f1-4a52-8f94-1a6d6771ea27 · outbound

This paper cites Zeroth-order optimization meets human feedback: Provable learning via ranking oracles.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Zeroth-order optimization meets human feedback: Provable learning via ranking oracles

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:20.474693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:20:13.828580Z digest=sha256:e915dfa927115952420e62efca1751aa14fa9f3f8019f73b35f21181bdf3758d

Observation 7340b7ca-af8f-47b1-826f-1103e648a73a · outbound

This paper cites Interactively optimizing information retrieval systems as a dueling bandits problem.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Interactively optimizing information retrieval systems as a dueling bandits problem

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:13.869262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:13.869262Z digest=sha256:6a3e8981ef8768adf440073965366822dd566a539e293b57c29070c592f6ae99

Observation 37158a94-e721-48dc-9a56-9d4451dae6fc · outbound

This paper cites Beat the mean bandit.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Beat the mean bandit

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:20.339023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:20:13.903719Z digest=sha256:863602c05d12fbc47180e6f6da5763a04a1e6184bb0dbc09518c0d0cfe077771

Observation 36498c12-90fc-466f-a433-0d8d4b0b5411 · outbound

This paper cites Generalized preference optimization: A unified approach to offline alignment.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Generalized preference optimization: A unified approach to offline alignment

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:20.260964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:20:13.939206Z digest=sha256:a9ed60315f0e55a9dad7e47b62316c466e2ec48aa031c8e01a2614ae6f640698

Observation e48c9491-b216-4f2c-8dce-4cfe88ce4f50 · outbound

This paper cites Human-in-the-loop: Provably efficient preference-based reinforcement learning with general function approximation.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Human-in-the-loop: Provably efficient preference-based reinforcement learning with general function approximation

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:20.110017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:20:13.972232Z digest=sha256:1e939056aa6dd4fb964f03c389d3306176c0fab38752a3a7e301672b26267459

Observation 5f9fc27b-a6fa-45d6-aceb-bc102ab9b3c0 · outbound

This paper cites A theoretical analysis of nash learning from human feedback under general kl-regularized preference.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function A theoretical analysis of nash learning from human feedback under general kl-regularized preference

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:20.005951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:20:13.994701Z digest=sha256:783966c49179d690add379598d50373baae37c937344fcf380e17bfd1dcc3bb6

Observation 5cd7e6c0-792b-4999-8f7b-1ef62dfc01ff · outbound

This paper cites Extragradient Preference Optimization (EGPO): Beyond Last-Iterate Convergence for Nash Learning from Human Feedback.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Extragradient Preference Optimization (EGPO): Beyond Last-Iterate Convergence for Nash Learning from Human Feedback

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:14.045330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:14.045330Z digest=sha256:57d041341cf2f266bfdfc4c0b977ba76166f6032396887291957e986011932d6

Observation e7f15eb2-1f01-4a41-b31c-5f4fe31cdcbc · outbound

This paper cites Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:14.109193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:14.109193Z digest=sha256:729859cbe728d14deee32f8998844af7a442973d4f224ad309678d8995f20c1e

Observation 7d609393-a485-48e9-831d-0d2108ca7be1 · outbound

This paper cites Iterative Nash Policy Optimization: Aligning LLMs with General Preferences via No-Regret Learning.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Iterative Nash Policy Optimization: Aligning LLMs with General Preferences via No-Regret Learning

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:14.166568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:14.166568Z digest=sha256:fb8ed7902409733eae06e5155e06d72caf1e502f9af57ebfa245a0893bcc0b4d

Observation d1155787-5804-4de0-aab3-ff4b133c206a · outbound

This paper cites Stochastic first-and zeroth-order methods for nonconvex stochastic programming.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Stochastic first-and zeroth-order methods for nonconvex stochastic programming

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:14.251100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:14.251100Z digest=sha256:8cc03cea19c1a411ee94ea20b02a02fb4291abc9b02baf0f2bb81db5a1626075

Observation b2cbb1d1-4a3a-4b1a-98d2-495f1ea29c8a · outbound

This paper cites Random gradient-free minimization of convex functions.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Random gradient-free minimization of convex functions

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:14.327923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:14.327923Z digest=sha256:d49ea926c543cafe70b4168332b1cd626482800defac8bcd8aa69678f3ac5ef3

Observation 4b2c1b49-9dbc-4dff-9042-fe52767810be · outbound

This paper cites Zeroth-order online alternating direction method of multipliers: Convergence analysis and applications.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Zeroth-order online alternating direction method of multipliers: Convergence analysis and applications

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:19.876255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:20:14.392308Z digest=sha256:d5e04b3f9aa10dd95da24f1bcec45ee484b9da403d2ccab3feb3205694d272f3

Observation 883be79c-b080-49d3-b44d-2ed6fc91247e · outbound

This paper cites Zeroth-order stochastic variance reduction for nonconvex optimization.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Zeroth-order stochastic variance reduction for nonconvex optimization

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:19.744327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:20:14.448455Z digest=sha256:270b2bdc9dcb3c5aec1cab29d3d1179b0e953fe678602136923cb22798bca23b

Observation 971ae2dd-6d2d-4f65-ab50-7b305d4b7b0f · outbound

This paper cites A zeroth-order block coordinate descent algorithm for huge-scale black-box optimization.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function A zeroth-order block coordinate descent algorithm for huge-scale black-box optimization

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:19.522274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:20:14.553872Z digest=sha256:3b990268449d9aac0f2f1c6ef5912973f94286355b4b59fc69186b263f639587

Observation 6aa4c8eb-31e8-4123-8fd8-86f3fa3e4b86 · outbound

This paper cites On the information-adaptive variants of the admm: an iteration complexity perspective.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function On the information-adaptive variants of the admm: an iteration complexity perspective

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:19.280656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:20:14.628579Z digest=sha256:6d81ba6c0f32e82c1d448bd76fd271cabe448e24c0f424b2dfa8d2d525932d40

Observation 32a34c98-7b60-4904-a31b-ac0bc3ab8a24 · outbound

This paper cites Fine-tuning language models with just forward passes.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Fine-tuning language models with just forward passes

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:19.029182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:20:14.679574Z digest=sha256:f31d6a4bf150e957f2066fa42c9079d8b3de916ccead8aba8c785a05c39e5ea0

Observation 07876c0c-5d52-47f5-8453-efbe4022bbcc · outbound

This paper cites Evolutionsstrategie.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Evolutionsstrategie

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:18.777335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:20:14.753674Z digest=sha256:8cde669cbd9031ae0f81038188ecbe2c5b238321d82a0d9967997d1f7e7a1e29

Observation f1de5ea6-2bb9-4c81-a9dd-ca48cfc2ac06 · outbound

This paper cites Evolution Strategies as a Scalable Alternative to Reinforcement Learning.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Evolution Strategies as a Scalable Alternative to Reinforcement Learning

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:14.770317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:14.770317Z digest=sha256:eb1557c35a8102104b12c75a8c503b3fe17a8c78af68b6e3acd063ab979babda

Observation 284ff06e-fe75-4924-ad8d-f9a1aeeae412 · outbound

This paper cites Improving exploration in evolution strategies for deep reinforcement learning via a population of novelty-seeking agents.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Improving exploration in evolution strategies for deep reinforcement learning via a population of novelty-seeking agents

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:18.519056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:20:14.834484Z digest=sha256:3b5b87af84be0e1aa9ede19b2547f92a439b6084370ed01bdc30cd6f4a528d7b

Observation b1c5defc-e852-4e0f-9cd1-8710f494396f · outbound

This paper cites o r \'e nyi, Paul Weng, Weiwei Cheng, and Eyke H \.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function o r \'e nyi, Paul Weng, Weiwei Cheng, and Eyke H \

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:18.225652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:20:14.893600Z digest=sha256:22bf18d52412390d9696a0799a19f9016ecbb942e2b826e690d5b70c34880eb6

Observation 4514c777-95fb-45ff-8643-5349d481cee0 · outbound

This paper cites Preference-based policy learning.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Preference-based policy learning

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:18.013767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:20:14.926887Z digest=sha256:b301e2bf7af025473e8974dcd32512a88319cbdbd6f41dc79e37da09348c04b9

Observation c1086b22-e42f-4c98-8177-ce41ecf25690 · outbound

This paper cites sign SGD : Compressed optimisation for non-convex problems.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function sign SGD : Compressed optimisation for non-convex problems

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:17.771367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:20:14.993052Z digest=sha256:1284ca69b9b069d6408f93728fe0aad534fed51d273a97d4196fe2933b8ebd6b

Observation 4624754b-48f7-4449-9214-3f6e92dee3e3 · outbound

This paper cites sign SGD via zeroth-order oracle.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function sign SGD via zeroth-order oracle

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:17.561389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:20:15.061239Z digest=sha256:491bbf2204e905be310edccab839d2c6bb2afdab812d12d87c846cbcf5ae80ce

Observation 1d23d772-18ad-4127-855e-212208825ea0 · outbound

This paper cites Thurstone.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Thurstone

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:17.425354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:20:15.109276Z digest=sha256:e261e036ede4a1e908d930047fbe22e6c0d79365df92a7acdc2f8e776b547110

Observation 6f98a202-ef5e-4072-ac1b-49f50c2593b3 · outbound

This paper cites an unresolved cited work.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Unresolved cited work

Reference 77

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:20:17.289639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:20:15.174397Z digest=sha256:13d1e937443d2965fa675a9c44ae8b135aa2468bd638144f71c374839486a338

Observation eea2b3cf-b74c-4cc7-b4b3-701a9358fa0b · outbound

This paper cites Reddi, Satyen Kale, and Sanjiv Kumar.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Reddi, Satyen Kale, and Sanjiv Kumar

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:15.228756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:15.228756Z digest=sha256:51c7f3d24d0077aa14c71cd5857d8d78e0ac89eac5c5ffcc8e45ece5112d2935

Observation a778a116-10e9-461d-acc1-169b01716d4a · outbound

This paper cites Online rl in linearly q^ -realizable mdps is as easy as in linear mdps if you learn what to ignore.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Online rl in linearly q^ -realizable mdps is as easy as in linear mdps if you learn what to ignore

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:17.145625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:20:15.262179Z digest=sha256:ec99aa480a164e8842dc9c49e9b6474f130536c016fe1bac3ff75ac74e1a3d13

Observation 35a2cadb-2800-410f-b344-516972a410a6 · outbound

This paper cites Sample-efficient reinforcement learning is feasible for linearly realizable mdps with limited revisiting.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Sample-efficient reinforcement learning is feasible for linearly realizable mdps with limited revisiting

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:16.927021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:20:15.331068Z digest=sha256:18c70a9f3df9d6e46fea768ba69ffaf94cc4d7aa3b14af07b8a16c62f24f935e

Observation ebf7a969-1673-4a56-a423-dfe76fb3dcca · outbound

This paper cites Provably efficient reinforcement learning with linear function approximation.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Provably efficient reinforcement learning with linear function approximation

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:16.706697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:20:15.383213Z digest=sha256:3977f893a94aeab862f036fd0d8d10f153aefd3d90e314aa806c86556d450a5d

Observation c2d9b706-27a8-433f-a24b-df58a46d0be7 · outbound

This paper cites High-dimensional probability: An introduction with applications in data science, volume 47.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function High-dimensional probability: An introduction with applications in data science, volume 47

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:15.443901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:15.443901Z digest=sha256:3e3af46bac9a09baf1285e546e08799c702d17e07782197654c418b426d4e7cd

Observation a1e85407-5e5e-4403-9f38-41b9ca43909b · outbound

This paper cites an unresolved cited work.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Unresolved cited work

Reference 83

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:20:16.605630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:20:15.515429Z digest=sha256:02a3fc251b7b2cf0194eed2274790cbd77ca461c2ed10e8ba937a6e74bcc60e4

Observation 23f9ea8d-a73f-4b8a-a799-0a9968480331 · outbound

This paper cites Direct Language Model Alignment from Online AI Feedback.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Direct Language Model Alignment from Online AI Feedback

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:15.560054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:15.560054Z digest=sha256:ad83d664ca9243238f75777a1b61e11e576b5490d63a15c15c77e444cb1eeb29

Observation 7e5ac7ae-8755-4084-8948-5a91bd1a59c4 · outbound

This paper cites Convex Optimization.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Convex Optimization

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:16.471296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:20:15.630602Z digest=sha256:133363a2127e773f4229ed8c9b165a4b037df8f785e04c1638858ce4ac3a8543

Observation ca26ef71-734b-4f3c-ad42-300c2dcf7080 · outbound

This paper cites Linear convergence of gradient and proximal-gradient methods under the polyak- ojasiewicz condition.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Linear convergence of gradient and proximal-gradient methods under the polyak- ojasiewicz condition

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:16.361093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:20:15.682673Z digest=sha256:1a1d1d6f6f6e759f6769c37c890874596d6bf6c96f2812a70e774b85153c418b

Observation 9d2fc4c2-5b85-45fa-8a99-90eafaa3e14a · outbound

This paper cites On the global convergence rates of softmax policy gradient methods.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function On the global convergence rates of softmax policy gradient methods

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:16.252160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:20:15.714867Z digest=sha256:e487d490f07d8fb5133a1a2a1dc0f74022a7a7835e12fd3e293e75c4b8307359

Pith citing papers

Observation c6d37a12-b035-4679-b920-5dea3a8d1499 · inbound

Efficient Federated RLHF via Zeroth-Order Policy Optimization cites this paper.

Efficient Federated RLHF via Zeroth-Order Policy Optimization Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-21T00:19:49.870688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T04:48:15.394329Z digest=sha256:ba6e58e678f3388af4861cb9045d1c6eb8929bf4b644871c559f312a95ea7c20