Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:20:15.714867Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 87 of 87 outbound references and 1 inbound Pith citation observation for arXiv:2506.03066.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:20:15.714867Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-10T04:48:15.394329Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-10T11:35:19.307998Z
87 of 87 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation e9703832-2747-4b73-8c47-4467da086c5e · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Reinforcement learning: An introduction, volume 1
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70fe8ede-870e-4e44-ae03-67b1798c8153 · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Controlled experiments on the web: survey and practical guide
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2de9bc8e-27eb-4749-95be-5ba5d2ac8ec5 · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Deep reinforcement learning from human preferences
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2916cde9-d8b5-465c-97bc-8968f3927110 · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Al Sallab, Senthil Yogamani, and Patrick Pérez
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 739bf0c4-c7c9-4cd1-bc45-3be231cf17ea · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Training language models to follow instructions with human feedback
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca01d7dd-fa85-480d-9975-8a0b310cb5f4 · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99c7fe35-9628-4b78-b4ab-abb6f85799f3 · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Dynamic programming and stochastic control processes
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 192aba53-7d83-4907-ad1b-7861438a4a85 · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Markov decision processes: discrete stochastic dynamic programming
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d6406f9-6fc8-452a-874a-8d9f581f79e9 · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Inverse reward design
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de54f59e-6712-4905-8a83-7fe6714d6b9d · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Reward Design with Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f2ccb55-f03d-4c47-9196-aa7b47ef13c6 · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Defining and characterizing reward gaming
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c50e7a09-cd8c-4f95-834e-d184db24cd46 · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31d800c5-7bad-4ed5-8efc-6e2d80ecf082 · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function GPT-4o System Card
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1534fe5b-74b5-4f5f-a2bb-9c1b5495e071 · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Direct preference optimization: Your language model is secretly a reward model
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5c0ff52-d75e-47a2-bb45-e07577dc1094 · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Zeroth-order policy gradient for reinforcement learning from human feedback without reward inference
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5def0db4-fbcf-475d-b4d3-3d100ea593d7 · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Proximal policy optimization algorithms, 2017
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a80d8297-66c6-4289-8a23-988584021b24 · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Open problems and fundamental limitations of reinforcement learning from human feedback
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2b6d06d7-162f-4354-abc4-664102270b38 · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Lee, Wotao Yin, Mingyi Hong, Zhangyang Wang, Sijia Liu, and Tianlong Chen
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4656b496-a7a3-473d-9590-2e947e847154 · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function From r to q^ * : Your language model is secretly a q-function
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b76e2247-cc93-4d0b-841a-15c87dfc1c31 · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Unresolved cited work
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2105833b-782a-4fa4-b5d9-68dbaa0f8459 · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Random utility theory for social choice
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7519c8fe-e391-4f82-b614-f4e0aa2e2cd9 · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Modeling ordered choices: A primer, 2010
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a1d649d2-8446-4572-87bf-486af6d3ef41 · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Econometric analysis 4th edition
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 571facf2-89e0-497e-8702-1fd5b464a369 · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Sensory evaluation of food: principles and practices
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 097214b3-cd0f-4089-a3e8-c302ce0d7ae6 · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Sensory evaluation techniques
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0b3b9f97-ef92-479c-bb26-deb82ab8b5de · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Nash learning from human feedback
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9cca57ca-66e0-4d12-bbd7-2c97d1726096 · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function A general theoretical paradigm to understand learning from human preferences
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f44cfb7f-350f-4cc4-ae86-c53b6055920e · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Models of human preference for learning reward functions
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1120e8b4-c18f-4faa-8d3b-7be8f969d57a · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Reinforcement learning from human feedback without reward inference: Model-free algorithm and instance-dependent analysis
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 16fc529c-6624-49cd-89d7-04165b8a6dc4 · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Preference-based online learning with dueling bandits: A survey
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 68c03123-55da-4bc6-a574-bbf61d4e86f5 · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Is RLHF More Difficult than Standard RL?
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 155b393a-2365-4814-8dc9-7b9c21daf1bc · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function A survey of reinforcement learning from human feedback
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce96e202-23b6-4bcb-aa22-30374648414c · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Scaling laws for reward model overoptimization
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation eb2419d6-f622-4eaa-b050-50fb1789f50c · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Model-free preference-based reinforcement learning
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 852239dc-55ea-4a60-bd60-6d5db5a9e88b · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Scaling Language Models: Methods, Analysis & Insights from Training Gopher
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47499b1b-0fdc-41be-b7ed-be6849a3b848 · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb008823-5efe-47a6-9bb5-5b36aaab89fe · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Iterative data smoothing: Mitigating reward overfitting and overoptimization in RLHF
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c53d914b-6a9f-429c-97e9-e5b27852a45f · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function RLHF Workflow: From Reward Modeling to Online RLHF
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72383b3d-0aa4-47c7-952e-3f082282bcbf · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Iterative preference learning from human feedback: Bridging theory and practice for RLHF under KL -constraint
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation faf8bccb-9cd4-44ef-92c4-992109603475 · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e900513-ea97-4973-99b3-e7ed056381b2 · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Unresolved cited work
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fde34b94-7148-45ca-9177-fabbbd68ac2d · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Dueling rl: Reinforcement learning with trajectory preferences
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d3376d72-f218-4f16-979d-2b5ac0b7f007 · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Lee, and Wen Sun
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c27ff57d-d834-4f70-9e04-45b415b4032f · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Unresolved cited work
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cb357bfe-a7f0-42b9-8cf9-09477257b24f · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Principled reinforcement learning with human feedback from pairwise or k-wise comparisons
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c998c1d4-4d99-46c4-bd9a-40bf5731e69c · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Provably feedback-efficient reinforcement learning via active reward learning
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0233317e-db04-4053-aae1-f918664243a0 · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Making RL with preference-based feedback efficient via randomization
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c477f12d-8642-45e5-9f7c-6d8ee20d0c30 · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Reinforcement Learning with Human Feedback: Learning Dynamic Choices via Pessimism
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd8c175d-60e3-4c22-a873-2d522bf430d3 · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function PARL : A unified framework for policy alignment in reinforcement learning from human feedback
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5e79b9c9-f0f6-461b-8c8b-9946d45997bc · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function A Theoretical Framework for Partially Observed Reward-States in RLHF
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed9d7dcf-c55c-4cba-b980-2322526f4b1e · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6690ffa4-8c86-4f9f-92e0-78b78d683504 · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Preference-based reinforcement learning with finite-time guarantees
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a625c456-97f1-4a52-8f94-1a6d6771ea27 · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Zeroth-order optimization meets human feedback: Provable learning via ranking oracles
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7340b7ca-af8f-47b1-826f-1103e648a73a · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Interactively optimizing information retrieval systems as a dueling bandits problem
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37158a94-e721-48dc-9a56-9d4451dae6fc · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Beat the mean bandit
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 36498c12-90fc-466f-a433-0d8d4b0b5411 · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Generalized preference optimization: A unified approach to offline alignment
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e48c9491-b216-4f2c-8dce-4cfe88ce4f50 · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Human-in-the-loop: Provably efficient preference-based reinforcement learning with general function approximation
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5f9fc27b-a6fa-45d6-aceb-bc102ab9b3c0 · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function A theoretical analysis of nash learning from human feedback under general kl-regularized preference
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5cd7e6c0-792b-4999-8f7b-1ef62dfc01ff · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Extragradient Preference Optimization (EGPO): Beyond Last-Iterate Convergence for Nash Learning from Human Feedback
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7f15eb2-1f01-4a41-b31c-5f4fe31cdcbc · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d609393-a485-48e9-831d-0d2108ca7be1 · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Iterative Nash Policy Optimization: Aligning LLMs with General Preferences via No-Regret Learning
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1155787-5804-4de0-aab3-ff4b133c206a · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Stochastic first-and zeroth-order methods for nonconvex stochastic programming
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2cbb1d1-4a3a-4b1a-98d2-495f1ea29c8a · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Random gradient-free minimization of convex functions
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b2c1b49-9dbc-4dff-9042-fe52767810be · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Zeroth-order online alternating direction method of multipliers: Convergence analysis and applications
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 883be79c-b080-49d3-b44d-2ed6fc91247e · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Zeroth-order stochastic variance reduction for nonconvex optimization
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 971ae2dd-6d2d-4f65-ab50-7b305d4b7b0f · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function A zeroth-order block coordinate descent algorithm for huge-scale black-box optimization
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6aa4c8eb-31e8-4123-8fd8-86f3fa3e4b86 · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function On the information-adaptive variants of the admm: an iteration complexity perspective
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 32a34c98-7b60-4904-a31b-ac0bc3ab8a24 · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Fine-tuning language models with just forward passes
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 07876c0c-5d52-47f5-8453-efbe4022bbcc · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Evolutionsstrategie
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f1de5ea6-2bb9-4c81-a9dd-ca48cfc2ac06 · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Evolution Strategies as a Scalable Alternative to Reinforcement Learning
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 284ff06e-fe75-4924-ad8d-f9a1aeeae412 · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Improving exploration in evolution strategies for deep reinforcement learning via a population of novelty-seeking agents
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b1c5defc-e852-4e0f-9cd1-8710f494396f · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function o r \'e nyi, Paul Weng, Weiwei Cheng, and Eyke H \
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4514c777-95fb-45ff-8643-5349d481cee0 · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Preference-based policy learning
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c1086b22-e42f-4c98-8177-ce41ecf25690 · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function sign SGD : Compressed optimisation for non-convex problems
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4624754b-48f7-4449-9214-3f6e92dee3e3 · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function sign SGD via zeroth-order oracle
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1d23d772-18ad-4127-855e-212208825ea0 · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Thurstone
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6f98a202-ef5e-4072-ac1b-49f50c2593b3 · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Unresolved cited work
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation eea2b3cf-b74c-4cc7-b4b3-701a9358fa0b · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Reddi, Satyen Kale, and Sanjiv Kumar
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a778a116-10e9-461d-acc1-169b01716d4a · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Online rl in linearly q^ -realizable mdps is as easy as in linear mdps if you learn what to ignore
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 35a2cadb-2800-410f-b344-516972a410a6 · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Sample-efficient reinforcement learning is feasible for linearly realizable mdps with limited revisiting
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ebf7a969-1673-4a56-a423-dfe76fb3dcca · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Provably efficient reinforcement learning with linear function approximation
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c2d9b706-27a8-433f-a24b-df58a46d0be7 · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function High-dimensional probability: An introduction with applications in data science, volume 47
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1e85407-5e5e-4403-9f38-41b9ca43909b · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Unresolved cited work
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 23f9ea8d-a73f-4b8a-a799-0a9968480331 · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Direct Language Model Alignment from Online AI Feedback
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e5ac7ae-8755-4084-8948-5a91bd1a59c4 · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Convex Optimization
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ca26ef71-734b-4f3c-ad42-300c2dcf7080 · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Linear convergence of gradient and proximal-gradient methods under the polyak- ojasiewicz condition
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9d2fc4c2-5b85-45fa-8a99-90eafaa3e14a · outbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function On the global convergence rates of softmax policy gradient methods
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c6d37a12-b035-4679-b920-5dea3a8d1499 · inbound
Efficient Federated RLHF via Zeroth-Order Policy Optimization Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.