Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T16:58:02.706784Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 80 of 80 outbound references and 6 inbound Pith citation observations for arXiv:2501.12735.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T16:58:02.706784Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T22:59:14.448812Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T04:17:37.298243Z
80 of 80 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation ef53fa48-6b45-4b67-8dd4-5df4ab33c1c1 · outbound
Online Preference Alignment for Language Models via Count-based Exploration write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa65da28-31df-4ceb-8871-a66f902761bb · outbound
Online Preference Alignment for Language Models via Count-based Exploration Improved algorithms for linear stochastic bandits
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37ebaddc-f5eb-4bb8-b750-5cc9186a9199 · outbound
Online Preference Alignment for Language Models via Count-based Exploration Reinforcement learning: Theory and algorithms
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 335e907c-2771-49c7-9467-6a51634de33f · outbound
Online Preference Alignment for Language Models via Count-based Exploration Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a205591e-4273-4ad1-9286-0445c7a496d4 · outbound
Online Preference Alignment for Language Models via Count-based Exploration A general theoretical paradigm to understand learning from human preferences
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3db2ce02-fa1f-4a4d-9889-dd930552cb92 · outbound
Online Preference Alignment for Language Models via Count-based Exploration Dynamic bottleneck for robust self-supervised exploration
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 61b7b2fc-bf1c-45b0-83c0-bb113d54659a · outbound
Online Preference Alignment for Language Models via Count-based Exploration Principled exploration via optimistic bootstrapping and backward induction
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3f01d01d-c28a-46fb-acec-1ec4fedcad04 · outbound
Online Preference Alignment for Language Models via Count-based Exploration Pessimistic bootstrapping for uncertainty-driven offline reinforcement learning
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5663f3e6-4208-4386-9523-0f3419e42898 · outbound
Online Preference Alignment for Language Models via Count-based Exploration Pessimistic value iteration for multi-task data sharing in offline reinforcement learning
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1b9ff02c-0539-4d21-b69d-ec8d6295a0cf · outbound
Online Preference Alignment for Language Models via Count-based Exploration Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22139940-ed4e-454f-88a4-306882a4fed1 · outbound
Online Preference Alignment for Language Models via Count-based Exploration Unifying count-based exploration and intrinsic motivation
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0298d3f8-8a7d-45ee-889c-57af64d913f9 · outbound
Online Preference Alignment for Language Models via Count-based Exploration Rank analysis of incomplete block designs: I
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91614892-6328-480f-83ad-5b5073a4cfc9 · outbound
Online Preference Alignment for Language Models via Count-based Exploration Exploration by Random Network Distillation
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5daa008d-20d7-4214-a794-a7d172dd9392 · outbound
Online Preference Alignment for Language Models via Count-based Exploration Value-Incentivized Preference Optimization: A Unified Approach to Online and Offline RLHF
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae3edc00-5f1d-441e-8640-e3d5dd19f24c · outbound
Online Preference Alignment for Language Models via Count-based Exploration Deep reinforcement learning from human preferences
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a83b5501-fa90-47cc-8437-fe75775a3151 · outbound
Online Preference Alignment for Language Models via Count-based Exploration Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73039118-e167-4d2e-9a6a-31bfef77f552 · outbound
Online Preference Alignment for Language Models via Count-based Exploration Training Verifiers to Solve Math Word Problems
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96c70b83-bb3e-4d7d-880a-138e758e79d4 · outbound
Online Preference Alignment for Language Models via Count-based Exploration UltraFeedback: Boosting Language Models with Scaled AI Feedback
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c77a1b4f-c896-445a-8633-bd811307e2fa · outbound
Online Preference Alignment for Language Models via Count-based Exploration RLHF Workflow: From Reward Modeling to Online RLHF
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0353558b-556c-4724-afc8-e0a726a9b717 · outbound
Online Preference Alignment for Language Models via Count-based Exploration The Llama 3 Herd of Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a1c866b-ff66-4914-a6a9-2bf484d45a3b · outbound
Online Preference Alignment for Language Models via Count-based Exploration Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf3c38bb-a4c6-4643-9338-a732163ca67a · outbound
Online Preference Alignment for Language Models via Count-based Exploration Efficient Exploration for LLMs
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90b6bc5f-ed61-4879-a5b2-7aa9de428e58 · outbound
Online Preference Alignment for Language Models via Count-based Exploration Kto: Model alignment as prospect theoretic optimization
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation bf25acec-14a5-45e5-96e7-ffe138407b89 · outbound
Online Preference Alignment for Language Models via Count-based Exploration KTO: Model Alignment as Prospect Theoretic Optimization
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70831cce-6fb2-4675-b1f0-a9132a2859b4 · outbound
Online Preference Alignment for Language Models via Count-based Exploration Scaling laws for reward model overoptimization
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4d4090a-b281-4761-b526-efe9d7227328 · outbound
Online Preference Alignment for Language Models via Count-based Exploration A framework for few-shot language model evaluation, 07 2024
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb455c0a-59fe-442c-8611-7e6a6b9419db · outbound
Online Preference Alignment for Language Models via Count-based Exploration Direct Language Model Alignment from Online AI Feedback
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f617c9d4-9cf5-475a-98cc-be6e812296e7 · outbound
Online Preference Alignment for Language Models via Count-based Exploration Exploration in deep reinforcement learning: From single-agent to multiagent domain
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 77bf73e0-b47e-4b04-abc8-3d37966d4aab · outbound
Online Preference Alignment for Language Models via Count-based Exploration VIME: variational information maximizing exploration
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f3290652-12c8-4769-b09c-31b9aadc1879 · outbound
Online Preference Alignment for Language Models via Count-based Exploration Lo RA : Low-rank adaptation of large language models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2f05f73-cd26-4f7a-a356-b5d5ecccb49e · outbound
Online Preference Alignment for Language Models via Count-based Exploration Unpacking dpo and ppo: Disentangling best practices for learning from preference feedback
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c7e4530f-8b0b-4cb2-83e8-46d35f5a7a31 · outbound
Online Preference Alignment for Language Models via Count-based Exploration LLM-Blender: Ensembling Large Language Models with Pairwise Ranking and Generative Fusion
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation afc77d64-92f3-4511-810d-a71c1d8ed02d · outbound
Online Preference Alignment for Language Models via Count-based Exploration Provably efficient reinforcement learning with linear function approximation
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1f1c2440-5c48-4ebb-8250-3e39785ed575 · outbound
Online Preference Alignment for Language Models via Count-based Exploration Kearns and Satinder P
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation de4c7455-0eea-4f23-876b-64f9db55da2b · outbound
Online Preference Alignment for Language Models via Count-based Exploration RewardBench: Evaluating Reward Models for Language Modeling
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b3b8d5d-bf58-40fc-ae56-2429ed449777 · outbound
Online Preference Alignment for Language Models via Count-based Exploration Bandit algorithms
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dee82772-2af4-4232-8af8-0e9167470206 · outbound
Online Preference Alignment for Language Models via Count-based Exploration A mechanistic understanding of alignment algorithms: A case study on dpo and toxicity
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0fcf4a93-0e7a-4cd1-9a7b-4f9298893c9b · outbound
Online Preference Alignment for Language Models via Count-based Exploration Aligning Large Language Models by On-Policy Self-Judgment
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2708c191-c90d-48ba-9d50-7948c84594fb · outbound
Online Preference Alignment for Language Models via Count-based Exploration TruthfulQA: Measuring How Models Mimic Human Falsehoods
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3cdb3c77-9ce4-4617-987a-4ab43bd70dce · outbound
Online Preference Alignment for Language Models via Count-based Exploration Flipping coins to estimate pseudocounts for exploration in reinforcement learning
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45194ecf-3380-4c65-8ebc-83eff819febe · outbound
Online Preference Alignment for Language Models via Count-based Exploration The sample complexity of exploration in the multi-armed bandit problem
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0aa02464-c6fb-408b-9cb4-382731f3dc7e · outbound
Online Preference Alignment for Language Models via Count-based Exploration SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91a4e0ba-81d1-414d-8fee-113b0b2aebab · outbound
Online Preference Alignment for Language Models via Count-based Exploration Introducing meta llama 3: The most capable openly available llm to date
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 983d111d-4900-4e47-bba8-dc096b0a7fbe · outbound
Online Preference Alignment for Language Models via Count-based Exploration Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3dadfa13-9fd9-4bbf-846e-7ebd4a259c5d · outbound
Online Preference Alignment for Language Models via Count-based Exploration Nash Learning from Human Feedback
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bde05da0-a5d9-4f72-8cf7-223df294b920 · outbound
Online Preference Alignment for Language Models via Count-based Exploration Count-based exploration with neural density models
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0cc0e5f-ebaf-4182-b347-dbc033d14cd2 · outbound
Online Preference Alignment for Language Models via Count-based Exploration Training language models to follow instructions with human feedback
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8995c09-eae0-4c6b-813d-d75ac4a71731 · outbound
Online Preference Alignment for Language Models via Count-based Exploration EQ-Bench: An Emotional Intelligence Benchmark for Large Language Models
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 925ec6a9-62f9-4792-bce0-e83dcb298f1f · outbound
Online Preference Alignment for Language Models via Count-based Exploration Curiosity-driven exploration by self-supervised prediction
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4860ad6f-df33-4618-ad38-765b80d8bf23 · outbound
Online Preference Alignment for Language Models via Count-based Exploration Contrastive ucb: Provably efficient contrastive self-supervised learning in online reinforcement learning
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3f86d261-1d25-49dd-b357-a0df6ebbcbdf · outbound
Online Preference Alignment for Language Models via Count-based Exploration Direct preference optimization: Your language model is secretly a reward model
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc65dccf-c32c-4e79-840d-5f81612232a5 · outbound
Online Preference Alignment for Language Models via Count-based Exploration Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd5effcb-c5e7-486e-9db5-38f54e5ee978 · outbound
Online Preference Alignment for Language Models via Count-based Exploration From $r$ to $Q^*$: Your Language Model is Secretly a Q-Function
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a117eff-091f-4b7d-a6bc-4bb6fe7826a2 · outbound
Online Preference Alignment for Language Models via Count-based Exploration Optimistic exploration even with a pessimistic initialisation
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f3b8509e-4cee-4633-a644-45c1cc7f43ba · outbound
Online Preference Alignment for Language Models via Count-based Exploration Proximal Policy Optimization Algorithms
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4974abf7-e451-4b7c-b3a2-52f85b3ce142 · outbound
Online Preference Alignment for Language Models via Count-based Exploration A Long Way to Go: Investigating Length Correlations in RLHF
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83b75950-a1ae-4a1c-81e5-8a156403d141 · outbound
Online Preference Alignment for Language Models via Count-based Exploration D2PO: Discriminator-Guided DPO with Response Evaluation Models
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3eab34b1-641e-424b-a777-9405db12d272 · outbound
Online Preference Alignment for Language Models via Count-based Exploration An analysis of model-based interval estimation for markov decision processes
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation fd4f4167-2e88-4f61-850d-ab1fbbdf79d0 · outbound
Online Preference Alignment for Language Models via Count-based Exploration A Minimaximalist Approach to Reinforcement Learning from Human Feedback
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8fd4d4e-54de-4202-9681-2119cceb56c2 · outbound
Online Preference Alignment for Language Models via Count-based Exploration \# exploration: A study of count-based exploration for deep reinforcement learning
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5d1e5d9-8723-4de7-97a1-bf97586244a1 · outbound
Online Preference Alignment for Language Models via Count-based Exploration Generalized Preference Optimization: A Unified Approach to Offline Alignment
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28198650-2280-43b9-b673-bd3300888985 · outbound
Online Preference Alignment for Language Models via Count-based Exploration Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11f1b9f5-4ff3-400a-9676-bd2544fa9720 · outbound
Online Preference Alignment for Language Models via Count-based Exploration Zephyr: Direct Distillation of LM Alignment
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90197c55-9f72-42fe-a34a-f099727d8e81 · outbound
Online Preference Alignment for Language Models via Count-based Exploration Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff9d91df-bbb1-4686-bece-ff3fa3c263c4 · outbound
Online Preference Alignment for Language Models via Count-based Exploration Iterative preference learning from human feedback: Bridging theory and practice for rlhf under kl-constraint
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b68010ec-8d77-469d-8b74-cc3561b50960 · outbound
Online Preference Alignment for Language Models via Count-based Exploration Is dpo superior to ppo for llm alignment? a comprehensive study
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a5c652fc-42fe-4a2b-8e5a-f41cf3cf9f65 · outbound
Online Preference Alignment for Language Models via Count-based Exploration Rorl: Robust offline reinforcement learning via conservative smoothing
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64c656ab-1840-42e8-b787-1f0d1cb7924d · outbound
Online Preference Alignment for Language Models via Count-based Exploration Yi: Open Foundation Models by 01.AI
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da437e9d-fa08-415d-8c39-920c81aaccf4 · outbound
Online Preference Alignment for Language Models via Count-based Exploration Regularized Conditional Diffusion Model for Multi-Task Preference Alignment
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 16e79d25-7dbb-489f-ba33-e2376f4b90c8 · outbound
Online Preference Alignment for Language Models via Count-based Exploration Self-Rewarding Language Models
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 158014e7-edc5-41f4-84c0-7e149477fe99 · outbound
Online Preference Alignment for Language Models via Count-based Exploration Preference Aligned Diffusion Planner for Quadrupedal Locomotion Control
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e93129e2-bb24-4e68-b060-75ec78a33bfd · outbound
Online Preference Alignment for Language Models via Count-based Exploration HellaSwag: Can a Machine Really Finish Your Sentence?
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8589d06f-41f9-4b44-9136-39958e26ad64 · outbound
Online Preference Alignment for Language Models via Count-based Exploration Self-Exploring Language Models: Active Preference Elicitation for Online Alignment
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ccd42246-84e3-40b2-ba18-505b2e9a2926 · outbound
Online Preference Alignment for Language Models via Count-based Exploration SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 219b48b0-ef9c-4a92-966c-d08c93b6e255 · outbound
Online Preference Alignment for Language Models via Count-based Exploration Judging llm-as-a-judge with mt-bench and chatbot arena
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1210001-b081-4da7-b0eb-90c876f76849 · outbound
Online Preference Alignment for Language Models via Count-based Exploration Principled reinforcement learning with human feedback from pairwise or k-wise comparisons
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7c8604ef-bd9c-4599-9b57-b61fe2bbadc3 · outbound
Online Preference Alignment for Language Models via Count-based Exploration Fine-Tuning Language Models from Human Preferences
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 335f4395-a3a5-471e-99aa-b46c6a07584b · outbound
Online Preference Alignment for Language Models via Count-based Exploration @esa (Ref
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f144e43-f8d6-4621-b292-9fe9213e1530 · outbound
Online Preference Alignment for Language Models via Count-based Exploration Unresolved cited work
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34a2b941-c2a4-40be-922c-02facace7c73 · outbound
Online Preference Alignment for Language Models via Count-based Exploration a&!Ï "l=-BpEMUs J5ū ?bvtCy O (^rڜH - Y`J* aH/'V t@Ys ;ӓj(u B FBa 竑 6 ^mN OB Y>X 5 >D Q=h .+' A Ի, _|k P(qd/T) nV P C/ۿA+ڽ.W b< ydQx> gQ`U Ԡ < a
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a61d8773-626d-46fe-b3bf-f7fa28628346 · inbound
Outcome-based Exploration for LLM Reasoning Online Preference Alignment for Language Models via Count-based Exploration
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0ebcc62-be89-4ba5-b101-e6d25f04a06e · inbound
CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models Online Preference Alignment for Language Models via Count-based Exploration
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b02aeed5-5e6d-4ec3-a9fd-2532a5403d52 · inbound
Representation-Based Exploration for Language Models: From Test-Time to Post-Training Online Preference Alignment for Language Models via Count-based Exploration
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82d9f26c-c454-4aa4-90cd-cef9da877dcd · inbound
DRIFT: Decoupled Rollouts and Importance-Weighted Fine-Tuning for Efficient Multi-Turn Optimization Online Preference Alignment for Language Models via Count-based Exploration
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7e4c3140-62f9-4b08-b63f-b3e808a90c04 · inbound
On Advantage Estimates for Max@K Policy Gradients Online Preference Alignment for Language Models via Count-based Exploration
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6ae9a656-4640-495b-853d-d45c2a6cb363 · inbound
N-GRPO: Embedding-Level Neighbor Mixing for Enhanced Policy Optimization Online Preference Alignment for Language Models via Count-based Exploration
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.