Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 13 inbound Pith citation observations for arXiv:2402.10500.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:33:51.144721Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-06-30T11:54:38.395848Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 3c1a0ce5-1a1c-4302-ad9c-3deb7fd487ab · inbound
Reinforcement Learning for Reasoning in Large Language Models with One Training Example Active Preference Optimization for Sample Efficient RLHF
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation fc84104b-b3ff-4fd6-aee8-135c2a4557f2 · inbound
FisherSFT: Data-Efficient Supervised Fine-Tuning of Language Models Using Information Gain Active Preference Optimization for Sample Efficient RLHF
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b56a1c6-51ce-444d-85c0-4266939f0d2e · inbound
Outcome-Based Online Reinforcement Learning: Algorithms and Fundamental Limits Active Preference Optimization for Sample Efficient RLHF
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5182579f-4527-4fea-820f-fafc04ade777 · inbound
Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts Active Preference Optimization for Sample Efficient RLHF
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0eed8347-9013-4e61-8a8c-85cc4566353c · inbound
Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges Active Preference Optimization for Sample Efficient RLHF
Reference 178
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 127f23ca-70cb-44e5-bfda-4e82a1ee6d50 · inbound
Beyond Variance: Prompt-Efficient RLVR via Rare-Event Amplification and Bidirectional Pairing Active Preference Optimization for Sample Efficient RLHF
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 83ad6cf0-5e9d-4a37-9fac-fa443207e2d3 · inbound
Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution Active Preference Optimization for Sample Efficient RLHF
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 654d8ec1-75c3-4deb-af72-177b3e5c065e · inbound
Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution Active Preference Optimization for Sample Efficient RLHF
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d63cda43-1747-4513-9716-71d4a312e267 · inbound
Reinforcement Learning from Human Feedback: A Statistical Perspective Active Preference Optimization for Sample Efficient RLHF
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e6b22369-3ace-4ff6-bdc8-ca9414f4c7a2 · inbound
Wasserstein Distributionally Robust Regret Optimization for Reinforcement Learning from Human Feedback Active Preference Optimization for Sample Efficient RLHF
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 650112f8-d2a3-42a5-aa76-307741b3e8eb · inbound
MASS-DPO: Multi-negative Active Sample Selection for Direct Policy Optimization Active Preference Optimization for Sample Efficient RLHF
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c01721f9-e770-48ed-b2e5-f67d73a59e35 · inbound
Spectral Souping: A Unified Framework for Online Preference Alignment Active Preference Optimization for Sample Efficient RLHF
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 91328711-5186-447a-aa88-0448c01b8788 · inbound
Active Learning for Stochastic Contextual Linear Bandits Active Preference Optimization for Sample Efficient RLHF
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.