Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2310.00212.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:16:35.711109Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T04:47:38.261871Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 0822a513-5769-4aa4-b39b-653923605c68 · inbound
ORPO: Monolithic Preference Optimization without Reference Model Pairwise Proximal Policy Optimization: Harnessing Relative Feedback for LLM Alignment
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d363611e-b4c9-47ae-8414-9e34bfd6599a · inbound
Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey Pairwise Proximal Policy Optimization: Harnessing Relative Feedback for LLM Alignment
Reference 162
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cb2943a0-b525-4bb0-9e86-0a48b51ea40b · inbound
Reinforcement Learning from Human Feedback Pairwise Proximal Policy Optimization: Harnessing Relative Feedback for LLM Alignment
Reference 140
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7f645668-9c8f-4882-b23b-0fe5ee70045a · inbound
Toward Scientific Reasoning in LLMs: Training from Expert Discussions via Reinforcement Learning Pairwise Proximal Policy Optimization: Harnessing Relative Feedback for LLM Alignment
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b7fd5a2-c199-4aaf-b7fb-24f6c33f8a3c · inbound
AlignCultura: Towards Culturally Aligned Large Language Models? Pairwise Proximal Policy Optimization: Harnessing Relative Feedback for LLM Alignment
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7d2c28ee-e210-4f24-9bea-aaea8b2cfe68 · inbound
Response Time Enhances Alignment with Heterogeneous Preferences Pairwise Proximal Policy Optimization: Harnessing Relative Feedback for LLM Alignment
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 62cb8b9e-2542-4a54-aa8b-506b9a52933f · inbound
Baseline-Free Policy Optimization for Neural Combinatorial Optimization Pairwise Proximal Policy Optimization: Harnessing Relative Feedback for LLM Alignment
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3e8c120f-eb0a-4277-be2c-af50c8101191 · inbound
Representation-Aware Advantage Estimation: Your Reward Model Provides More Than A Scalar Output Pairwise Proximal Policy Optimization: Harnessing Relative Feedback for LLM Alignment
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.