Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2203.10050.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:13:00.332041Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T19:40:06.249099Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 35e6f199-18ce-4b1f-b74f-06ac3a279846 · inbound
The Unreasonable Effectiveness of Entropy Minimization in LLM Reasoning SURF: Semi-supervised Reward Learning with Data Augmentation for Feedback-efficient Preference-based Reinforcement Learning
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cba3728a-2163-4aaa-8eee-1e95afce2404 · inbound
CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries SURF: Semi-supervised Reward Learning with Data Augmentation for Feedback-efficient Preference-based Reinforcement Learning
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0ed7339-82d8-44be-a7b7-a506a191fd40 · inbound
PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning SURF: Semi-supervised Reward Learning with Data Augmentation for Feedback-efficient Preference-based Reinforcement Learning
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60a89451-5820-4f34-b2dd-015d8b93da35 · inbound
SENIOR: Efficient Query Selection and Preference-Guided Exploration in Preference-based Reinforcement Learning SURF: Semi-supervised Reward Learning with Data Augmentation for Feedback-efficient Preference-based Reinforcement Learning
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a33cc633-ab9e-40c0-8d3f-d278959f06ca · inbound
Residual Reward Models for Preference-based Reinforcement Learning SURF: Semi-supervised Reward Learning with Data Augmentation for Feedback-efficient Preference-based Reinforcement Learning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a7a11f4-574d-433d-a85e-a8c1ccd9e8f5 · inbound
Learning Process Rewards via Success Visitation Matching for Efficient RL SURF: Semi-supervised Reward Learning with Data Augmentation for Feedback-efficient Preference-based Reinforcement Learning
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a642df2a-df81-4add-97a5-b3b176611927 · inbound
Themis: An explainable AI-enabled framework for Reinforcement Learning with Human Feedback SURF: Semi-supervised Reward Learning with Data Augmentation for Feedback-efficient Preference-based Reinforcement Learning
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 45516930-8444-476e-9e21-69b2c7f7a3fa · inbound
MAPL: Multi-Objective Preference Learning for Robot Locomotion SURF: Semi-supervised Reward Learning with Data Augmentation for Feedback-efficient Preference-based Reinforcement Learning
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.