Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2010.05848.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-11T20:25:14.589831Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-11T02:33:28.313220Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 63266383-dc18-4525-a337-b70291d8de0c · inbound
Direct Preference Optimization: Your Language Model is Secretly a Reward Model Human-centric Dialog Training via Offline Reinforcement Learning
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation be89228f-fb0b-45fc-b0d1-a2e725cbf520 · inbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Human-centric Dialog Training via Offline Reinforcement Learning
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c13c74ba-0510-45eb-b4da-2f6fa9eccded · inbound
Data Diversification Methods In Alignment Enhance Math Performance In LLMs Human-centric Dialog Training via Offline Reinforcement Learning
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 113730e2-dce4-4225-89cd-f8a6f13ae0e9 · inbound
Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges Human-centric Dialog Training via Offline Reinforcement Learning
Reference 215
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4998f26b-2ced-4ce9-90d1-62de6e372346 · inbound
Wavelet Fourier Diffuser: Frequency-Aware Diffusion Model for Reinforcement Learning Human-centric Dialog Training via Offline Reinforcement Learning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0a5731f-92a9-4aa8-864f-01ef077481da · inbound
Efficient Preference Poisoning Attack on Offline RLHF Human-centric Dialog Training via Offline Reinforcement Learning
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.