Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2405.14655.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:32:30.406644Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-17T12:04:10.473003Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation be71c4ed-012f-4b6b-9f99-37eee45a950b · inbound
Training Language Models to Self-Correct via Reinforcement Learning Multi-turn Reinforcement Learning from Preference Human Feedback
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b51c2a23-0629-41b7-baf2-ebe83f2b8bfd · inbound
Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models Multi-turn Reinforcement Learning from Preference Human Feedback
Reference 128
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 06c03cc1-5efd-4109-ad8b-30f0011fe3dc · inbound
Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Multi-turn Reinforcement Learning from Preference Human Feedback
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be2dfc51-47b6-4556-814c-b0588fa2897e · inbound
PAG: Multi-Turn Reinforced LLM Self-Correction with Policy as Generative Verifier Multi-turn Reinforcement Learning from Preference Human Feedback
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3327ff83-8f88-4819-8567-34e633969dde · inbound
Robust pid sliding mode control for dc servo motor speed control Multi-turn Reinforcement Learning from Preference Human Feedback
Reference 2010
Source-reported events for the cited work
Unavailable: canonical work link unavailable.