Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T14:44:15.804741Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 0 inbound Pith citation observations for arXiv:2509.23730.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T14:44:15.804741Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
17 of 17 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 62e5f660-7ad1-4646-bbbe-0adcc5b60b7c · outbound
EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance Ask the Right Questions: Active Question Reformulation with Reinforcement Learning
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2f5499a-556a-4cb9-b502-68a2b81a64ec · outbound
EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25a5375b-e003-47f6-9baa-ed17c4240004 · outbound
EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5584dfa7-3459-46be-a689-c1addf41cc4a · outbound
EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance Tulu 3: Pushing Frontiers in Open Language Model Post-Training
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c27b8e1e-b6de-48bb-b66c-bb51be424721 · outbound
EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance Learning from Peers in Reasoning Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e1918c7-94ee-40ac-9f42-91e812559d95 · outbound
EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance Self-Consistency Improves Chain of Thought Reasoning in Language Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20978142-4479-40fd-9c6a-255df20375c9 · outbound
EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94a462c1-b579-4624-9661-c2d3ac48cf7a · outbound
EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance Conformalized Interactive Imitation Learning: Handling Expert Shift and Intermittent Feedback
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d43ee754-cc80-4702-8db6-99c0c8b8987a · outbound
EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance expert_id
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f63d06b-0b17-4988-93e5-828fe488091a · outbound
EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance Kickstarting Deep Reinforcement Learning
Reference 2015
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42c1ec05-700d-4cb7-896e-f6bbea0b035f · outbound
EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance Reinforcement Learning from LLM Feedback to Counteract Goal Misgeneralization
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85a9474b-8731-4776-94e3-93d4f31e777e · outbound
EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance Mixture-of-Agents Enhances Large Language Model Capabilities
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b595ddec-1e1b-43db-a522-d74a07f306cf · outbound
EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance Countering Reward Over-optimization in LLM with Demonstration-Guided Reinforcement Learning
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef26b841-6810-43db-a517-287c105c63db · outbound
EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance Training Language Models to Self-Correct via Reinforcement Learning
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88d3d7c3-21f4-4525-8edf-18976e896cfc · outbound
EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 638047fa-cab6-415b-9a0d-e9249938123c · outbound
EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance Efficient Active Imitation Learning with Random Network Distillation
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96ef28e7-a304-45d7-97b6-74c9564ba4ed · outbound
EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance Full Parameter Fine-tuning for Large Language Models with Limited Resources
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.