Pith. sign in

Paper Citation Record · LEDGER

EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance

As of 13 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 0 inbound Pith citation observations for arXiv:2509.23730.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.23730 v2

Coverage vector

measured 17 of 17 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T14:44:15.804741Z

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

17 of 17 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 62e5f660-7ad1-4646-bbbe-0adcc5b60b7c · outbound

This paper cites Ask the Right Questions: Active Question Reformulation with Reinforcement Learning.

EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance Ask the Right Questions: Active Question Reformulation with Reinforcement Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T14:44:13.454741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:44:13.454741Z digest=sha256:871e2421bdf8f00a67e8dc697810d78dbc747909f04c8da20a98961a7aa906e3

Observation a2f5499a-556a-4cb9-b502-68a2b81a64ec · outbound

This paper cites Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback.

EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T14:44:13.604742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:44:13.604742Z digest=sha256:e8d59f358965f73d367bfb77fb99ecad67f24920f53db7875b1babd2badb6c82

Observation 25a5375b-e003-47f6-9baa-ed17c4240004 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T14:44:13.895720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:44:13.895720Z digest=sha256:78df95d899f356f13d22eabafcd34944ee9bd0bf21d79e09286eafd2f077ecef

Observation 5584dfa7-3459-46be-a689-c1addf41cc4a · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T14:44:14.374747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:44:14.374747Z digest=sha256:0e2635facb8068779064c7d2552d2df2d4aa58871b5bf1b2af079fa4e83a1c1d

Observation c27b8e1e-b6de-48bb-b66c-bb51be424721 · outbound

This paper cites Learning from Peers in Reasoning Models.

EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance Learning from Peers in Reasoning Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T14:44:14.634746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:44:14.634746Z digest=sha256:57f1f63c5130ecfb9592abbbcc656fd35c9ded5f6703a66b491df7f940d0bbd3

Observation 8e1918c7-94ee-40ac-9f42-91e812559d95 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T14:44:15.364737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:44:15.364737Z digest=sha256:371ec28724fbc2584d253225674c1e55c75d9a7cbb5cb9bfacd2f4bf4f6274be

Observation 20978142-4479-40fd-9c6a-255df20375c9 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T14:44:15.514854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:44:15.514854Z digest=sha256:2edfbc387ab0230c5f470d4d01029c371f1722b046ff4ce414290ee7f3e18b09

Observation 94a462c1-b579-4624-9661-c2d3ac48cf7a · outbound

This paper cites Conformalized Interactive Imitation Learning: Handling Expert Shift and Intermittent Feedback.

EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance Conformalized Interactive Imitation Learning: Handling Expert Shift and Intermittent Feedback

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T14:44:15.684741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:44:15.684741Z digest=sha256:2dc8851e70aeb08e60bfa4808b7b91ac50654a4c978ff2c67c4a511bf688a0b9

Observation d43ee754-cc80-4702-8db6-99c0c8b8987a · outbound

This paper cites expert_id.

EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance expert_id

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T14:44:15.804741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:44:15.804741Z digest=sha256:b09ab004e334c473037b6cd805646cb7e37efae90a83ab2b1073c87d7d4380f1

Observation 1f63d06b-0b17-4988-93e5-828fe488091a · outbound

This paper cites Kickstarting Deep Reinforcement Learning.

EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance Kickstarting Deep Reinforcement Learning

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-04T14:44:15.083396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:44:15.083396Z digest=sha256:90e5e36a823d05bf9fa553116a5234eb45d2fccd71c67eedbba9c96204ae8310

Observation 42c1ec05-700d-4cb7-896e-f6bbea0b035f · outbound

This paper cites Reinforcement Learning from LLM Feedback to Counteract Goal Misgeneralization.

EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance Reinforcement Learning from LLM Feedback to Counteract Goal Misgeneralization

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-04T14:44:13.214742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:44:13.214742Z digest=sha256:b47213678f63c4576bc93c087b3dee74f53965cac5fff4cbd9933400682b54d5

Observation 85a9474b-8731-4776-94e3-93d4f31e777e · outbound

This paper cites Mixture-of-Agents Enhances Large Language Model Capabilities.

EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance Mixture-of-Agents Enhances Large Language Model Capabilities

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-04T14:44:15.199364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:44:15.199364Z digest=sha256:1736141d98d78d0118ccfc2f18cf48d2071c7db75a65e28c33aa68971f1f029d

Observation b595ddec-1e1b-43db-a522-d74a07f306cf · outbound

This paper cites Countering Reward Over-optimization in LLM with Demonstration-Guided Reinforcement Learning.

EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance Countering Reward Over-optimization in LLM with Demonstration-Guided Reinforcement Learning

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-04T14:44:14.954740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:44:14.954740Z digest=sha256:5097ec89afd537801922349b56c39cd4f09511101340c36e05f3f1b5d39f79e0

Observation ef26b841-6810-43db-a517-287c105c63db · outbound

This paper cites Training Language Models to Self-Correct via Reinforcement Learning.

EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance Training Language Models to Self-Correct via Reinforcement Learning

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-04T14:44:14.079997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:44:14.079997Z digest=sha256:74f687287721fa0c17a50e714817521022b9c3073280f17239bd5cb402168afa

Observation 88d3d7c3-21f4-4525-8edf-18976e896cfc · outbound

This paper cites CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing.

EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T14:44:13.758213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:44:13.758213Z digest=sha256:ef86b658a643c2cd54b9048268d980623a54871d62b791302616f79a79cfdd7e

Observation 638047fa-cab6-415b-9a0d-e9249938123c · outbound

This paper cites Efficient Active Imitation Learning with Random Network Distillation.

EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance Efficient Active Imitation Learning with Random Network Distillation

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T14:44:13.324744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:44:13.324744Z digest=sha256:e7aea807de43701195fe7d8be1b127c04bf686de7a7a03ea0e62d20f6936c9fc

Observation 96ef28e7-a304-45d7-97b6-74c9564ba4ed · outbound

This paper cites Full Parameter Fine-tuning for Large Language Models with Limited Resources.

EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance Full Parameter Fine-tuning for Large Language Models with Limited Resources

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T14:44:14.782933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:44:14.782933Z digest=sha256:76d806e7119eff42d956d7cccb4e468fb4badecb34f143843a06c6ed188c5b10

Pith citing papers

No inbound Pith citation observations are available.