Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2409.17401.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-10T21:41:39.597198Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T17:08:43.686216Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 70bfd106-b668-4668-8b44-1f90a18f919f · inbound
ElasticZO: A Memory-Efficient On-Device Learning with Combined Zeroth- and First-Order Optimization Zeroth-Order Policy Gradient for Reinforcement Learning from Human Feedback without Reward Inference
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd4d33d2-afbe-4190-931c-843e0bbb0e49 · inbound
Distributed primal-dual algorithm for constrained multi-agent reinforcement learning under coupled policies Zeroth-Order Policy Gradient for Reinforcement Learning from Human Feedback without Reward Inference
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6aa46c19-5430-4ba7-a8ed-971f3a9aba5a · inbound
AI-Driven Stabilization in Power Grids through Controlling Line Admittances Zeroth-Order Policy Gradient for Reinforcement Learning from Human Feedback without Reward Inference
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1fa882c-a460-4bbc-aeb4-701505604e1d · inbound
Policy Gradient Primal-Dual Method for Safe Reinforcement Learning from Human Feedback Zeroth-Order Policy Gradient for Reinforcement Learning from Human Feedback without Reward Inference
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation bc84b6c3-59a8-48a0-8caa-ad1fa77ff559 · inbound
Distributed Zeroth-Order Policy Gradient for Networked Multi-agent Reinforcement Learning from Human Feedback Zeroth-Order Policy Gradient for Reinforcement Learning from Human Feedback without Reward Inference
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation dbbfb27b-d55a-4fb3-971b-153a94fcef5c · inbound
Zero-order Parameter-free Optimization for LMO-based Methods: Novel Approach for Efficient Fine-tuning Zeroth-Order Policy Gradient for Reinforcement Learning from Human Feedback without Reward Inference
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.