Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:45:01.725671Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 0 inbound Pith citation observations for arXiv:2505.17714.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:45:01.725671Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
20 of 20 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation ed48a1ba-0985-4002-8884-99638701bedf · outbound
PPO-BR: Dual-Signal Entropy-Reward Adaptation for Trust Region Policy Optimization Human-level control through deep reinforcement learning,
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba071adb-7693-41f6-b689-43307b421dc9 · outbound
PPO-BR: Dual-Signal Entropy-Reward Adaptation for Trust Region Policy Optimization Benchmarking deep reinforcement learning for continuous control,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9fa4db05-1286-48d5-982b-b0542b980360 · outbound
PPO-BR: Dual-Signal Entropy-Reward Adaptation for Trust Region Policy Optimization Proximal Policy Optimization Algorithms
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aaea74fd-0dfa-4cf0-80f6-8ce220d945b5 · outbound
PPO-BR: Dual-Signal Entropy-Reward Adaptation for Trust Region Policy Optimization Discriminator-actor-critic: Addressing sample inefficiency and reward bias in adversarial imitation learning,
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation af84fa71-8172-418a-a21f-51ccf5973597 · outbound
PPO-BR: Dual-Signal Entropy-Reward Adaptation for Trust Region Policy Optimization Trust region policy optimization,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 1ea400c7-b34f-496a-8f83-5b92ddac29bc · outbound
PPO-BR: Dual-Signal Entropy-Reward Adaptation for Trust Region Policy Optimization Annealed policy optimization for deep reinforcement learning,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e826bed6-4128-40b3-8976-794773bad0f1 · outbound
PPO-BR: Dual-Signal Entropy-Reward Adaptation for Trust Region Policy Optimization Understanding the impact of entropy on policy optimization,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9f4fba45-8f02-4dd0-a811-d838d83d1529 · outbound
PPO-BR: Dual-Signal Entropy-Reward Adaptation for Trust Region Policy Optimization Normalized policy gradients for reinforcement learning,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d435a8a6-a97a-4815-8819-1cf31dee32e2 · outbound
PPO-BR: Dual-Signal Entropy-Reward Adaptation for Trust Region Policy Optimization Sutton and A
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation db774f95-8495-4d18-87a8-156d86dc4837 · outbound
PPO-BR: Dual-Signal Entropy-Reward Adaptation for Trust Region Policy Optimization Reinforcement learning: A survey,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 49f59f67-a684-43b3-9e07-d88295652ad2 · outbound
PPO-BR: Dual-Signal Entropy-Reward Adaptation for Trust Region Policy Optimization Simple statistical gradient-following algorithms for connectionist reinforcement learning,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ce536db6-f0aa-47ec-bbae-acfece1719f7 · outbound
PPO-BR: Dual-Signal Entropy-Reward Adaptation for Trust Region Policy Optimization High-dimensional continuous control using generalized advantage estimation,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 350014f5-19c8-4dc2-b356-18a77b7a37c4 · outbound
PPO-BR: Dual-Signal Entropy-Reward Adaptation for Trust Region Policy Optimization Grandmaster level in StarCraft II using multi-agent reinforcement learning,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 7dc67ac9-c92c-44ff-88af-46b978cdf3e7 · outbound
PPO-BR: Dual-Signal Entropy-Reward Adaptation for Trust Region Policy Optimization Annealed policy optimization,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f77c92cb-f40a-42c5-b9bd-a3ef5461038d · outbound
PPO-BR: Dual-Signal Entropy-Reward Adaptation for Trust Region Policy Optimization Noisy networks for exploration,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 725dfd50-53b0-4a33-abe4-67a6f5dbd49b · outbound
PPO-BR: Dual-Signal Entropy-Reward Adaptation for Trust Region Policy Optimization Deep Reinforcement Learning: An Overview,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5a5b95cf-715c-4aaf-8a2d-a4aa84a5f2d4 · outbound
PPO-BR: Dual-Signal Entropy-Reward Adaptation for Trust Region Policy Optimization Constrained Policy Optimization,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 6e852697-3be3-4650-a1f9-4a74d4c7c120 · outbound
PPO-BR: Dual-Signal Entropy-Reward Adaptation for Trust Region Policy Optimization A Study on Overfitting in Deep Reinforcement Learning,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b155d807-98c2-4b2d-967a-679e0e6b73db · outbound
PPO-BR: Dual-Signal Entropy-Reward Adaptation for Trust Region Policy Optimization Optimizing Customer Satisfaction Through Sentiment Analysis: A BERT-Based Machine Learning Approach to Extract Insights,
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec1f8a1a-a535-4f03-9ffd-7d932f0ab305 · outbound
PPO-BR: Dual-Signal Entropy-Reward Adaptation for Trust Region Policy Optimization Context-Aware Semantic Segmentation: Enhancing Pixel-Level Understanding with Large Language Models for Advanced Vision Applications
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.