Pith. sign in

Paper Citation Record · LEDGER

PPO-BR: Dual-Signal Entropy-Reward Adaptation for Trust Region Policy Optimization

As of 19 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 0 inbound Pith citation observations for arXiv:2505.17714.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.17714 v1

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:45:01.725671Z

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

20 of 20 outbound references displayed

  • verified exact0
  • verified fuzzy15
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ed48a1ba-0985-4002-8884-99638701bedf · outbound

This paper cites Human-level control through deep reinforcement learning,.

PPO-BR: Dual-Signal Entropy-Reward Adaptation for Trust Region Policy Optimization Human-level control through deep reinforcement learning,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:44:59.795938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:44:59.795938Z digest=sha256:d14eb710023801c3e504fde5ce476bbf7a01cbecef159f8673925fad37a8bd12

Observation ba071adb-7693-41f6-b689-43307b421dc9 · outbound

This paper cites Benchmarking deep reinforcement learning for continuous control,.

PPO-BR: Dual-Signal Entropy-Reward Adaptation for Trust Region Policy Optimization Benchmarking deep reinforcement learning for continuous control,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:45:06.440862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:44:59.896292Z digest=sha256:b19e50d5cdad8f83a4f45a7d746e976f033ffdf87a67de9fb30b34accc2367e4

Observation 9fa4db05-1286-48d5-982b-b0542b980360 · outbound

This paper cites Proximal Policy Optimization Algorithms.

PPO-BR: Dual-Signal Entropy-Reward Adaptation for Trust Region Policy Optimization Proximal Policy Optimization Algorithms

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:45:00.026191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:45:00.026191Z digest=sha256:01c4386ffce90aa8086146e6463952dd0dbe2710a4dfeed6c47fd8adf629a816

Observation aaea74fd-0dfa-4cf0-80f6-8ce220d945b5 · outbound

This paper cites Discriminator-actor-critic: Addressing sample inefficiency and reward bias in adversarial imitation learning,.

PPO-BR: Dual-Signal Entropy-Reward Adaptation for Trust Region Policy Optimization Discriminator-actor-critic: Addressing sample inefficiency and reward bias in adversarial imitation learning,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:45:06.094173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:45:00.118894Z digest=sha256:342a74e0262b6179a2718ba881cfb8ab0db98a8420914c9e04ae593c381c86b4

Observation af84fa71-8172-418a-a21f-51ccf5973597 · outbound

This paper cites Trust region policy optimization,.

PPO-BR: Dual-Signal Entropy-Reward Adaptation for Trust Region Policy Optimization Trust region policy optimization,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:45:05.758067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:45:00.173993Z digest=sha256:3a4c795a38db43466b3eb5f3f2899fe807e7ddcacc0df14f936bc81c4fbb486b

Observation 1ea400c7-b34f-496a-8f83-5b92ddac29bc · outbound

This paper cites Annealed policy optimization for deep reinforcement learning,.

PPO-BR: Dual-Signal Entropy-Reward Adaptation for Trust Region Policy Optimization Annealed policy optimization for deep reinforcement learning,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:45:05.421447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:45:00.226610Z digest=sha256:81c481f2e161aa267ac71c4f8897d89ad8698da4016f0b4bdb074b59010867f9

Observation e826bed6-4128-40b3-8976-794773bad0f1 · outbound

This paper cites Understanding the impact of entropy on policy optimization,.

PPO-BR: Dual-Signal Entropy-Reward Adaptation for Trust Region Policy Optimization Understanding the impact of entropy on policy optimization,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:45:05.066997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:45:00.328576Z digest=sha256:18903ecea0bbfcf1cb09e6c3fa00837e2ff2aa66543e1de6a972a0a54c38a841

Observation 9f4fba45-8f02-4dd0-a811-d838d83d1529 · outbound

This paper cites Normalized policy gradients for reinforcement learning,.

PPO-BR: Dual-Signal Entropy-Reward Adaptation for Trust Region Policy Optimization Normalized policy gradients for reinforcement learning,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:45:04.720748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:45:00.425882Z digest=sha256:ea63c6cc4a70a12c3300ca086333d65eec39488cf66564a26b1fe2c3bc4c6cfe

Observation d435a8a6-a97a-4815-8819-1cf31dee32e2 · outbound

This paper cites Sutton and A.

PPO-BR: Dual-Signal Entropy-Reward Adaptation for Trust Region Policy Optimization Sutton and A

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:45:04.445147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:45:00.489904Z digest=sha256:e92a0c4f4c6a4314bc724435226da855a3d6c310111efb0e0e70953f931ffed1

Observation db774f95-8495-4d18-87a8-156d86dc4837 · outbound

This paper cites Reinforcement learning: A survey,.

PPO-BR: Dual-Signal Entropy-Reward Adaptation for Trust Region Policy Optimization Reinforcement learning: A survey,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:45:04.108939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:45:00.593465Z digest=sha256:17b822b45e462a53df4fd407218a9ee9678773d282cc70de7add98afc346e189

Observation 49f59f67-a684-43b3-9e07-d88295652ad2 · outbound

This paper cites Simple statistical gradient-following algorithms for connectionist reinforcement learning,.

PPO-BR: Dual-Signal Entropy-Reward Adaptation for Trust Region Policy Optimization Simple statistical gradient-following algorithms for connectionist reinforcement learning,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:45:03.861939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:45:00.667690Z digest=sha256:ed7998729dd6448c4cd980a6dabea0b9e6019ce39fe733a525c8dba0e928f68a

Observation ce536db6-f0aa-47ec-bbae-acfece1719f7 · outbound

This paper cites High-dimensional continuous control using generalized advantage estimation,.

PPO-BR: Dual-Signal Entropy-Reward Adaptation for Trust Region Policy Optimization High-dimensional continuous control using generalized advantage estimation,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:45:03.634034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:45:00.768484Z digest=sha256:e81820c289f39d070a5dceadb990241a9d9a77cef19a11c755e54a3cc78299f0

Observation 350014f5-19c8-4dc2-b356-18a77b7a37c4 · outbound

This paper cites Grandmaster level in StarCraft II using multi-agent reinforcement learning,.

PPO-BR: Dual-Signal Entropy-Reward Adaptation for Trust Region Policy Optimization Grandmaster level in StarCraft II using multi-agent reinforcement learning,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:45:03.466030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:45:00.891898Z digest=sha256:2d232d3cbde7a329558e9374760ea8f90e8f2f0cdc27e55d3874c454e24a4faf

Observation 7dc67ac9-c92c-44ff-88af-46b978cdf3e7 · outbound

This paper cites Annealed policy optimization,.

PPO-BR: Dual-Signal Entropy-Reward Adaptation for Trust Region Policy Optimization Annealed policy optimization,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:45:03.117391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:45:01.006215Z digest=sha256:7ce5e8309cf086330f9a62204d9b782c2e27e59672557b89b4e4ab0e6d120465

Observation f77c92cb-f40a-42c5-b9bd-a3ef5461038d · outbound

This paper cites Noisy networks for exploration,.

PPO-BR: Dual-Signal Entropy-Reward Adaptation for Trust Region Policy Optimization Noisy networks for exploration,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:45:02.878463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:45:01.112097Z digest=sha256:34b9ba74476c67db263097a46539adc46dc92abc5471eb28bf250f79cbcfbfa1

Observation 725dfd50-53b0-4a33-abe4-67a6f5dbd49b · outbound

This paper cites Deep Reinforcement Learning: An Overview,.

PPO-BR: Dual-Signal Entropy-Reward Adaptation for Trust Region Policy Optimization Deep Reinforcement Learning: An Overview,

Reference 16

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T14:45:02.226161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:45:01.223279Z digest=sha256:c4056f18ee5aef985ec452ea19a2b8e1c922180e4a8793e000edfa7afeded3bd

Observation 5a5b95cf-715c-4aaf-8a2d-a4aa84a5f2d4 · outbound

This paper cites Constrained Policy Optimization,.

PPO-BR: Dual-Signal Entropy-Reward Adaptation for Trust Region Policy Optimization Constrained Policy Optimization,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:45:02.628187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:45:01.373170Z digest=sha256:dace27ea643e63efe3b699251a540d3082fa5ad1f8f5b1bca81bc045dc271e3c

Observation 6e852697-3be3-4650-a1f9-4a74d4c7c120 · outbound

This paper cites A Study on Overfitting in Deep Reinforcement Learning,.

PPO-BR: Dual-Signal Entropy-Reward Adaptation for Trust Region Policy Optimization A Study on Overfitting in Deep Reinforcement Learning,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:45:02.404638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:45:01.506161Z digest=sha256:cd77f26d07ba65d884eb60b6fbeb09a059f1c7f29c5367e4cf6822215710cda7

Observation b155d807-98c2-4b2d-967a-679e0e6b73db · outbound

This paper cites Optimizing Customer Satisfaction Through Sentiment Analysis: A BERT-Based Machine Learning Approach to Extract Insights,.

PPO-BR: Dual-Signal Entropy-Reward Adaptation for Trust Region Policy Optimization Optimizing Customer Satisfaction Through Sentiment Analysis: A BERT-Based Machine Learning Approach to Extract Insights,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:45:01.598606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:45:01.598606Z digest=sha256:d941d911532cf33698a1b3395845e4b142ce4653139698ec64fba62c64320c46

Observation ec1f8a1a-a535-4f03-9ffd-7d932f0ab305 · outbound

This paper cites Context-Aware Semantic Segmentation: Enhancing Pixel-Level Understanding with Large Language Models for Advanced Vision Applications.

PPO-BR: Dual-Signal Entropy-Reward Adaptation for Trust Region Policy Optimization Context-Aware Semantic Segmentation: Enhancing Pixel-Level Understanding with Large Language Models for Advanced Vision Applications

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:45:01.725671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:45:01.725671Z digest=sha256:d12c74d8966c7970c80bef688856102896a3b165d0df4f2c75ab17fd886bbcac

Pith citing papers

No inbound Pith citation observations are available.