Pith. sign in

Paper Citation Record · LEDGER

Monte Carlo Beam Search for Actor-Critic Reinforcement Learning in Continuous Control

As of 18 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 0 inbound Pith citation observations for arXiv:2505.09029.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.09029 v1

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:45:35.150536Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

24 of 24 outbound references displayed

  • verified exact0
  • verified fuzzy14
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b214486c-e58f-4253-8f3b-f0f7564568bf · outbound

This paper cites ”Reinforcement learning: An introduction.

Monte Carlo Beam Search for Actor-Critic Reinforcement Learning in Continuous Control ”Reinforcement learning: An introduction

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:45:35.547841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:45:35.040358Z digest=sha256:6289df8c4f3ae138b66f3e60c4ec877a5cc0212eb192fb6fa5c138c37907d273

Observation cae4ee93-2c40-4817-881e-53a88a3a888f · outbound

This paper cites ”Reinforcement learning algorithms: A brief survey.” Expert Systems with Applications 231 (2023): 120495.

Monte Carlo Beam Search for Actor-Critic Reinforcement Learning in Continuous Control ”Reinforcement learning algorithms: A brief survey.” Expert Systems with Applications 231 (2023): 120495

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:45:35.533228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:45:35.045488Z digest=sha256:1c628ee8baccc1ece74b8ce8193a0c0318ac47cad7775cf73b5762d986c3b745

Observation 9e7cba20-a200-4130-adee-723c7eba24cd · outbound

This paper cites ”Using reinforcement learning for load testing of video games.” Proceedings of the 44th international conference on software engineering.

Monte Carlo Beam Search for Actor-Critic Reinforcement Learning in Continuous Control ”Using reinforcement learning for load testing of video games.” Proceedings of the 44th international conference on software engineering

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:45:35.517836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:45:35.050151Z digest=sha256:a073a334f8368b9c8a61352033f34fed4ae259f06273888d9cc5ba725d0efaa0

Observation 6c3251f6-180d-4827-8cb1-ca76e7cd2591 · outbound

This paper cites ”A comprehensive survey of research towards AI-enabled unmanned aerial systems in pre-, active-, and post-wildfire management.” Information Fusion (2024): 102369.

Monte Carlo Beam Search for Actor-Critic Reinforcement Learning in Continuous Control ”A comprehensive survey of research towards AI-enabled unmanned aerial systems in pre-, active-, and post-wildfire management.” Information Fusion (2024): 102369

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:45:35.503172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:45:35.054810Z digest=sha256:78ffac4f095b9dd5c8e412a472a448c4283e22ba3367c84256c376a04c8f49ac

Observation 30fd499a-75dc-4bcf-b16c-6153aecc0ae8 · outbound

This paper cites an unresolved cited work.

Monte Carlo Beam Search for Actor-Critic Reinforcement Learning in Continuous Control Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:45:35.488101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:45:35.060042Z digest=sha256:7222fdf5905cd4758baeb677909f323815c41dc995edbe1ef252dbf1bde2e18f

Observation 0184d00b-5bf5-4fab-ac08-c62caccdaf01 · outbound

This paper cites MuJoCo Playground.

Monte Carlo Beam Search for Actor-Critic Reinforcement Learning in Continuous Control MuJoCo Playground

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T21:45:35.064770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:45:35.064770Z digest=sha256:f7522ec1b6b3195e991b2ddeb7d05d8a38bb7131ac9bddaf09129fabafe01919

Observation af29866b-345d-4b86-af39-4cf23f056e9f · outbound

This paper cites ”Continuous control actions learning and adaptation for robotic manipulation through reinforcement learning.” Autonomous Robots 46.3 (2022): 483-498.

Monte Carlo Beam Search for Actor-Critic Reinforcement Learning in Continuous Control ”Continuous control actions learning and adaptation for robotic manipulation through reinforcement learning.” Autonomous Robots 46.3 (2022): 483-498

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:45:35.472673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:45:35.070259Z digest=sha256:560a78d4b8801524585a2ce25bcd15a730c00465eb324b7141ba5825cf4c91e2

Observation b5f35a32-d0f8-405b-92c3-84355e64a4ed · outbound

This paper cites ”Deep deterministic policy gradient algorithm: A systematic review.” Heliyon (2024).

Monte Carlo Beam Search for Actor-Critic Reinforcement Learning in Continuous Control ”Deep deterministic policy gradient algorithm: A systematic review.” Heliyon (2024)

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:45:35.455950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:45:35.074744Z digest=sha256:459fe1c4e166f575ae23dadee595b74fe96a3810562dba6cf9eaa50ffcb1e96d

Observation 503de0fe-97cc-423e-8862-ca89dfbd3465 · outbound

This paper cites ”Addressing function approximation error in actor-critic methods.” International conference on machine learning.

Monte Carlo Beam Search for Actor-Critic Reinforcement Learning in Continuous Control ”Addressing function approximation error in actor-critic methods.” International conference on machine learning

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:45:35.439305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:45:35.079309Z digest=sha256:5aeb6b1aff7cf1c7d1bc9e2ce9998967649545d5398a575fda2d1721e1064a6b

Observation 2020f5a2-5dac-4227-b83e-ced049cbbba8 · outbound

This paper cites ”Stable-baselines3: Reliable reinforcement learning implementations.” Journal of machine learning research 22.268 (2021): 1-8.

Monte Carlo Beam Search for Actor-Critic Reinforcement Learning in Continuous Control ”Stable-baselines3: Reliable reinforcement learning implementations.” Journal of machine learning research 22.268 (2021): 1-8

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:45:35.421564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:45:35.083839Z digest=sha256:f0c820b24a29388046c43e09f443cc514f6bacdf57c11add2025594232009d06

Observation c9a48528-1cfa-4590-a1cd-9d4ee9ef4617 · outbound

This paper cites Eyes on the Environment: AI-Driven Analysis for Fire and Smoke Classification, Segmentation, and Detection.

Monte Carlo Beam Search for Actor-Critic Reinforcement Learning in Continuous Control Eyes on the Environment: AI-Driven Analysis for Fire and Smoke Classification, Segmentation, and Detection

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T21:45:35.088192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:45:35.088192Z digest=sha256:8af26101a9b0bf81508ca5698e985cc85b6e25ae83c05ddcebc004b81bb06b91

Observation 0ce57820-9eae-4921-ad16-ec3d785ba5c2 · outbound

This paper cites Deep Reinforcement Learning Hands-On: Apply modern RL methods, with deep Q-networks, value iteration, policy gradients, TRPO, AlphaGo Zero and more.

Monte Carlo Beam Search for Actor-Critic Reinforcement Learning in Continuous Control Deep Reinforcement Learning Hands-On: Apply modern RL methods, with deep Q-networks, value iteration, policy gradients, TRPO, AlphaGo Zero and more

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:45:35.404273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:45:35.092822Z digest=sha256:66ae81528b1ddd1ac09101f56fb9c19d3c33fc6484a7f9fea71a77c8b50ef650

Observation eda5f139-b497-4678-85be-b4cc802f87e9 · outbound

This paper cites ”AlphaZero.” Deep Reinforce- ment Learning: Fundamentals, Research and Applications (2020): 391- 415.

Monte Carlo Beam Search for Actor-Critic Reinforcement Learning in Continuous Control ”AlphaZero.” Deep Reinforce- ment Learning: Fundamentals, Research and Applications (2020): 391- 415

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:45:35.387544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:45:35.097469Z digest=sha256:151d85e1a1de2985e652f7b46a02ac23c10c3a3561af43f383d87db31c8683aa

Observation 866b4b9a-3db6-4e01-aa06-335aa0eef24b · outbound

This paper cites Automatic Prompt Optimization with "Gradient Descent" and Beam Search.

Monte Carlo Beam Search for Actor-Critic Reinforcement Learning in Continuous Control Automatic Prompt Optimization with "Gradient Descent" and Beam Search

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T21:45:35.101965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:45:35.101965Z digest=sha256:4bab5065bbe36533e9219be8cdbd81e1985b0e3af711971fcea5c0d3d640d9a3

Observation c8ccdd19-8117-4251-94a3-3953d7dbfa9b · outbound

This paper cites ”Simulation-guided beam search for neural com- binatorial optimization.” Advances in Neural Information Processing Systems 35 (2022): 8760-8772.

Monte Carlo Beam Search for Actor-Critic Reinforcement Learning in Continuous Control ”Simulation-guided beam search for neural com- binatorial optimization.” Advances in Neural Information Processing Systems 35 (2022): 8760-8772

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:45:35.371794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:45:35.106933Z digest=sha256:c16c736d78f08e1eebcb95311ceed7c7f7c738a641838b8ff3ebc45207988679

Observation 43d36889-2ea1-4806-b25b-77e502231a41 · outbound

This paper cites ”Monte Carlo tree search: A review of recent modifications and applications.” Artificial Intelligence Review 56.3 (2023): 2497-2562.

Monte Carlo Beam Search for Actor-Critic Reinforcement Learning in Continuous Control ”Monte Carlo tree search: A review of recent modifications and applications.” Artificial Intelligence Review 56.3 (2023): 2497-2562

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:45:35.356588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:45:35.111345Z digest=sha256:19329b8043b70d0f144274381c3c65c8dc470b630df9c93d33bb7eddd0d333ba

Observation ae813dbc-8b2a-44ac-b6b5-de339479dfe9 · outbound

This paper cites Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning.

Monte Carlo Beam Search for Actor-Critic Reinforcement Learning in Continuous Control Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T21:45:35.115772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:45:35.115772Z digest=sha256:40bacd76c2e6a537c5ba60877b983547f88264bcda2e6dfc61b4fe233488f122

Observation 0315eb13-f23d-43cc-b771-cc935a2af31d · outbound

This paper cites ”Beyond greedy search: Tracking by multi-agent reinforcement learning-based beam search.” IEEE Transactions on Image Processing 31 (2022): 6239-6254.

Monte Carlo Beam Search for Actor-Critic Reinforcement Learning in Continuous Control ”Beyond greedy search: Tracking by multi-agent reinforcement learning-based beam search.” IEEE Transactions on Image Processing 31 (2022): 6239-6254

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:45:35.341152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:45:35.120482Z digest=sha256:3fd3fc93b308c779871c72809ea9cf6e4fc4db2f0c3ac403d98d2b82e3692fd8

Observation b68f52c9-d7ec-4cfe-9954-f33d97a001dd · outbound

This paper cites ”RL Baselines3 Zoo.” GitHub repository (2020).

Monte Carlo Beam Search for Actor-Critic Reinforcement Learning in Continuous Control ”RL Baselines3 Zoo.” GitHub repository (2020)

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:45:35.323678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:45:35.125457Z digest=sha256:dede67b886811a32445ba2c7bfb276d82528dcf0058af27f288c46040d83af05

Observation d34cd3fe-9512-440a-9675-026d56de3c96 · outbound

This paper cites OpenAI Gym.

Monte Carlo Beam Search for Actor-Critic Reinforcement Learning in Continuous Control OpenAI Gym

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T21:45:35.130093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:45:35.130093Z digest=sha256:ca8d8a47d4a1e85b19cf77a08157ce583625ddc8c719e8462543644d1ee3cbb5

Observation a7f1d474-7272-4cd3-8e9b-0d075b2ae5aa · outbound

This paper cites Soft Actor-Critic Algorithms and Applications.

Monte Carlo Beam Search for Actor-Critic Reinforcement Learning in Continuous Control Soft Actor-Critic Algorithms and Applications

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T21:45:35.134938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:45:35.134938Z digest=sha256:e459aadb7c1b9702c2820f4c0f82a4ce6412190959c9832d4007521bb9677efd

Observation 845e944e-000f-4355-ab96-3c7494df062e · outbound

This paper cites A2C is a special case of PPO.

Monte Carlo Beam Search for Actor-Critic Reinforcement Learning in Continuous Control A2C is a special case of PPO

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T21:45:35.140853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:45:35.140853Z digest=sha256:8a4e59f606dce294f0d55db85af783bff3fd4f50e395f30487085d3d1d06e559

Observation 6dbc1c6b-53ec-47bd-be34-3cf5cfb4fa7a · outbound

This paper cites Proximal Policy Optimization Algorithms.

Monte Carlo Beam Search for Actor-Critic Reinforcement Learning in Continuous Control Proximal Policy Optimization Algorithms

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T21:45:35.146067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:45:35.146067Z digest=sha256:94760dade3a366d62290f42a5c8f446e94f01cb616656ef0d535583d92f4bb28

Observation ff361cf6-bfb1-4f83-8c13-456afcea3f27 · outbound

This paper cites VisionGPT: LLM-Assisted Real-Time Anomaly Detection for Safe Visual Navigation.

Monte Carlo Beam Search for Actor-Critic Reinforcement Learning in Continuous Control VisionGPT: LLM-Assisted Real-Time Anomaly Detection for Safe Visual Navigation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T21:45:35.150536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:45:35.150536Z digest=sha256:16bb3719c359524161831961b852f334f97c2e3d12ce2f97e17884a6ab37bac4

Pith citing papers

No inbound Pith citation observations are available.