Pith. sign in

Paper Citation Record · LEDGER

Reinforcement Learning with Random Time Horizons

As of 20 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2506.00962.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.00962 v2

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:04:00.844121Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

31 of 31 outbound references displayed

  • verified exact2
  • verified fuzzy21
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c17865b5-f16c-417e-a61e-14af8991c27a · outbound

This paper cites M., and Sun, W.

Reinforcement Learning with Random Time Horizons M., and Sun, W

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:04:08.314379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T12:03:57.285325Z digest=sha256:d9063a4ddf444ae3d40e2505f936cf24184e5fce1a27376c4aa0785c6db120af

Observation 3d79a602-1722-42b5-8dcc-109581632a6f · outbound

This paper cites An optimal control perspective on diffusion-based generative modeling.

Reinforcement Learning with Random Time Horizons An optimal control perspective on diffusion-based generative modeling

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:04:07.943354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T12:03:57.399726Z digest=sha256:4a35d1a52c59c45a695080bd276a6d526c6744107109ed5b0872ccec83758169

Observation 11848eba-7792-4a3a-8349-4313dd1673b4 · outbound

This paper cites Steady state analysis of episodic reinforcement learning.

Reinforcement Learning with Random Time Horizons Steady state analysis of episodic reinforcement learning

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:04:07.587976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T12:03:57.573753Z digest=sha256:cb623dfa8eabea87eca1fe99accd5afbc59e9a63c76986b2b8851f35663aa01e

Observation 7d63d11e-ac53-4719-9996-b3efa119e559 · outbound

This paper cites Finite-Sample Analysis of the Monte Carlo Exploring Starts Algorithm for Reinforcement Learning.

Reinforcement Learning with Random Time Horizons Finite-Sample Analysis of the Monte Carlo Exploring Starts Algorithm for Reinforcement Learning

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:04:01.429448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T12:03:57.759352Z digest=sha256:6437e6821cd5ee54e68f1b385663281fae8ac0836887df76c64a33277aacfacf

Observation 6def04ef-1c65-4891-bbd3-dd906e61c7cd · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Random Time Horizons Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:04:07.258358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T12:03:57.847987Z digest=sha256:00cc46d9ce2b35c790b1b02972a9a28e42e2c18ed66679ff045d35259e64e45b

Observation bb57821f-0633-4c4a-bde4-23972733b118 · outbound

This paper cites Finite state Markovian decision processes.

Reinforcement Learning with Random Time Horizons Finite state Markovian decision processes

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:04:06.807588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T12:03:57.919722Z digest=sha256:befe5d6446d6a297cdce6488137bfacae4b82ddd396ea3dfda37a1cca457b58d

Observation 1d87063d-636e-4e13-9170-a787eb8f0b5f · outbound

This paper cites and Sch \"u tte, C.

Reinforcement Learning with Random Time Horizons and Sch \"u tte, C

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:04:06.403025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T12:03:58.057033Z digest=sha256:92736d559cf48f382b8eda093bd89c95334ed71fc05cea101bf624125a2f58d2

Observation 0438410c-f251-4002-a688-757897fd52a0 · outbound

This paper cites Variational characterization of free energy: Theory and algorithms.

Reinforcement Learning with Random Time Horizons Variational characterization of free energy: Theory and algorithms

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:04:06.142821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T12:03:58.160551Z digest=sha256:8b4cd5e1977886849629b9637b5cebfd594b9fa27be1b3bea205d98ac2891e32

Observation d7fc37ce-0105-45ca-a973-6348e1cdcb16 · outbound

This paper cites Fr\'{e}chet derivatives of expected functionals of solutions to stochastic differential equations.

Reinforcement Learning with Random Time Horizons Fr\'{e}chet derivatives of expected functionals of solutions to stochastic differential equations

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:04:01.138144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T12:03:58.220862Z digest=sha256:a70ea092be1b1855e9fe465ec5eabb63bede6b9df32a1d4cd78e88082c67857d

Observation 40fdb2c6-5c4c-4d16-97fb-d3442c12ae7c · outbound

This paper cites Continuous control with deep reinforcement learning.

Reinforcement Learning with Random Time Horizons Continuous control with deep reinforcement learning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:04:05.961251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T12:03:58.355556Z digest=sha256:03ac73a672b9ddda084d7a0ec2e67a7d1303749352f6a02c5cecaad2ce201cf8

Observation e76d062e-9886-463a-884d-8ed878b4972b · outbound

This paper cites Online reinforcement learning with uncertain episode lengths.

Reinforcement Learning with Random Time Horizons Online reinforcement learning with uncertain episode lengths

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:04:05.662630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T12:03:58.543579Z digest=sha256:b17af0a8c34a7ee23a02ddc577e6523ca01a7ba9a227f3c4400aa6b4dfe94bac

Observation ff03ed89-d27d-491b-a17e-dde9b5b8bc02 · outbound

This paper cites Playing Atari with Deep Reinforcement Learning.

Reinforcement Learning with Random Time Horizons Playing Atari with Deep Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:58.652529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:03:58.652529Z digest=sha256:fb5bbb5cb4c32e419bf66ccc47248e867a97bf267f563b3ff94835e69a8a5b3b

Observation fb4803f5-2282-4267-bdea-77254ea5029f · outbound

This paper cites and Thomas, P.

Reinforcement Learning with Random Time Horizons and Thomas, P

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:04:05.402213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T12:03:58.796246Z digest=sha256:547d259bd217abc0f3984a150797dc12bcb57c61ff42875287be2a3a0fe8b7e5

Observation f83b191f-1f2c-4bbc-aedc-fded12b0ad2d · outbound

This paper cites and Richter, L.

Reinforcement Learning with Random Time Horizons and Richter, L

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:04:05.128034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T12:03:58.952385Z digest=sha256:0d487e021ffbbb1c4c6c1f1c7dc3fd9e4f9038063e57bcb5e82a791a47fe4700

Observation 285f59f2-a65c-4643-9515-7d892ef44a37 · outbound

This paper cites Stochastic control foundations of autonomous behavior.

Reinforcement Learning with Random Time Horizons Stochastic control foundations of autonomous behavior

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:04:04.952143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T12:03:59.075138Z digest=sha256:a4021cd321a370534c27be347b07f0b6c5a6fccdac3dfc0e4b1fe0a5c2cd72df

Observation fea4c62f-45d4-416b-aed7-d4b9cf0e2e99 · outbound

This paper cites Continuous-time stochastic control and optimization with financial applications, volume 61.

Reinforcement Learning with Random Time Horizons Continuous-time stochastic control and optimization with financial applications, volume 61

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:04:04.648176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T12:03:59.188297Z digest=sha256:a6709e60f0757ca7da8377329cbb334cb44ada929db33d02a63845499f7aa37d

Observation 620359c4-2a76-4563-93ca-4493dc8d1dee · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Random Time Horizons Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:59.268173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:03:59.268173Z digest=sha256:0f5f93d8dcc27006929d75dc7ea14837787eb3b76da196a760e51836e1c65557

Observation ef440227-290e-422e-bb6f-112ee4e62c7e · outbound

This paper cites and Ribera Borrell, E.

Reinforcement Learning with Random Time Horizons and Ribera Borrell, E

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:04:04.333644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T12:03:59.385860Z digest=sha256:6b2186866b9a2cad59b11fde57f38551dc971054c6476467ac4d18125a2626a3

Observation ea9911ac-36ec-477e-8832-9d4d67f16afd · outbound

This paper cites Improving control based importance sampling strategies for metastable diffusions via adapted metadynamics.

Reinforcement Learning with Random Time Horizons Improving control based importance sampling strategies for metastable diffusions via adapted metadynamics

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:04:03.911120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T12:03:59.489656Z digest=sha256:375bf895bee6458225b5988b7efaa555af4c250e51e45712b616b4996492eec2

Observation f08cd45a-5b85-4937-8c03-6786bf0f78e3 · outbound

This paper cites Trust region policy optimization.

Reinforcement Learning with Random Time Horizons Trust region policy optimization

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:04:03.590012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T12:03:59.583527Z digest=sha256:9fb51a185bde04ea825b5b0055dc1f22b2d2a907629dff22010b1515703e005f

Observation 48911a4b-217a-475a-967a-300416e5dee8 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Reinforcement Learning with Random Time Horizons Proximal Policy Optimization Algorithms

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:59.660830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:03:59.660830Z digest=sha256:da2c4427906ba315c24217d4dc0546e04f5e17ba74dc4272cc96b4133ab0b00f

Observation 26949471-6f45-4a5c-b5a4-1cd2faf1e4a7 · outbound

This paper cites Overcoming the timescale barrier in molecular dynamics: Transfer operators, variational principles and machine learning.

Reinforcement Learning with Random Time Horizons Overcoming the timescale barrier in molecular dynamics: Transfer operators, variational principles and machine learning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:04:03.235858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T12:03:59.730277Z digest=sha256:6a8062ed629f55c5f17285c9ebb2aa8ef2a57d9304197e9bc52e29f4bcc1602a

Observation 0d10f094-e17d-4716-a164-1eab4003f514 · outbound

This paper cites Deterministic policy gradient algorithms.

Reinforcement Learning with Random Time Horizons Deterministic policy gradient algorithms

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:04:03.022097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T12:03:59.821771Z digest=sha256:ec71f7c5455c4f32ebc573add948f7f18efe658325a435de53f4f5aeb27695d4

Observation 512c57e3-5cc5-4688-a4b3-e7b465b16bc0 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Random Time Horizons Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:04:02.819558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T12:03:59.947299Z digest=sha256:6261e0e1635bd8cdbe09d0ba036e8fbdc8e9aea736520045a7e27027bf75797b

Observation 7646e382-a65e-4fb0-aff2-83722a1e2c91 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Random Time Horizons Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:04:02.645742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T12:04:00.091133Z digest=sha256:b0db0b9288838eababc435f5cf4e291b4850d559c0d17dded1bfd6ef65872ee2

Observation 11121e87-018f-471b-a95d-7e2d24782b50 · outbound

This paper cites S., McAllester, D., Singh, S., and Mansour, Y.

Reinforcement Learning with Random Time Horizons S., McAllester, D., Singh, S., and Mansour, Y

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:04:00.247441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:04:00.247441Z digest=sha256:e949207b02e971e72a1f897d440fc1e5d42ebeba092f3375357169ab4478a717

Observation 08383ad2-08e0-4ec9-99b2-ab3491473147 · outbound

This paper cites U., De Cola, G., Deleu, T., Goulão, M., Kallinteris, A., Krimmel, M., KG, A., Perez-Vicente, R., Pierré, A., Schulhoff, S., Tai, J.

Reinforcement Learning with Random Time Horizons U., De Cola, G., Deleu, T., Goulão, M., Kallinteris, A., Krimmel, M., KG, A., Perez-Vicente, R., Pierré, A., Schulhoff, S., Tai, J

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:04:02.457682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T12:04:00.334819Z digest=sha256:01ec551bf20b4ffe3d5943cfbf07f84a360ba0ae2e9f38f0164f6ec29c4d42d3

Observation 49aa1aca-32ea-4a76-8e4a-7c88dab3e0fc · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Random Time Horizons Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:04:02.298815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T12:04:00.441945Z digest=sha256:9fe3d67d3c6a97fe55259a4bc94938eb182b1f34d8be04714297c8a66c9613d5

Observation eb154fd5-b6ac-41cd-bb99-d20188244c60 · outbound

This paper cites Unifying task specification in reinforcement learning.

Reinforcement Learning with Random Time Horizons Unifying task specification in reinforcement learning

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:04:02.124196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T12:04:00.568325Z digest=sha256:9b2d95a6e6fb03f80dee3d74551ec108077c2021c54de3fa2248207f9a92a2ba

Observation 066d052c-99c8-4bcb-a21f-1bb88685fb30 · outbound

This paper cites Global convergence of policy gradient methods to (almost) locally optimal policies.

Reinforcement Learning with Random Time Horizons Global convergence of policy gradient methods to (almost) locally optimal policies

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:04:01.909939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T12:04:00.723411Z digest=sha256:3123903ba7985e9e2ce2a0285077fa2700f43f781919f74ca3f5ce96c983a006

Observation 53440e3e-290b-41e0-bdcc-7bd52c7f6e3f · outbound

This paper cites Actor-critic method for high dimensional static H amilton-- J acobi-- B ellman partial differential equations based on neural networks.

Reinforcement Learning with Random Time Horizons Actor-critic method for high dimensional static H amilton-- J acobi-- B ellman partial differential equations based on neural networks

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:04:01.708906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T12:04:00.844121Z digest=sha256:383b7ac5dcd66e81c48e900983ad5ecf9ecfea65512fd70f9ede6e90b824ced9

Pith citing papers

No inbound Pith citation observations are available.