Pith. sign in

Paper Citation Record · LEDGER

Reinforcement Learning with Random Time Horizons

As of 8 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2506.00962.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.00962 v2

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:04:00.844121Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

31 of 31 outbound references displayed

  • verified exact2
  • verified fuzzy21
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c17865b5-f16c-417e-a61e-14af8991c27a · outbound

This paper cites M., and Sun, W.

Reinforcement Learning with Random Time Horizons M., and Sun, W

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:04:08.314379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:03:57.285325Z digest=sha256:dcb1f22ff41689fbea992c9e11ac545a7e0959cd71d18df855f39df1b859d7cf

Observation 3d79a602-1722-42b5-8dcc-109581632a6f · outbound

This paper cites An optimal control perspective on diffusion-based generative modeling.

Reinforcement Learning with Random Time Horizons An optimal control perspective on diffusion-based generative modeling

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:04:07.943354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:03:57.399726Z digest=sha256:4aed7e08272b5642e5b4e01a54600609c791fd00d7a9047e40ddfaee8de43c99

Observation 11848eba-7792-4a3a-8349-4313dd1673b4 · outbound

This paper cites Steady state analysis of episodic reinforcement learning.

Reinforcement Learning with Random Time Horizons Steady state analysis of episodic reinforcement learning

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:04:07.587976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:03:57.573753Z digest=sha256:fdf2682a25f6a6f29a1d54a4a59a33f693b5f6cc9eda38421a304086b37ce8b8

Observation 7d63d11e-ac53-4719-9996-b3efa119e559 · outbound

This paper cites Finite-Sample Analysis of the Monte Carlo Exploring Starts Algorithm for Reinforcement Learning.

Reinforcement Learning with Random Time Horizons Finite-Sample Analysis of the Monte Carlo Exploring Starts Algorithm for Reinforcement Learning

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:04:01.429448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:03:57.759352Z digest=sha256:3aa1c96f1f86ec3c0c72654b246c4ba91283ed7b1762cbfa69619db34681f98f

Observation 6def04ef-1c65-4891-bbd3-dd906e61c7cd · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Random Time Horizons Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:04:07.258358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:03:57.847987Z digest=sha256:7cf178a926e75da40bf20558785e6284593718c7af43669aef7d92b03b2fbf67

Observation bb57821f-0633-4c4a-bde4-23972733b118 · outbound

This paper cites Finite state Markovian decision processes.

Reinforcement Learning with Random Time Horizons Finite state Markovian decision processes

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:04:06.807588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:03:57.919722Z digest=sha256:13c8f825e7a70d8c1559d58e485b84d5ae45b09256266ae094a0100dadc41947

Observation 1d87063d-636e-4e13-9170-a787eb8f0b5f · outbound

This paper cites and Sch \"u tte, C.

Reinforcement Learning with Random Time Horizons and Sch \"u tte, C

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:04:06.403025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:03:58.057033Z digest=sha256:996b05584af83df4c6e3973e6b125dcfa7b9ebac56c7cf4a9e4474642ddb904f

Observation 0438410c-f251-4002-a688-757897fd52a0 · outbound

This paper cites Variational characterization of free energy: Theory and algorithms.

Reinforcement Learning with Random Time Horizons Variational characterization of free energy: Theory and algorithms

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:04:06.142821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:03:58.160551Z digest=sha256:fa24d871131af9a6e61956c1ba4a43827b9b8f060c5f3acbf69b4d0e488bd6c7

Observation d7fc37ce-0105-45ca-a973-6348e1cdcb16 · outbound

This paper cites Fr\'{e}chet derivatives of expected functionals of solutions to stochastic differential equations.

Reinforcement Learning with Random Time Horizons Fr\'{e}chet derivatives of expected functionals of solutions to stochastic differential equations

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:04:01.138144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:03:58.220862Z digest=sha256:4a443a5a3eb130337e6e53c79839b10c1d74f1d170d6c4b523a760d7188c6cd3

Observation 40fdb2c6-5c4c-4d16-97fb-d3442c12ae7c · outbound

This paper cites Continuous control with deep reinforcement learning.

Reinforcement Learning with Random Time Horizons Continuous control with deep reinforcement learning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:04:05.961251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:03:58.355556Z digest=sha256:031f37564ad20a84c230f3baa89e9c7c3b3b7ea03de5bea6f54d7d285152c9b6

Observation e76d062e-9886-463a-884d-8ed878b4972b · outbound

This paper cites Online reinforcement learning with uncertain episode lengths.

Reinforcement Learning with Random Time Horizons Online reinforcement learning with uncertain episode lengths

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:04:05.662630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:03:58.543579Z digest=sha256:ff14b1f7dbdcea9f6554516b8ec938bcaaee348f267b4fe7f0f704766978c631

Observation ff03ed89-d27d-491b-a17e-dde9b5b8bc02 · outbound

This paper cites Playing Atari with Deep Reinforcement Learning.

Reinforcement Learning with Random Time Horizons Playing Atari with Deep Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:58.652529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:03:58.652529Z digest=sha256:1b3d447cae1fff4882958987429ef717cc33b4f86bebb1ea335945a15df769dc

Observation fb4803f5-2282-4267-bdea-77254ea5029f · outbound

This paper cites and Thomas, P.

Reinforcement Learning with Random Time Horizons and Thomas, P

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:04:05.402213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:03:58.796246Z digest=sha256:f5194c377e8f12b96c98c20167c3b2e3b61d2f22b4d00885f6f648e1c3004f87

Observation f83b191f-1f2c-4bbc-aedc-fded12b0ad2d · outbound

This paper cites and Richter, L.

Reinforcement Learning with Random Time Horizons and Richter, L

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:04:05.128034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:03:58.952385Z digest=sha256:124fb405dc597a6c4f410e1adcca9d2099e6eecfcbab63fc4c7bee6442a7e6ab

Observation 285f59f2-a65c-4643-9515-7d892ef44a37 · outbound

This paper cites Stochastic control foundations of autonomous behavior.

Reinforcement Learning with Random Time Horizons Stochastic control foundations of autonomous behavior

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:04:04.952143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:03:59.075138Z digest=sha256:b204398d2c19eb4f35641146c2e48e2e117dc800832b2f15ceb68b1090d66af8

Observation fea4c62f-45d4-416b-aed7-d4b9cf0e2e99 · outbound

This paper cites Continuous-time stochastic control and optimization with financial applications, volume 61.

Reinforcement Learning with Random Time Horizons Continuous-time stochastic control and optimization with financial applications, volume 61

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:04:04.648176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:03:59.188297Z digest=sha256:bf468a45df677b80cafff59536f237fde53fea77c6014bb5d80803dbdfb202bf

Observation 620359c4-2a76-4563-93ca-4493dc8d1dee · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Random Time Horizons Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:59.268173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:03:59.268173Z digest=sha256:d0e3cb8e1308e8edc52e5c0c7df65fc643f125d44b037895614c7cce2b9d8db5

Observation ef440227-290e-422e-bb6f-112ee4e62c7e · outbound

This paper cites and Ribera Borrell, E.

Reinforcement Learning with Random Time Horizons and Ribera Borrell, E

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:04:04.333644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:03:59.385860Z digest=sha256:c22163127f21c1a5ab74b3286fb4990166a44429b392273cac9e1558321f1b17

Observation ea9911ac-36ec-477e-8832-9d4d67f16afd · outbound

This paper cites Improving control based importance sampling strategies for metastable diffusions via adapted metadynamics.

Reinforcement Learning with Random Time Horizons Improving control based importance sampling strategies for metastable diffusions via adapted metadynamics

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:04:03.911120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:03:59.489656Z digest=sha256:c6d5995ce6e5f56447eb973ccb605abe37cf2e607f575e63799cf89ccbed5d63

Observation f08cd45a-5b85-4937-8c03-6786bf0f78e3 · outbound

This paper cites Trust region policy optimization.

Reinforcement Learning with Random Time Horizons Trust region policy optimization

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:04:03.590012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:03:59.583527Z digest=sha256:829fad9c0b7d34cf30c357fae0f2eda7d1d77c49ace57225eeba6e1929865d0a

Observation 48911a4b-217a-475a-967a-300416e5dee8 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Reinforcement Learning with Random Time Horizons Proximal Policy Optimization Algorithms

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:59.660830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:03:59.660830Z digest=sha256:043fcff71dae606a409427b1fde033d36dfea093cf385c07a149e6d7ce46cee1

Observation 26949471-6f45-4a5c-b5a4-1cd2faf1e4a7 · outbound

This paper cites Overcoming the timescale barrier in molecular dynamics: Transfer operators, variational principles and machine learning.

Reinforcement Learning with Random Time Horizons Overcoming the timescale barrier in molecular dynamics: Transfer operators, variational principles and machine learning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:04:03.235858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:03:59.730277Z digest=sha256:362587b1484df4a9359979850b6b49abd64b073211e8f27af99eb04099fd42fa

Observation 0d10f094-e17d-4716-a164-1eab4003f514 · outbound

This paper cites Deterministic policy gradient algorithms.

Reinforcement Learning with Random Time Horizons Deterministic policy gradient algorithms

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:04:03.022097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:03:59.821771Z digest=sha256:8bf996610e6772a35bd5e9944480703a39aa3846198ed88cc16b5fa97139131e

Observation 512c57e3-5cc5-4688-a4b3-e7b465b16bc0 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Random Time Horizons Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:04:02.819558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:03:59.947299Z digest=sha256:078fd588bbad81cdf848c71ce4fef64a7845d1b47e481c263084c3e14388d917

Observation 7646e382-a65e-4fb0-aff2-83722a1e2c91 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Random Time Horizons Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:04:02.645742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:04:00.091133Z digest=sha256:25cc897d03f848b08a851480868b1bf7c282e4b8a58854516ad63ef7cd55a69a

Observation 11121e87-018f-471b-a95d-7e2d24782b50 · outbound

This paper cites S., McAllester, D., Singh, S., and Mansour, Y.

Reinforcement Learning with Random Time Horizons S., McAllester, D., Singh, S., and Mansour, Y

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:04:00.247441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:04:00.247441Z digest=sha256:4aee26fe681d75d14231d3c51ea9afd7c66a48d8c63edcbac687f8f47a4c0523

Observation 08383ad2-08e0-4ec9-99b2-ab3491473147 · outbound

This paper cites U., De Cola, G., Deleu, T., Goulão, M., Kallinteris, A., Krimmel, M., KG, A., Perez-Vicente, R., Pierré, A., Schulhoff, S., Tai, J.

Reinforcement Learning with Random Time Horizons U., De Cola, G., Deleu, T., Goulão, M., Kallinteris, A., Krimmel, M., KG, A., Perez-Vicente, R., Pierré, A., Schulhoff, S., Tai, J

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:04:02.457682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:04:00.334819Z digest=sha256:2439f5bd48f84847a782c762a0e13dc9b26dabfdbb43625b2ffb1d9d7b2327d2

Observation 49aa1aca-32ea-4a76-8e4a-7c88dab3e0fc · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Random Time Horizons Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:04:02.298815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:04:00.441945Z digest=sha256:59db387d0078fc6c1887193509b7bd19975bf39e9af50e5e6c50e54a5b519bc2

Observation eb154fd5-b6ac-41cd-bb99-d20188244c60 · outbound

This paper cites Unifying task specification in reinforcement learning.

Reinforcement Learning with Random Time Horizons Unifying task specification in reinforcement learning

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:04:02.124196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:04:00.568325Z digest=sha256:67990889858ff6ff0d2f23689aedac0476aab798cc3dac987758fa6ec7501b5e

Observation 066d052c-99c8-4bcb-a21f-1bb88685fb30 · outbound

This paper cites Global convergence of policy gradient methods to (almost) locally optimal policies.

Reinforcement Learning with Random Time Horizons Global convergence of policy gradient methods to (almost) locally optimal policies

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:04:01.909939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:04:00.723411Z digest=sha256:bdbbb1a997c85f5933f33c51f47b99caacf12c8cf19510c75d07d4bf1b0e0f1e

Observation 53440e3e-290b-41e0-bdcc-7bd52c7f6e3f · outbound

This paper cites Actor-critic method for high dimensional static H amilton-- J acobi-- B ellman partial differential equations based on neural networks.

Reinforcement Learning with Random Time Horizons Actor-critic method for high dimensional static H amilton-- J acobi-- B ellman partial differential equations based on neural networks

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:04:01.708906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:04:00.844121Z digest=sha256:bce4b8288c8e60f61147731cd1661085d4c5435fa9a3b20277db1615fabd1fb7

Pith citing papers

No inbound Pith citation observations are available.