Pith. sign in

Paper Citation Record · LEDGER

Reinforcement Learning for Reachability: Guaranteeing Asymptotic Optimality

As of 6 August 2026, this Paper Citation Record lists 11 of 11 outbound references and 0 inbound Pith citation observations for arXiv:2605.24740.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.24740 v1

Coverage vector

measured 11 of 11 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-30T14:15:45.030536Z

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

11 of 11 outbound references displayed

  • verified exact3
  • verified fuzzy7
  • unresolved0
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation dc37688e-fad1-4a27-b913-fcfa6d0a94f4 · outbound

This paper cites and Henzinger, M.

Reinforcement Learning for Reachability: Guaranteeing Asymptotic Optimality and Henzinger, M

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T23:35:46.407810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T14:15:45.030536Z digest=sha256:0777a137d06c4a90048174c11339603bcd04d05a25285e3b18617d2bfab82297

Observation 6b5e19c7-318f-4292-9c11-3b8442c99c18 · outbound

This paper cites Faster algorithm for turn-based stochastic games with bounded treewidth.

Reinforcement Learning for Reachability: Guaranteeing Asymptotic Optimality Faster algorithm for turn-based stochastic games with bounded treewidth

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T23:35:46.400422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T14:15:45.030536Z digest=sha256:4efb94d9e0e4f175104cd01ecf530bd70f55ada295cdf6b0e177cb96a52b0917

Observation 2674e2d9-34b6-4938-9cde-0dd0bc93af4d · outbound

This paper cites Logically-Constrained Reinforcement Learning.

Reinforcement Learning for Reachability: Guaranteeing Asymptotic Optimality Logically-Constrained Reinforcement Learning

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-06-30T14:24:45.265320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T14:15:45.030536Z digest=sha256:19e5f055913ab48bd77244e9ec5ba9a41a158195dd1037990a73e382d03325f8

Observation 3925b654-030c-4fcf-94b1-ce8a278426c3 · outbound

This paper cites A PAC Learning Algorithm for LTL and Omega-regular Objectives in MDPs.

Reinforcement Learning for Reachability: Guaranteeing Asymptotic Optimality A PAC Learning Algorithm for LTL and Omega-regular Objectives in MDPs

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-06-30T14:24:45.267833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T14:15:45.030536Z digest=sha256:301650888407190e3efddcea3eb1bdfdd8fedff6cb1e23fad6e34dd77b1937d6

Observation 428610bf-446d-44d5-8a7e-ecf0d9f56e65 · outbound

This paper cites On the complexity of omega-automata.

Reinforcement Learning for Reachability: Guaranteeing Asymptotic Optimality On the complexity of omega-automata

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T23:35:46.402304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T14:15:45.030536Z digest=sha256:1730daf01afc6e8f13cc7bdd97a29de81c4477f7054a08614038be116be5d7b5

Observation 42cc6dc3-819c-40b0-87fa-d18166e2366b · outbound

This paper cites Reinforcement learning from reachability specifications: PAC guaran- tees with expected conditional distance.

Reinforcement Learning for Reachability: Guaranteeing Asymptotic Optimality Reinforcement learning from reachability specifications: PAC guaran- tees with expected conditional distance

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T23:35:46.406030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T14:15:45.030536Z digest=sha256:ea36541754246f8fb25d815b97ba41980e7c2798768d62d3f2597fee8f124851

Observation e1a92c9d-f4d2-41d6-ae61-2b33d703ebf1 · outbound

This paper cites Modular Deep Reinforcement Learning with Temporal Logic Specifications.

Reinforcement Learning for Reachability: Guaranteeing Asymptotic Optimality Modular Deep Reinforcement Learning with Temporal Logic Specifications

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-06-30T14:24:45.270482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T14:15:45.030536Z digest=sha256:8417071d520c77b4d4592a0b84e1adbd421dd51454c39fc5f45766757be04718

Observation 83e3207b-b765-44b2-8e29-f07ae4d6cdbe · outbound

This paper cites In the case of transition probabilities, the empirical and true mean translate to the empirical and true probability of a transition’s probability of occurring.

Reinforcement Learning for Reachability: Guaranteeing Asymptotic Optimality In the case of transition probabilities, the empirical and true mean translate to the empirical and true probability of a transition’s probability of occurring

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T23:35:46.396870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T14:15:45.030536Z digest=sha256:caf258e10cb971b1b4ef49c028f514806a216102bbf3c45299b3567cc601ab1f

Observation 4fd95968-cfc9-451d-8298-26696dcf9a19 · outbound

This paper cites The algorithm runs BVI on the collapsed MDP, which is derived from the discovered partial MDP, as mentioned in Section 3.2.

Reinforcement Learning for Reachability: Guaranteeing Asymptotic Optimality The algorithm runs BVI on the collapsed MDP, which is derived from the discovered partial MDP, as mentioned in Section 3.2

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T23:35:46.398648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T14:15:45.030536Z digest=sha256:3e9f585ea2c44afedca875627b65c79c96483f6c12e969066b3ec77dc14d008e

Observation 2779ab9a-8a80-4a1c-ac12-8d86de9b3511 · outbound

This paper cites By this definition V( ˆMC) is the limit of the best lower bound, L(sC,0), which BVI can obtain using ˆP.

Reinforcement Learning for Reachability: Guaranteeing Asymptotic Optimality By this definition V( ˆMC) is the limit of the best lower bound, L(sC,0), which BVI can obtain using ˆP

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T23:35:46.404204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T14:15:45.030536Z digest=sha256:cc6d10ba8d4f7741e355b5e3c1b91b083208b9adb7c92293756d449a5dc4cd42

Observation 9c27cfb6-76a6-45d5-8bf2-dffe75ce1282 · outbound

This paper cites We define Ds,a = lcms′∈S(q(s,a),s′), where lcm is the lowest common multiple of inputs.

Reinforcement Learning for Reachability: Guaranteeing Asymptotic Optimality We define Ds,a = lcms′∈S(q(s,a),s′), where lcm is the lowest common multiple of inputs

Reference 11

Resolution
malformed identifier
raw_fallback, observed 2026-07-08T23:35:46.413907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T14:15:45.030536Z digest=sha256:2b9d16f17e02cf57ab7b59f1d48da3b0fd95a9acd7187ae40ab0449859ff535b

Pith citing papers

No inbound Pith citation observations are available.