Pith. sign in

Paper Citation Record · LEDGER

ExploRLLM: Guiding Exploration in Reinforcement Learning with Large Language Models

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2403.09583.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.09583 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:47:41.236915Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T02:26:26.974524Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 41423878-e643-4d71-a570-f1945e5c1435 · inbound

Training-free Generation of Temporally Consistent Rewards from VLMs cites this paper.

Training-free Generation of Temporally Consistent Rewards from VLMs ExploRLLM: Guiding Exploration in Reinforcement Learning with Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:41.236915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:47:41.236915Z digest=sha256:1f4ad05c75f58a360df2037681c532b70c0d208bb8aa5be0220e8983f5a61b62

Observation e599a50b-cceb-4295-ae3c-d470fff19073 · inbound

Accelerating Reinforcement Learning Algorithms Convergence using Pre-trained Large Language Models as Tutors With Advice Reusing cites this paper.

Accelerating Reinforcement Learning Algorithms Convergence using Pre-trained Large Language Models as Tutors With Advice Reusing ExploRLLM: Guiding Exploration in Reinforcement Learning with Large Language Models

Reference 106

Resolution
unresolved
no resolver link, observed 2026-08-04T20:49:37.963470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:49:37.963470Z digest=sha256:7b58fa8da5f80c7a3a880eb11deaf892c57d98d5875eaa6cf85aaae0e31948c2

Observation 83fa3052-c748-4340-a909-f342273640ea · inbound

Curriculum-Based Multi-Tier Semantic Exploration via Deep Reinforcement Learning cites this paper.

Curriculum-Based Multi-Tier Semantic Exploration via Deep Reinforcement Learning ExploRLLM: Guiding Exploration in Reinforcement Learning with Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T19:17:00.606863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:17:00.606863Z digest=sha256:c8aa5444fc308986df3729d7f6857277010ce178ef0cc83472408e987fbda1fd

Observation 53881d46-9304-4121-87e9-313486866ddb · inbound

Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning cites this paper.

Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning ExploRLLM: Guiding Exploration in Reinforcement Learning with Large Language Models

Reference 80

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:15:48.960469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T19:15:27.406778Z digest=sha256:093169096d9355081fb40b5a398247baa2133dce62ccffad6d98e57314e08020

Observation f7b47828-442e-4f1b-acea-ba885c1410aa · inbound

Reinforcement Learning from Cross-domain Videos with Video Prediction Model cites this paper.

Reinforcement Learning from Cross-domain Videos with Video Prediction Model ExploRLLM: Guiding Exploration in Reinforcement Learning with Large Language Models

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T02:26:26.976160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T10:54:49.582908Z digest=sha256:1d17d4ca02b36487b5e2cb84199297de5ec73f24f518553946974c9d8e5a8d68