Pith. sign in

Paper Citation Record · LEDGER

Dealing with Sparse Rewards in Reinforcement Learning

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:1910.09281.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1910.09281 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:18:52.713790Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:29:55.594907Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 3dcc38c7-143c-4d6c-9845-a260871e95c9 · inbound

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL cites this paper.

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL Dealing with Sparse Rewards in Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:52.713790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:52.713790Z digest=sha256:1c891ff59b947567a1226b6d27d7ec77143bb6edaaf8af8b85b7b3625b9b6faa

Observation dc2e7ed3-6589-478f-8935-61ef15aaacef · inbound

SCAR: Shapley Credit Assignment for More Efficient RLHF cites this paper.

SCAR: Shapley Credit Assignment for More Efficient RLHF Dealing with Sparse Rewards in Reinforcement Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:49.225204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:49.225204Z digest=sha256:f9ba22e93c8c9456663c236fbf1371cfa863246b00fa580fd082653eb1129528

Observation ce1b4e95-c112-4510-8d56-4328a4dbab63 · inbound

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? cites this paper.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Dealing with Sparse Rewards in Reinforcement Learning

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:48:57.311552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:b59281587ce114535b5bc0d81f86a75c6f19f15c0c9eca9ca45ca999e0fa2fd1

Observation 824aeed7-4275-48ab-ba4a-539b19db6c86 · inbound

Mesh-RL: Coupled subgrid reinforcement learning cites this paper.

Mesh-RL: Coupled subgrid reinforcement learning Dealing with Sparse Rewards in Reinforcement Learning

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:29:55.596964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T01:37:31.834106Z digest=sha256:79a38ed82d01d4ee20b06f31bcdc093fcecfacd8944b576fffa98b3e06cda0c3

Observation d234f2db-73f8-4d88-8010-ee730468adda · inbound

Learning Gait-Aware Quadruped Locomotion with Temporal Logic Specifications cites this paper.

Learning Gait-Aware Quadruped Locomotion with Temporal Logic Specifications Dealing with Sparse Rewards in Reinforcement Learning

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:06:55.368456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-02T11:59:04.103002Z digest=sha256:f43e8fc2d1d14a6802e26c9dd0144c3e0d86f31142560a43e9b5f94b4b05eb4e