Pith. sign in

Paper Citation Record · LEDGER

Continuously Discovering Novel Strategies via Reward-Switching Policy Optimization

As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 2 inbound Pith citation observations for arXiv:2204.02246.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2204.02246 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 2 of 2 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T00:46:16.638153Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

2
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 6a105918-acd3-4fbf-aed2-097036a68bd8 · inbound

Reward Prediction Error Prioritisation in Experience Replay: The RPE-PER Method cites this paper.

Reward Prediction Error Prioritisation in Experience Replay: The RPE-PER Method Continuously Discovering Novel Strategies via Reward-Switching Policy Optimization

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T00:46:16.638153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:46:16.638153Z digest=sha256:8585f2c79c51557a4f89292e0f9a092adf351f70ef82e4a7102b3f0660e3c642

Observation b62eeae8-63be-42a8-8379-3b568772b613 · inbound

Robust Instruction Compliance in Cooperative Multi-Agent Reinforcement Learning cites this paper.

Robust Instruction Compliance in Cooperative Multi-Agent Reinforcement Learning Continuously Discovering Novel Strategies via Reward-Switching Policy Optimization

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-06-30T22:15:05.811146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-30T22:11:35.277901Z digest=sha256:0d0ac413655e42a7f365500c0308f960de4a6d7e23b5b44d01f248e82a4cc0be