Pith. sign in

Paper Citation Record · LEDGER

Transductive Off-policy Proximal Policy Optimization

As of 11 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2406.03894.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.03894 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T21:10:19.704141Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T02:36:26.594232Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 4fcde9f4-100b-4ac9-86fb-e9cbeb5bc50c · inbound

Revisiting Group Relative Policy Optimization: Insights into On-Policy and Off-Policy Training cites this paper.

Revisiting Group Relative Policy Optimization: Insights into On-Policy and Off-Policy Training Transductive Off-policy Proximal Policy Optimization

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:53.980768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:53.980768Z digest=sha256:29786c3ad709479df3c1445d69526cbdbd56ea00eb7d3ee6956f062c16068a85

Observation fc7b049a-f12a-47d2-ae21-52829fd08589 · inbound

Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges cites this paper.

Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges Transductive Off-policy Proximal Policy Optimization

Reference 230

Resolution
unresolved
no resolver link, observed 2026-08-06T14:13:07.137205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:13:07.137205Z digest=sha256:df27f3a90614fbc3e5680df68d1086791abac8b95f5faeeaced129bb52c3c110

Observation ef5529c6-ec8a-431e-aa8e-3ee4f848e2e5 · inbound

Dissecting Discrete Soft Actor-Critic: Limitations and Principled Alternatives cites this paper.

Dissecting Discrete Soft Actor-Critic: Limitations and Principled Alternatives Transductive Off-policy Proximal Policy Optimization

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-18T17:06:39.755118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-18T17:05:36.100114Z digest=sha256:caf3d8da1758d2cae716462b5a7273787dd011bbd225c6aafdc155983c98b974

Observation 3afa1339-ebe9-41d6-b94f-dff8ad93dcba · inbound

Local Guidance, Global Impact: Gaussian-Reshaped Trust Region Unlocks Behavior Transitions cites this paper.

Local Guidance, Global Impact: Gaussian-Reshaped Trust Region Unlocks Behavior Transitions Transductive Off-policy Proximal Policy Optimization

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T02:36:26.595926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-28T10:51:30.546134Z digest=sha256:bd45ee931718525d468e3e0b637e8c10df2430fa9fba467ddf403f024813e4e5

Observation 31ecbaae-2295-410d-91a4-e33d612dc1cd · inbound

Explore or Converge? Stage-Guided Per-Step Optimization for Diffusion Models cites this paper.

Explore or Converge? Stage-Guided Per-Step Optimization for Diffusion Models Transductive Off-policy Proximal Policy Optimization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:19.704141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:19.704141Z digest=sha256:730da5e4e8be2cfab0a83e16b0da6c70bc669254803e872f226e9bea040da43e