Pith. sign in

Paper Citation Record · LEDGER

Efficient RLHF: Reducing the Memory Usage of PPO

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2309.00754.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2309.00754 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:13:18.743392Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-13T06:17:23.152582Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation beff939c-dcde-4dc1-a8e3-497cdbd1638e · inbound

HybridFlow: A Flexible and Efficient RLHF Framework cites this paper.

HybridFlow: A Flexible and Efficient RLHF Framework Efficient RLHF: Reducing the Memory Usage of PPO

Reference 75

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T07:53:38.837980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T07:53:38.715353Z digest=sha256:a5c2afcdc9607e772b775fb0e7836c008669704402a01a27f3136b18a87a8e82

Observation 9d141c4d-440b-4d20-8b30-a8982e54cb9d · inbound

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race cites this paper.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Efficient RLHF: Reducing the Memory Usage of PPO

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:18.743392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:18.743392Z digest=sha256:a1096b3433cbb50577cb2526d9c9de05c599ca6354907828d35d40812d492bb9

Observation fb625c15-5994-4293-911c-4a98e22c47b9 · inbound

Aligning Large Language Models with Implicit Preferences from User-Generated Content cites this paper.

Aligning Large Language Models with Implicit Preferences from User-Generated Content Efficient RLHF: Reducing the Memory Usage of PPO

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.123936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.123936Z digest=sha256:778c816b1270b11fa898cf543b6b0cd037fa5324a0e064a858e3385250e24313

Observation a86b0c55-969d-4b60-92a1-e5a81a318488 · inbound

A Technical Survey of Reinforcement Learning Techniques for Large Language Models cites this paper.

A Technical Survey of Reinforcement Learning Techniques for Large Language Models Efficient RLHF: Reducing the Memory Usage of PPO

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T19:59:32.527866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:59:32.527866Z digest=sha256:d77cbc78a8dd2c59a19f544aff4df384f3fda9d71b27a78e11383f861c67e945

Observation 50bc0b6a-c2dd-4d41-a0cf-da219e2fa9ae · inbound

Representation-Guided Parameter-Efficient LLM Unlearning cites this paper.

Representation-Guided Parameter-Efficient LLM Unlearning Efficient RLHF: Reducing the Memory Usage of PPO

Reference 202

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T06:06:19.194939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T06:01:46.885030Z digest=sha256:f44b4734ce4cc97cf61f6108808205c74f56e2156345896b6b3d963d816e8936

Observation 9964e064-08c2-48bf-b2b2-1a7de21bdffe · inbound

Reference-Sampled Boltzmann Projection for KL-Regularized RLVR: Target-Matched Weighted SFT, Finite One-Shot Gaps, and Policy Mirror Descent cites this paper.

Reference-Sampled Boltzmann Projection for KL-Regularized RLVR: Target-Matched Weighted SFT, Finite One-Shot Gaps, and Policy Mirror Descent Efficient RLHF: Reducing the Memory Usage of PPO

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-09T05:45:23.216943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T19:34:22.546508Z digest=sha256:1f4a960afd246607e56a49a419ac2b3a626beaa3ae05728529d4c62891b6c612

Observation 786db28a-13ad-4303-b908-504c09691404 · inbound

Reinforcement Learning for Scalable and Trustworthy Intelligent Systems cites this paper.

Reinforcement Learning for Scalable and Trustworthy Intelligent Systems Efficient RLHF: Reducing the Memory Usage of PPO

Reference 96

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:51:37.926751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T01:47:40.772146Z digest=sha256:ffdafb0ef13a7e919154d953deae7ceea4c9e1ea2fe846ee3f1ae85415c9da9b

Observation d5f26f3d-2527-426f-bff9-191c97966a0b · inbound

MARLaaS: Multi-Tenant Asynchronous Reinforcement Learning as a Service cites this paper.

MARLaaS: Multi-Tenant Asynchronous Reinforcement Learning as a Service Efficient RLHF: Reducing the Memory Usage of PPO

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:01:28.609302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-12T01:24:22.084179Z digest=sha256:ad6199cefa4b1aae96abd6b96c2c12ebe47773fceac7b0ec1f7d1b3fe7021649

Observation a70cb779-6bbd-47c5-b285-e32cf2ec6730 · inbound

Not How Many, But Which: Parameter Placement in Low-Rank Adaptation cites this paper.

Not How Many, But Which: Parameter Placement in Low-Rank Adaptation Efficient RLHF: Reducing the Memory Usage of PPO

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:17:23.154491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T06:13:55.497799Z digest=sha256:12353271fe28134d4be806f4a581fe392da71fdfc7fc8ea8663f2e54955e3a59