Pith. sign in

Paper Citation Record · LEDGER

Trust Region Preference Approximation: A simple and stable reinforcement learning algorithm for LLM reasoning

As of 4 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2504.04524.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.04524 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T06:47:16.873682Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T08:39:42.243396Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 8a53ea4e-2d10-467b-868c-15af3475e4ae · inbound

Learning to Reason under Off-Policy Guidance cites this paper.

Learning to Reason under Off-Policy Guidance Trust Region Preference Approximation: A simple and stable reinforcement learning algorithm for LLM reasoning

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-15T23:17:02.863067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T23:17:02.701393Z digest=sha256:51c6685757679bd6fdbf116f24baa34ff558f0220a0e8c9d0c4e6e974ac99c9a

Observation 18f62dfc-7d9f-4739-b673-f9e314a89751 · inbound

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards cites this paper.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Trust Region Preference Approximation: A simple and stable reinforcement learning algorithm for LLM reasoning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T19:38:22.422879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:38:22.422879Z digest=sha256:9540aa29780bc6e198fa465d4ad85be4e1d6955468f6036852eae31e8e875338

Observation d354e495-9260-4fa8-a6ce-32bd312f9be2 · inbound

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards cites this paper.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Trust Region Preference Approximation: A simple and stable reinforcement learning algorithm for LLM reasoning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.873682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.873682Z digest=sha256:f144470c4b0d87bb98b9a33d32c5ca2e1aeefb7b7acdeef649803afe5500e4fe

Observation fccb151c-0cc3-4f93-83c2-cf182470efc2 · inbound

Revisiting Reinforcement Learning with Verifiable Rewards from a Contrastive Perspective cites this paper.

Revisiting Reinforcement Learning with Verifiable Rewards from a Contrastive Perspective Trust Region Preference Approximation: A simple and stable reinforcement learning algorithm for LLM reasoning

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:22:54.809006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T20:21:26.640493Z digest=sha256:e370eedb6922f1088532841aee058d35f50c302d31893bf8d86367189b5a9119

Observation 19fedf20-ab9c-4e66-b1ab-2428440dc814 · inbound

Revisiting Reinforcement Learning with Verifiable Rewards from a Contrastive Perspective cites this paper.

Revisiting Reinforcement Learning with Verifiable Rewards from a Contrastive Perspective Trust Region Preference Approximation: A simple and stable reinforcement learning algorithm for LLM reasoning

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-20T21:43:45.426051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T21:42:49.452347Z digest=sha256:16c9672f4a6f20a91e107719ad916af6b13538ba7488ce1f92361038adb377b6

Observation 7ef419e3-2fbe-40db-8337-6893563729c0 · inbound

Curriculum Reinforcement Learning Can Incentivize Reasoning Capacity in LLMs Beyond the Base Model cites this paper.

Curriculum Reinforcement Learning Can Incentivize Reasoning Capacity in LLMs Beyond the Base Model Trust Region Preference Approximation: A simple and stable reinforcement learning algorithm for LLM reasoning

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:39:42.245665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T11:10:56.755558Z digest=sha256:6d01688aabb760c1e13a536fbfb1cead819d77fd61c1bbbc370ca9a00786dca2