Pith. sign in

Paper Citation Record · LEDGER

Multi-turn Reinforcement Learning from Preference Human Feedback

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2405.14655.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.14655 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:32:30.406644Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-17T12:04:10.473003Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation be71c4ed-012f-4b6b-9f99-37eee45a950b · inbound

Training Language Models to Self-Correct via Reinforcement Learning cites this paper.

Training Language Models to Self-Correct via Reinforcement Learning Multi-turn Reinforcement Learning from Preference Human Feedback

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-17T12:04:10.476073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-17T12:04:10.210508Z digest=sha256:9c37c03a9c91d137d24b05d4b7b7860beca2e8ca0364c76feccf4465c752bf9d

Observation b51c2a23-0629-41b7-baf2-ebe83f2b8bfd · inbound

Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models cites this paper.

Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models Multi-turn Reinforcement Learning from Preference Human Feedback

Reference 128

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T21:20:59.256774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T21:20:59.128986Z digest=sha256:96adc07723706ff0bedf03e6310c0cfda7a233bf2c8d811398a7f63f026e7589

Observation 06c03cc1-5efd-4109-ad8b-30f0011fe3dc · inbound

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective cites this paper.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Multi-turn Reinforcement Learning from Preference Human Feedback

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:30.406644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:30.406644Z digest=sha256:7d2e043be8bf76e92f03f78ae56bbcb7c3d9c7e4163c4ea20b0594b76999a2d3

Observation be2dfc51-47b6-4556-814c-b0588fa2897e · inbound

PAG: Multi-Turn Reinforced LLM Self-Correction with Policy as Generative Verifier cites this paper.

PAG: Multi-Turn Reinforced LLM Self-Correction with Policy as Generative Verifier Multi-turn Reinforcement Learning from Preference Human Feedback

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T04:39:35.492469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:39:35.492469Z digest=sha256:eb449b0eeb379cabee0c311d3921033c6e23db41e5ffcba577bd9eb9b3973f7a

Observation 3327ff83-8f88-4819-8567-34e633969dde · inbound

Robust pid sliding mode control for dc servo motor speed control cites this paper.

Robust pid sliding mode control for dc servo motor speed control Multi-turn Reinforcement Learning from Preference Human Feedback

Reference 2010

Resolution
unresolved
no resolver link, observed 2026-08-05T23:42:04.428118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:42:04.428118Z digest=sha256:d49887106a55b65ec748b00cd64c516884a2d7f6335d85c15d67d9260df0beb4