Pith. sign in

Paper Citation Record · LEDGER

Principled Reinforcement Learning with Human Feedback from Pairwise or $K$-wise Comparisons

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2301.11270.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2301.11270 v5

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:33:55.702118Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T13:09:51.317186Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 8b16b78e-9cf1-419d-94fc-e7029f1716da · inbound

FisherSFT: Data-Efficient Supervised Fine-Tuning of Language Models Using Information Gain cites this paper.

FisherSFT: Data-Efficient Supervised Fine-Tuning of Language Models Using Information Gain Principled Reinforcement Learning with Human Feedback from Pairwise or $K$-wise Comparisons

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:55.702118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:33:55.702118Z digest=sha256:9ad4be036ce67c474c9bc4c05a9f5fa567037166745df4f0f78defbefec55c05

Observation 11729123-2e3c-4b59-a987-8f59a4fd8bb4 · inbound

Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges cites this paper.

Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges Principled Reinforcement Learning with Human Feedback from Pairwise or $K$-wise Comparisons

Reference 156

Resolution
unresolved
no resolver link, observed 2026-08-06T14:13:06.505284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:13:06.505284Z digest=sha256:1b0577195e2d9112452c77333b5d83f122c6321ec51a040676c155ff2f902c9b

Observation 27248681-92ff-4935-a2c9-6dac5146875f · inbound

On the optimization dynamics of RLVR: Gradient gap and step size thresholds cites this paper.

On the optimization dynamics of RLVR: Gradient gap and step size thresholds Principled Reinforcement Learning with Human Feedback from Pairwise or $K$-wise Comparisons

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-18T08:36:07.323261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T08:34:36.543874Z digest=sha256:69e3591131220a79c6f89d89fd659478d1586e4c8f735e4cd5141bb465a0c80b

Observation 31e02f26-4df5-4e55-a224-5a10e8999d7d · inbound

Minimizing the Hidden Cost of Scales: Graph-Guided Ultra-Low-Bit Quantization for Large Language Models cites this paper.

Minimizing the Hidden Cost of Scales: Graph-Guided Ultra-Low-Bit Quantization for Large Language Models Principled Reinforcement Learning with Human Feedback from Pairwise or $K$-wise Comparisons

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-02T08:16:48.231712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T06:09:42.838355Z digest=sha256:f07a7b495d2174b4e5a7fa0d1cb7c062e3cf4217a42e364de1ad967ffcfb84e7

Observation dd5c430b-a540-42ca-b56a-c11c3f90c757 · inbound

Finding Stationary Points by Comparisons cites this paper.

Finding Stationary Points by Comparisons Principled Reinforcement Learning with Human Feedback from Pairwise or $K$-wise Comparisons

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T13:09:51.318748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T05:26:07.720218Z digest=sha256:d0a0784a8b7af44eab183fe7ed692002caa20bedd09249ef9a3f7bb3a5725638

Observation db8ed887-2c8a-4b74-8092-fc3333feb96e · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Principled Reinforcement Learning with Human Feedback from Pairwise or $K$-wise Comparisons

Reference 265

Resolution
unresolved
no resolver link, observed 2026-07-11T13:53:36.775836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T13:53:36.775836Z digest=sha256:c07cc9b79d084af967e81952d0b5776021ab636babe500308dfbc5acae2cdd7e

Observation 0521904d-992f-4593-9abf-b00cc1bb4a9c · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Principled Reinforcement Learning with Human Feedback from Pairwise or $K$-wise Comparisons

Reference 266

Resolution
unresolved
no resolver link, observed 2026-08-02T08:41:03.373334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:41:03.373334Z digest=sha256:d5773c3c15d6f51051de57a85a0f85ce8fa6abd01dcb1d4e1bd49dce7eb3d796