Pith. sign in

Paper Citation Record · LEDGER

Principled Reinforcement Learning with Human Feedback from Pairwise or $K$-wise Comparisons

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2301.11270.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2301.11270 v5

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:33:55.702118Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T13:09:51.317186Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 8b16b78e-9cf1-419d-94fc-e7029f1716da · inbound

FisherSFT: Data-Efficient Supervised Fine-Tuning of Language Models Using Information Gain cites this paper.

FisherSFT: Data-Efficient Supervised Fine-Tuning of Language Models Using Information Gain Principled Reinforcement Learning with Human Feedback from Pairwise or $K$-wise Comparisons

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:55.702118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:33:55.702118Z digest=sha256:4cd19547fcfdf5213ca09b6502e11abbace7d29b5d2c1b4788308d40befc5b85

Observation 11729123-2e3c-4b59-a987-8f59a4fd8bb4 · inbound

Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges cites this paper.

Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges Principled Reinforcement Learning with Human Feedback from Pairwise or $K$-wise Comparisons

Reference 156

Resolution
unresolved
no resolver link, observed 2026-08-06T14:13:06.505284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:13:06.505284Z digest=sha256:48e38136612dbcb76e3825a95bf92abbce52de1fbd0d8d14bea48df83c575959

Observation 27248681-92ff-4935-a2c9-6dac5146875f · inbound

On the optimization dynamics of RLVR: Gradient gap and step size thresholds cites this paper.

On the optimization dynamics of RLVR: Gradient gap and step size thresholds Principled Reinforcement Learning with Human Feedback from Pairwise or $K$-wise Comparisons

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-18T08:36:07.323261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T08:34:36.543874Z digest=sha256:9b6ee254798104850d2214a6bd799aebfb9fdd5e16bc9ffaacdab2e3900f6ec2

Observation 31e02f26-4df5-4e55-a224-5a10e8999d7d · inbound

Minimizing the Hidden Cost of Scales: Graph-Guided Ultra-Low-Bit Quantization for Large Language Models cites this paper.

Minimizing the Hidden Cost of Scales: Graph-Guided Ultra-Low-Bit Quantization for Large Language Models Principled Reinforcement Learning with Human Feedback from Pairwise or $K$-wise Comparisons

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-02T08:16:48.231712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T06:09:42.838355Z digest=sha256:2e1ef4ee83b2c2408f7c9ec94af19ba393ec11d573ca8240043d6e554a01407a

Observation dd5c430b-a540-42ca-b56a-c11c3f90c757 · inbound

Finding Stationary Points by Comparisons cites this paper.

Finding Stationary Points by Comparisons Principled Reinforcement Learning with Human Feedback from Pairwise or $K$-wise Comparisons

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T13:09:51.318748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T05:26:07.720218Z digest=sha256:219b61a1ef3a98cedf21f7ec5e58cecd04d0c52b23f7335045413b9a1e38fb6e

Observation db8ed887-2c8a-4b74-8092-fc3333feb96e · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Principled Reinforcement Learning with Human Feedback from Pairwise or $K$-wise Comparisons

Reference 265

Resolution
unresolved
no resolver link, observed 2026-07-11T13:53:36.775836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T13:53:36.775836Z digest=sha256:36bbb5615f083e1fbad8c6ec7e689c2c63d8d395f36ccdb01877249546c74767

Observation 0521904d-992f-4593-9abf-b00cc1bb4a9c · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Principled Reinforcement Learning with Human Feedback from Pairwise or $K$-wise Comparisons

Reference 266

Resolution
unresolved
no resolver link, observed 2026-08-02T08:41:03.373334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:41:03.373334Z digest=sha256:28653c6e14a972a4a2edca5ac2e210f1f01e6cfdf910c7abe16bc17a2f431b55