Pith. sign in

Paper Citation Record · LEDGER

How to Evaluate Reward Models for RLHF

As of 5 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 4 inbound Pith citation observations for arXiv:2410.14872.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.14872 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 4 of 4 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T08:15:17.876896Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-23T06:25:27.821937Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 1be5e7df-2b31-4e5b-af58-5138884f6000 · inbound

Qwen2.5 Technical Report cites this paper.

Qwen2.5 Technical Report How to Evaluate Reward Models for RLHF

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-23T06:25:27.824781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T06:25:00.376073Z digest=sha256:13396a95d5c81708fb5ab7948b13afeb4976aacfc9b4b81f5ee0094387bb3acf

Observation 32242d9b-6592-4ddc-8418-e2dfbe177a3e · inbound

RewardBench 2: Advancing Reward Model Evaluation cites this paper.

RewardBench 2: Advancing Reward Model Evaluation How to Evaluate Reward Models for RLHF

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T11:22:16.817153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:18:03.965711Z digest=sha256:9787661d3cc71d7a9e1747e1fc6cb610745fa4c224d61e51b7a3f863e92a997c

Observation 03dbc1b8-32d0-41cb-8622-38e8d95efb1e · inbound

A Survey on Evaluating Quality and Trustworthiness in LLM-Generated Data cites this paper.

A Survey on Evaluating Quality and Trustworthiness in LLM-Generated Data How to Evaluate Reward Models for RLHF

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-03T08:15:17.876896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T08:15:17.876896Z digest=sha256:31abbda95bce2f3efd764c892497bdbf43472e816b79c9392ce949276b167d9d

Observation 195b989d-6a45-47a2-9d0e-40a5a1bab2c3 · inbound

CompliBench: Benchmarking LLM Judges for Compliance Violation Detection in Dialogue Systems cites this paper.

CompliBench: Benchmarking LLM Judges for Compliance Violation Detection in Dialogue Systems How to Evaluate Reward Models for RLHF

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T11:26:02.353928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T14:56:00.449776Z digest=sha256:3cc20be15d34ea159a2f24652b8948ed6e60ea42ccfd6825f5697e8efc0c4f5d