Pith. sign in

Paper Citation Record · LEDGER

Agentic Reward Modeling: Integrating Human Preferences with Verifiable Correctness Signals for Reliable Reward Systems

As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2502.19328.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.19328 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-22T21:39:49.832151Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T21:42:10.983751Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 9fc06d98-347d-431a-a51f-739512f98e4e · inbound

From System 1 to System 2: A Survey of Reasoning Large Language Models cites this paper.

From System 1 to System 2: A Survey of Reasoning Large Language Models Agentic Reward Modeling: Integrating Human Preferences with Verifiable Correctness Signals for Reliable Reward Systems

Reference 169

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:36:24.198198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:36:23.845366Z digest=sha256:008ec5f6bae10fac6e76f1040f45caa5eb0983392461645c7d7e8e6cff500c50

Observation 791788dc-5c19-4aed-82c4-c4bb840c344e · inbound

Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems cites this paper.

Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems Agentic Reward Modeling: Integrating Human Preferences with Verifiable Correctness Signals for Reliable Reward Systems

Reference 132

Resolution
verified exact
arxiv_id, observed 2026-05-22T21:42:10.986459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T21:39:49.832151Z digest=sha256:81af0f32be1634dd2f1d5688897bf66d8e8ba12fe648eec157f9b4f9b864944d

Observation dd73b719-13d4-43ea-a8c9-0f9e40c8d112 · inbound

Visual-ERM: Reward Modeling for Visual Equivalence cites this paper.

Visual-ERM: Reward Modeling for Visual Equivalence Agentic Reward Modeling: Integrating Human Preferences with Verifiable Correctness Signals for Reliable Reward Systems

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-15T11:19:57.899096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T11:19:42.002790Z digest=sha256:af7cc1fc53d255441b558834aaa35927fa6c33e039d5e3025ddaea36f6ae9810

Observation 157d32dd-a0d8-4df5-9ddb-b096e4f19198 · inbound

SPREG: Structured Plan Repair with Entropy-Guided Test-Time Intervention for Large Language Model Reasoning cites this paper.

SPREG: Structured Plan Repair with Entropy-Guided Test-Time Intervention for Large Language Model Reasoning Agentic Reward Modeling: Integrating Human Preferences with Verifiable Correctness Signals for Reliable Reward Systems

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T11:51:02.791253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-10T04:33:14.807721Z digest=sha256:a4ab7e6c053fa9cc9722a087b9ed96d76023567e86de93a4be661f8ca29bf80a

Observation 8189beed-aa38-406d-96fc-d76ca7036d29 · inbound

StoryAlign: Evaluating and Training Reward Models for Story Generation cites this paper.

StoryAlign: Evaluating and Training Reward Models for Story Generation Agentic Reward Modeling: Integrating Human Preferences with Verifiable Correctness Signals for Reliable Reward Systems

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:31:07.649828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T17:29:13.549559Z digest=sha256:61e566e7e0294b7ad87f1c8a02e138fd2a47664123db82af8097eb4a98de2e7a