Pith. sign in

Paper Citation Record · LEDGER

A Comprehensive Survey of Reward Models: Taxonomy, Applications, Challenges, and Future

As of 4 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2504.12328.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.12328 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T05:56:17.734874Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T05:46:41.076602Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 4f5f303e-8ec3-4e6a-8cec-2d4640f9a9f7 · inbound

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey cites this paper.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey A Comprehensive Survey of Reward Models: Taxonomy, Applications, Challenges, and Future

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:21:48.183065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:a60bddf07ce96b1436650268ad52289559344ad1c622be742b4510538968ae55

Observation 91b27b97-6836-4701-b574-f0eba351c387 · inbound

Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation cites this paper.

Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation A Comprehensive Survey of Reward Models: Taxonomy, Applications, Challenges, and Future

Reference 125

Resolution
unresolved
no resolver link, observed 2026-08-03T03:04:45.216822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:04:45.216822Z digest=sha256:d2af139d4b5084283c025ad143b2ec2e38162b7732a8ae674c1d52c45cb20f3e

Observation c9f226c9-a0cd-4f28-9a2f-94e4b5493b28 · inbound

Toward Robust LLM-Based Judges: Taxonomic Bias Evaluation and Debiasing Optimization cites this paper.

Toward Robust LLM-Based Judges: Taxonomic Bias Evaluation and Debiasing Optimization A Comprehensive Survey of Reward Models: Taxonomy, Applications, Challenges, and Future

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T05:56:17.734874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:56:17.734874Z digest=sha256:7285a5981f2625808efb26aa69d1a881506559bfe8ffae580cdab0aca9c42073

Observation b80ce6e4-1d60-4892-af25-9ed48cca04e3 · inbound

StoryAlign: Evaluating and Training Reward Models for Story Generation cites this paper.

StoryAlign: Evaluating and Training Reward Models for Story Generation A Comprehensive Survey of Reward Models: Taxonomy, Applications, Challenges, and Future

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:31:07.645525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T17:29:13.549559Z digest=sha256:bcc5b9ba4327ea08191e4e2768932cb5bec3824b3dd1051b03a1e98d9eb66e52

Observation 35a22826-9bed-4c54-ad0b-4b26450c99e0 · inbound

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification cites this paper.

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification A Comprehensive Survey of Reward Models: Taxonomy, Applications, Challenges, and Future

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:01:24.230952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T04:41:44.833354Z digest=sha256:5b64e81e87c8e3ba1bf6ca85a4c0003c851c46e71c20170f66070c6d27710149

Observation 39a11b7d-857e-4c3a-8c55-237c94b4d1fd · inbound

GEAR: Granularity-Adaptive Advantage Reweighting for LLM Agents via Self-Distillation cites this paper.

GEAR: Granularity-Adaptive Advantage Reweighting for LLM Agents via Self-Distillation A Comprehensive Survey of Reward Models: Taxonomy, Applications, Challenges, and Future

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:47:26.720310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T06:44:30.069444Z digest=sha256:18b713a49168aebfddea57a8def11ac0c6a5240e69b4e4de7acba8e1826ca782

Observation f89c9da0-0fbe-4a96-a8a7-c30ff9d84c95 · inbound

GEAR: Granularity-Adaptive Advantage Reweighting for LLM Agents via Self-Distillation cites this paper.

GEAR: Granularity-Adaptive Advantage Reweighting for LLM Agents via Self-Distillation A Comprehensive Survey of Reward Models: Taxonomy, Applications, Challenges, and Future

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:59:48.637981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T05:59:23.949124Z digest=sha256:6ec7c41c1f05dca608517c70fac3828b0c7fadf6e50878aea407f47922f4bcc3

Observation 900a3756-db76-42c6-9c21-a578f0cea061 · inbound

Scalable Token-Level Hallucination Detection in Large Language Models cites this paper.

Scalable Token-Level Hallucination Detection in Large Language Models A Comprehensive Survey of Reward Models: Taxonomy, Applications, Challenges, and Future

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:52:22.393010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:49:23.534294Z digest=sha256:fdeb4f6c3f876e3bf81d816228a5810e8d8bf13100a56b9309b8ba5374f3e35d

Observation 61303caa-455b-4c9e-92bd-11e43e669e5c · inbound

SocialCoach: Personalized Social Skill Learning with RL-based Agentic Tutoring and Practice cites this paper.

SocialCoach: Personalized Social Skill Learning with RL-based Agentic Tutoring and Practice A Comprehensive Survey of Reward Models: Taxonomy, Applications, Challenges, and Future

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:46:41.078116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T08:04:17.224987Z digest=sha256:6ea9ac216b1e0ce8e47c2f10e662bcf44cb225624317e325dee64007ca0848dd

Observation 0b1b94c1-dcdb-4cd0-8d61-8945a6a2d2d5 · inbound

Agents Don't Just Agree, They Remember: Benchmarking Persistent Sycophancy in Stateful Personal Agents cites this paper.

Agents Don't Just Agree, They Remember: Benchmarking Persistent Sycophancy in Stateful Personal Agents A Comprehensive Survey of Reward Models: Taxonomy, Applications, Challenges, and Future

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-14T11:04:44.592375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T11:04:44.592375Z digest=sha256:cb8fb5cb8ffa224122dc01045157c8678f7a9563b8af8264f0e6ad63cbfc135c