Pith. sign in

Paper Citation Record · LEDGER

Self-Generated Critiques Boost Reward Modeling for Language Models

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2411.16646.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.16646 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T18:23:50.481768Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T12:25:43.079840Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 4d8777be-64db-4555-bd5a-c50a6fe447bf · inbound

In Context Learning and Reasoning for Symbolic Regression with Large Language Models cites this paper.

In Context Learning and Reasoning for Symbolic Regression with Large Language Models Self-Generated Critiques Boost Reward Modeling for Language Models

Reference 95

Resolution
verified exact
arxiv_id, observed 2026-05-23T18:53:21.144572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-23T18:50:40.378720Z digest=sha256:40181634ed042af88c502f62171d0f752db72d9e2a4700e7a37f875c84d4832c

Observation 199020c2-2a2a-4000-8eb8-0509df32500a · inbound

Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs cites this paper.

Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs Self-Generated Critiques Boost Reward Modeling for Language Models

Reference 278

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T15:51:29.406469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-13T15:51:29.022336Z digest=sha256:a6bd58a60a94ffa7f9a6ffe95ba40e00435d6cb6dd8a70ca6c58433a28dddc6b

Observation d4229abe-225d-4a77-996c-286beb3372a5 · inbound

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment cites this paper.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Self-Generated Critiques Boost Reward Modeling for Language Models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:50.481768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:50.481768Z digest=sha256:469c2f23f6c9bfd6fe575048f8d247f6e81e8ddd43f4512560220a661756188a

Observation faf6a9ce-0c08-4f47-8de0-a77a1dec0b40 · inbound

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models cites this paper.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Self-Generated Critiques Boost Reward Modeling for Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:54.445894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:54.445894Z digest=sha256:b330ebb37fdcfcdeb0e8128f5dfcbffc5a380363c205b1144eeec96e402e476b

Observation a978274c-3e04-4b41-9561-e696d770c2d0 · inbound

RewardAnything: Generalizable Principle-Following Reward Models cites this paper.

RewardAnything: Generalizable Principle-Following Reward Models Self-Generated Critiques Boost Reward Modeling for Language Models

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:07.139918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:07.139918Z digest=sha256:23f9bae731a10057ea6be67c8d2ed66b71ac6810d221e1ad25841121f46af840

Observation 422ef1bf-506e-4adc-bdba-de513b822b03 · inbound

Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges cites this paper.

Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges Self-Generated Critiques Boost Reward Modeling for Language Models

Reference 282

Resolution
unresolved
no resolver link, observed 2026-08-06T14:13:07.294496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:13:07.294496Z digest=sha256:4e13890a4248e58988bd539747ed9ec1157a02815fb7a60c9a511b90f84baac4

Observation 58c58590-3bd8-48e6-b25b-a0b1448d2c79 · inbound

Building a Precise Video Language with Human-AI Oversight cites this paper.

Building a Precise Video Language with Human-AI Oversight Self-Generated Critiques Boost Reward Modeling for Language Models

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:04.461763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:43f5fae936d134aa75c94a6482b3fa3fbca6ead1f3b7a0fef04cb3b7d6bb85e2

Observation 0dfaad6f-0025-45e8-99fb-15e8ced842ba · inbound

Test-Time Verification for Text-to-SQL via Outcome Reward Models cites this paper.

Test-Time Verification for Text-to-SQL via Outcome Reward Models Self-Generated Critiques Boost Reward Modeling for Language Models

Reference 83

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T12:25:43.081624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-07-01T02:10:13.533450Z digest=sha256:4d0e21c4041f5dcc55402ee1ae820dfac221a0e5310eef0e0467f9bbb3a34d64