Pith. sign in

Paper Citation Record · LEDGER

Critique-out-Loud Reward Models

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 34 inbound Pith citation observations for arXiv:2408.11791.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2408.11791 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 34 of 34 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 34 of 34 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:09:54.211281Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 29829418-edc8-4424-97e7-1a55043c0341 · inbound

A Survey on LLM-as-a-Judge cites this paper.

A Survey on LLM-as-a-Judge Critique-out-Loud Reward Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-23T17:35:44.196635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-23T17:33:13.394338Z digest=sha256:de0bd7244c94f6fc1579ea458b0a53aa17255d3ea7a0f44d1e151c5b0d8177c8

Observation 4f674538-c9d1-4632-97fa-bcf9647e8a61 · inbound

Reinforcement Learning from Human Feedback cites this paper.

Reinforcement Learning from Human Feedback Critique-out-Loud Reward Models

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-05-22T19:32:01.313460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T19:27:40.991325Z digest=sha256:ec9d26a6d1cd3a68077256c666c1a9e4e7517b35af33847cf105dc2a8d5fc7c6

Observation 73873e99-f590-43c3-8f88-36992110828b · inbound

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models cites this paper.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Critique-out-Loud Reward Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:54.211281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:54.211281Z digest=sha256:3c3dd3e973932445125a9b5879928f41251eaa74c2f045906520d4c8f2994dbc

Observation 053d29f9-a696-4f8b-bf2c-3e99b8a48106 · inbound

Generative RLHF-V: Learning Principles from Multi-modal Human Preference cites this paper.

Generative RLHF-V: Learning Principles from Multi-modal Human Preference Critique-out-Loud Reward Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:53.519748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:53.519748Z digest=sha256:71579743a4cacebf0f2ee8da21c29c292f1ea2da2dfa374e785ba52b0c996772

Observation 387c1d45-c326-44e0-a065-300f432830c8 · inbound

When Slower Isn't Truer: Inverse Scaling Law of Truthfulness in Multimodal Reasoning cites this paper.

When Slower Isn't Truer: Inverse Scaling Law of Truthfulness in Multimodal Reasoning Critique-out-Loud Reward Models

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-19T12:52:18.131065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T12:47:36.228366Z digest=sha256:52eed1acafac4e3d03965d743e21e9e7a5bd9ed5835a066e4853330514bb57b4

Observation 50165dfc-02ec-42c5-b4e2-4589e29b9d64 · inbound

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness cites this paper.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Critique-out-Loud Reward Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:41.917221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:41.917221Z digest=sha256:556e832ddff1a04770b147ffca06bdfca4b432d2328ce200d326e8d421e1c84b

Observation 61de2095-0d17-444c-88a6-01b3badfa06f · inbound

Socratic-PRMBench: Benchmarking Process Reward Models with Systematic Reasoning Patterns cites this paper.

Socratic-PRMBench: Benchmarking Process Reward Models with Systematic Reasoning Patterns Critique-out-Loud Reward Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:08.718847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:08.718847Z digest=sha256:5795e3535e8dc10922d4a958e7f976122403f44974421443135deac9a9ab7eca

Observation 90199c7f-2665-4981-b314-6dc5fbad52c5 · inbound

Datasheets Aren't Enough: DataRubrics for Automated Quality Metrics and Accountability cites this paper.

Datasheets Aren't Enough: DataRubrics for Automated Quality Metrics and Accountability Critique-out-Loud Reward Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:36:43.442054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:36:43.442054Z digest=sha256:af6099a93cf7a9d896d0c75d8f86eefdaf5b58113c6ac55af1cf6f30f73ff06e

Observation 3a461104-60f6-4eef-8872-5c714efec7d4 · inbound

RewardBench 2: Advancing Reward Model Evaluation cites this paper.

RewardBench 2: Advancing Reward Model Evaluation Critique-out-Loud Reward Models

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:22:16.798084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T11:18:03.965711Z digest=sha256:8fe3f7a96e9fbf377374cc2c42248747a62ae46e24f66ab9fb36775559b61bb9

Observation e2782c99-9211-4367-ad11-4a63a5f59a4f · inbound

RewardAnything: Generalizable Principle-Following Reward Models cites this paper.

RewardAnything: Generalizable Principle-Following Reward Models Critique-out-Loud Reward Models

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:07.144776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:07.144776Z digest=sha256:036cb5641a30f7e500dafc3120c13f6c5f4bce3b960f0453fa8927203877e514

Observation 47b801cc-3afb-4edc-95fa-edd60c593ab1 · inbound

Unlocking Recursive Thinking of LLMs: Alignment via Refinement cites this paper.

Unlocking Recursive Thinking of LLMs: Alignment via Refinement Critique-out-Loud Reward Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:01.983009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:01.983009Z digest=sha256:de4f06004957ad3e42286afde13caef5a921b147760258aac768db0542ae550a

Observation b7185dab-cc29-44f0-aed7-77af05663516 · inbound

CyberV: Cybernetics for Test-time Scaling in Video Understanding cites this paper.

CyberV: Cybernetics for Test-time Scaling in Video Understanding Critique-out-Loud Reward Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:46.991810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:26:46.991810Z digest=sha256:3f559c04ae4d735dd7e6b612aa47048b3f4635c9e978841c66832aa614278271

Observation 88600e08-3aad-440a-b8fe-7493dd2e2d5f · inbound

A Survey on Large Language Models for Mathematical Reasoning cites this paper.

A Survey on Large Language Models for Mathematical Reasoning Critique-out-Loud Reward Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T05:14:44.256223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:14:44.256223Z digest=sha256:861574aef356b796650948e8ba3340e254e1904a0ece27071a317114e7ef3e90

Observation 717b82e3-abeb-4939-8f27-cbe14a2d3d99 · inbound

GFRIEND: Generative Few-shot Reward Inference through EfficieNt DPO cites this paper.

GFRIEND: Generative Few-shot Reward Inference through EfficieNt DPO Critique-out-Loud Reward Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:59.046421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:59.046421Z digest=sha256:d76c639ee9732c866401bbaa4acd5b6ede3159a57e89a0a3cb33c7172c3b2a69

Observation 54cf2e96-04ce-4fd7-abac-42db1f557aa2 · inbound

Large Reasoning Models are not thinking straight: on the unreliability of thinking trajectories cites this paper.

Large Reasoning Models are not thinking straight: on the unreliability of thinking trajectories Critique-out-Loud Reward Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T21:13:12.739806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:13:12.739806Z digest=sha256:80b51b0ba06e4b5f38a30c7aadeaa0ed2091b6cefb191b64c9c01ba684c48edd

Observation f312f2fd-44f8-487b-a850-46a7ca56ff50 · inbound

Breaking the Myth: Can Small Models Infer Postconditions Too? cites this paper.

Breaking the Myth: Can Small Models Infer Postconditions Too? Critique-out-Loud Reward Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T17:41:41.971957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:41:41.971957Z digest=sha256:d91d7932e7c76d652867ce7740e2dd94abe949bf7a655ea4e6d993588cb3ab27

Observation 6f7414c9-713f-489a-a2f9-ed4b1c21aeac · inbound

RefCritic: Training Long Chain-of-Thought Critic Models with Refinement Feedback cites this paper.

RefCritic: Training Long Chain-of-Thought Critic Models with Refinement Feedback Critique-out-Loud Reward Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T15:47:30.588128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:47:30.588128Z digest=sha256:bd92477ed37506276b9cadd54322ee5b258cf607e34ea9e32904609bef807f84

Observation 1c6ce263-51aa-46b7-9c5f-c73c385aa4e9 · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models Critique-out-Loud Reward Models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:02:24.845131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:f8cb4df3fbcf847eb5a096c50635c74d60838ed1b3d33228aba9b80a0dc99c85

Observation 2871b6e8-113e-4917-92d1-70c386faed27 · inbound

PaTaRM: Bridging Pairwise and Pointwise Signals via Preference-Aware Task-Adaptive Reward Modeling cites this paper.

PaTaRM: Bridging Pairwise and Pointwise Signals via Preference-Aware Task-Adaptive Reward Modeling Critique-out-Loud Reward Models

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T03:20:49.228790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T03:15:57.744706Z digest=sha256:1f85e70a7e2e03fbff00bc5b768e199ebdec58e427a57bbb6ce40b69b5131857

Observation 2fd3859e-1fec-4c0e-ad7c-7d64b65f3511 · inbound

Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation cites this paper.

Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation Critique-out-Loud Reward Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T03:04:43.039235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:04:43.039235Z digest=sha256:a38fd996d9683efb7df8c30110f808537c44aba2219869797efe3128d6e647d0

Observation 1b3a7623-2bc3-4be7-863b-fb344aa79ecf · inbound

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges cites this paper.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Critique-out-Loud Reward Models

Reference 132

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:00:28.649424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:a6e2b84f9822c50a906268da8eb944bcf320f317868606b7257ec9958f6481f8

Observation 94806a1e-9d9c-4889-be33-349810620731 · inbound

Text-to-Distribution Prediction with Quantile Tokens and Neighbor Context cites this paper.

Text-to-Distribution Prediction with Quantile Tokens and Neighbor Context Critique-out-Loud Reward Models

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:59:48.734938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T00:58:34.750132Z digest=sha256:bd6584422e175be681120ad9579ae2c6d5cdd05ca98a99b6c674246ad406b0c3

Observation e9eedcf7-ea0e-4e25-a36b-53136ad5baff · inbound

Building a Precise Video Language with Human-AI Oversight cites this paper.

Building a Precise Video Language with Human-AI Oversight Critique-out-Loud Reward Models

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:04.566261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:5ab8932fec0d79263359ee6d87784146e9643837abe4828fd9e9b0ae85584eae

Observation c06dd0f4-de11-46f0-b672-6127bd5eda1f · inbound

POSTCONDBENCH: Benchmarking Correctness and Completeness in Formal Postcondition Inference cites this paper.

POSTCONDBENCH: Benchmarking Correctness and Completeness in Formal Postcondition Inference Critique-out-Loud Reward Models

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T23:56:12.299635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-07T16:04:48.394294Z digest=sha256:31d5e7ab533c7e707a9e15288fbc32dcdd191b099bee7cd16e992efc04f36430

Observation 3950b2ef-8a37-4881-a12a-ed5e522102e3 · inbound

Auto-Rubric as Reward: From Implicit Preferences to Explicit Multimodal Generative Criteria cites this paper.

Auto-Rubric as Reward: From Implicit Preferences to Explicit Multimodal Generative Criteria Critique-out-Loud Reward Models

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:36:24.356872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T00:58:53.626310Z digest=sha256:6664757a2bac6553bc5413b92d676f396471eef0ff6704542f0b107979a3fbfb

Observation 167c9dd6-8b63-46e6-9737-97b2558b2275 · inbound

Think Twice, Act Once: Verifier-Guided Action Selection For Embodied Agents cites this paper.

Think Twice, Act Once: Verifier-Guided Action Selection For Embodied Agents Critique-out-Loud Reward Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:19:28.896031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-14T21:04:05.443622Z digest=sha256:a5d817b3bf21fe78813c54519828e1520fc6b92da8924aa0a455bb00115f7dc9

Observation b610004c-d2ae-49c4-9b69-3783db5b7733 · inbound

Teaching Large Language Models When Not to Know: Learning Temporal Critique for Ex-Ante Reasoning cites this paper.

Teaching Large Language Models When Not to Know: Learning Temporal Critique for Ex-Ante Reasoning Critique-out-Loud Reward Models

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-06-30T20:55:04.641566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T20:46:31.368628Z digest=sha256:d432b495cf9d495dfb3c21551b291386c89572ebc5df8810d1a0ceb027a406be

Observation 32b7c71e-b8e2-43cf-9ba2-8038e38c8035 · inbound

Will It Go Viral? Grounding Micro-Video Popularity Prediction on the Open Web cites this paper.

Will It Go Viral? Grounding Micro-Video Popularity Prediction on the Open Web Critique-out-Loud Reward Models

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-20T00:57:54.234885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T00:55:59.857014Z digest=sha256:9dba09e01ee45290236522f5be8ef47d6b78812c2f507f670dd602624fe9037d

Observation 44aa9132-73d4-4734-a09e-d5693b32022d · inbound

STRIDE: Learnable Stepwise Language Feedback for LLM Reasoning cites this paper.

STRIDE: Learnable Stepwise Language Feedback for LLM Reasoning Critique-out-Loud Reward Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-20T20:49:00.763493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T20:47:16.236629Z digest=sha256:f3ab41fd74f6087b14f157bc240319763f3120ad5d96d19737126ca4e9204a53

Observation 60e72776-fd66-40b5-ad67-d78891cd5a0b · inbound

Weak Critics Make Strong Learners: On-Policy Critique Distillation for Scalable Oversight cites this paper.

Weak Critics Make Strong Learners: On-Policy Critique Distillation for Scalable Oversight Critique-out-Loud Reward Models

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T19:56:11.026051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T21:55:58.645062Z digest=sha256:a9743c4b3218ffa61ef0b214da0cb0bfd8faf67f7e577352a04b1a511a9bdf52

Observation ecdb5959-f451-4728-8209-00316980fc7a · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation Critique-out-Loud Reward Models

Reference 150

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T20:46:14.361106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:4d5fa98767bf8819fab1af498c514989540cbade63924ca267dc712bff8dd9fa

Observation e5fdd755-982f-4f9a-a992-a42aef3bd337 · inbound

Support Vector Rubrics: Closing the Gap Between Self-Generated and Human Rubrics cites this paper.

Support Vector Rubrics: Closing the Gap Between Self-Generated and Human Rubrics Critique-out-Loud Reward Models

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T20:47:23.356231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T20:06:12.523067Z digest=sha256:ded361278701566d603e58fce98c3e5aa921437735cbfae4307af0b9c389a58d

Observation 519f152e-7f22-4a9f-8f18-3e4c80e4fed9 · inbound

LLM-as-a-Coach: Experiential Learning for Non-Verifiable Tasks cites this paper.

LLM-as-a-Coach: Experiential Learning for Non-Verifiable Tasks Critique-out-Loud Reward Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T16:07:59.714666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T16:07:59.714666Z digest=sha256:bfbd0d7686d920e22512d607ee1cf094df48154d826c4de7f3dff03e5f7c58e4

Observation 485eac96-9550-4488-868a-acade5e44020 · inbound

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End cites this paper.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End Critique-out-Loud Reward Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:05.397076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:05.397076Z digest=sha256:c3c91e677dedfe087a34821c1c568b927ed77a2992f05c3931d0eda73f4c06e1