Pith. sign in

Paper Citation Record · LEDGER

CriticBench: Benchmarking LLMs for Critique-Correct Reasoning

As of 19 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 13 inbound Pith citation observations for arXiv:2402.14809.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.14809 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:33:53.442371Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d88a59ea-bd8b-4bac-b3bb-283a39dd8045 · inbound

ProcessBench: Identifying Process Errors in Mathematical Reasoning cites this paper.

ProcessBench: Identifying Process Errors in Mathematical Reasoning CriticBench: Benchmarking LLMs for Critique-Correct Reasoning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T19:37:21.572926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:37:21.572926Z digest=sha256:74d23b3a3ce871ff63565e72626ae66a6a1575402e13987762521bdda26dd1c0

Observation e4b7a4b8-e042-4425-be3a-058db3c31911 · inbound

Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training cites this paper.

Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training CriticBench: Benchmarking LLMs for Critique-Correct Reasoning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T18:22:02.533704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:22:02.533704Z digest=sha256:2f696b2985223ed1ed3ece080d4c6ec15be54a3b4ce8f46509dd73e4b96eb6b1

Observation 68ea7ff4-bd91-4c7b-ac97-382d98d49848 · inbound

Improving Video Generation with Human Feedback cites this paper.

Improving Video Generation with Human Feedback CriticBench: Benchmarking LLMs for Critique-Correct Reasoning

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-13T15:30:02.737394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-13T15:30:02.578430Z digest=sha256:2b361c7909ef309b8e69a784e3a543e6d14bcc0bc68ae57497cfd8c4fa256301

Observation 6acbb55f-3aed-4f68-b65f-083e46dd1ea8 · inbound

LLMs can be easily Confused by Instructional Distractions cites this paper.

LLMs can be easily Confused by Instructional Distractions CriticBench: Benchmarking LLMs for Critique-Correct Reasoning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-09T10:50:12.337684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:50:12.337684Z digest=sha256:cd239461a8bba4c838872d6853241df555f5a5f0013923790753ed30148f45ea

Observation e31cffcc-7df2-4230-b851-9b85a6631491 · inbound

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators cites this paper.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators CriticBench: Benchmarking LLMs for Critique-Correct Reasoning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.442371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.442371Z digest=sha256:edecb484325c22ff04d816c9a1a69fbb50d62a350ab2a20675ae598b32b38579

Observation 7a419c50-c06a-4eac-9b93-1900ea188b43 · inbound

Step-Wise Formal Verification for LLM-Based Mathematical Problem Solving cites this paper.

Step-Wise Formal Verification for LLM-Based Mathematical Problem Solving CriticBench: Benchmarking LLMs for Critique-Correct Reasoning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:56.468279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:48:56.468279Z digest=sha256:1ecd214cab3df848d7a5273db92678c85ca16a5a5fddc550428e7c4957821921

Observation 1bca06c8-7f32-4bb1-a072-615a8862484e · inbound

From Standalone LLMs to Integrated Intelligence: A Survey of Compound Al Systems cites this paper.

From Standalone LLMs to Integrated Intelligence: A Survey of Compound Al Systems CriticBench: Benchmarking LLMs for Critique-Correct Reasoning

Reference 97

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:52:16.201852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T11:49:36.574471Z digest=sha256:18172c9f71125f5bb965212e0d71de2388baeab34ce5654a1212410981ae3452

Observation 7bf96a30-5c1b-49ba-9e5c-4cc6c5e0f159 · inbound

VIS-Shepherd: Constructing Critic for LLM-based Data Visualization Generation cites this paper.

VIS-Shepherd: Constructing Critic for LLM-based Data Visualization Generation CriticBench: Benchmarking LLMs for Critique-Correct Reasoning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T00:41:33.813200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:41:33.813200Z digest=sha256:61416f59bbca8dff7697979ace091ef7bab91d2d7f118cfd5f905829f28a33e0

Observation 9a806e1b-7c91-41a8-8c31-fd9bfdcef0b0 · inbound

CriticLean: Critic-Guided Reinforcement Learning for Mathematical Formalization cites this paper.

CriticLean: Critic-Guided Reinforcement Learning for Mathematical Formalization CriticBench: Benchmarking LLMs for Critique-Correct Reasoning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T19:14:16.816999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:14:16.816999Z digest=sha256:c1541ad5bc237021c997db07ed2aca5e9ed1578e0bf6258dbff2e77b6a47aa2a

Observation 6d143611-5771-45c2-b162-1a9dcf1eea7c · inbound

RefCritic: Training Long Chain-of-Thought Critic Models with Refinement Feedback cites this paper.

RefCritic: Training Long Chain-of-Thought Critic Models with Refinement Feedback CriticBench: Benchmarking LLMs for Critique-Correct Reasoning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T15:47:31.805450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:47:31.805450Z digest=sha256:e5e749aef38ffe3a1139a36e51aa9c9e1816c7b0c44f4ae3cc6945b11b06e2c5

Observation daf54d4c-c5c4-4190-b04a-cbd843e76f0b · inbound

Can Large Multimodal Models Actively Recognize Faulty Inputs? A Systematic Evaluation Framework of Their Input Scrutiny Ability cites this paper.

Can Large Multimodal Models Actively Recognize Faulty Inputs? A Systematic Evaluation Framework of Their Input Scrutiny Ability CriticBench: Benchmarking LLMs for Critique-Correct Reasoning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T01:00:27.481589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T01:00:27.481589Z digest=sha256:9f62295f5de983f75728ef7e752e0d028c92fae30e60cf03005590e0721f33ba

Observation 2086b781-cc8b-44aa-9d5e-0e6140c01720 · inbound

DR-MMSearchAgent: Deepening Reasoning in Multimodal Search Agents cites this paper.

DR-MMSearchAgent: Deepening Reasoning in Multimodal Search Agents CriticBench: Benchmarking LLMs for Critique-Correct Reasoning

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T12:31:07.916093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-10T03:21:30.732925Z digest=sha256:1675612f55ea90931d7311a494d5b1a09bddc29462f4eefd9fe0c695f0c88ff0

Observation 2578d05a-2c2a-45e4-9621-83c45bd0c98d · inbound

QVal: Cheaply Evaluating Dense Supervision Signals for Long-Horizon LLM Agents cites this paper.

QVal: Cheaply Evaluating Dense Supervision Signals for Long-Horizon LLM Agents CriticBench: Benchmarking LLMs for Critique-Correct Reasoning

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T06:05:29.030156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T05:59:13.631078Z digest=sha256:e399263474f94c3140d4184ef8f6d28ece9a93f19a39ca49f5c4165cbbb2594c