Pith. sign in

Paper Citation Record · LEDGER

LLM Critics Help Catch LLM Bugs

As of 19 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 60 inbound Pith citation observations for arXiv:2407.00215.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.00215 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 60 of 60 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T04:42:07.455016Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

8
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 704ca59b-50c4-4424-a10c-a3888922ef0d · inbound

AIGS: Generating Science from AI-Powered Automated Falsification cites this paper.

AIGS: Generating Science from AI-Powered Automated Falsification LLM Critics Help Catch LLM Bugs

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-12T19:02:19.303413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:02:19.303413Z digest=sha256:d4433b6a1b7c941196ed7fd0656c289950464ef56e5682ec6236e4ea9ba24f5b

Observation 6bba815f-51b3-421d-a58c-56362845ad24 · inbound

Self-Generated Critiques Boost Reward Modeling for Language Models cites this paper.

Self-Generated Critiques Boost Reward Modeling for Language Models LLM Critics Help Catch LLM Bugs

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T12:58:29.816635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:58:29.816635Z digest=sha256:cc413dd1f5db3963dfaf88ba55cc4a52ea7b41cb8d854856bd1415f8b39ae40e

Observation 8d63ecfa-947c-4889-b4fe-80eec2bde86e · inbound

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning cites this paper.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning LLM Critics Help Catch LLM Bugs

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.259690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.259690Z digest=sha256:fd0c336352bffa208e8b129021863c0b3d2bcf7f2097d5207279a1640118ef58

Observation 4600b532-afe7-4fae-b2e0-e91787dd4a7c · inbound

ProcessBench: Identifying Process Errors in Mathematical Reasoning cites this paper.

ProcessBench: Identifying Process Errors in Mathematical Reasoning LLM Critics Help Catch LLM Bugs

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T19:37:21.583311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:37:21.583311Z digest=sha256:84990a475b3d6024b43dd3df9b501f4d5dcd167c5dcbfaff18cbadb03e0d5f2f

Observation 2b09ff6f-28a8-459d-a149-4e8803cf23fd · inbound

SpearBot: Leveraging Large Language Models in a Generative-Critique Framework for Spear-Phishing Email Generation cites this paper.

SpearBot: Leveraging Large Language Models in a Generative-Critique Framework for Spear-Phishing Email Generation LLM Critics Help Catch LLM Bugs

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T15:21:48.397078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:21:48.397078Z digest=sha256:560c30f6e2809593d28891dfecd319967731aed2f83bf344569482bb1d08f032

Observation b66629ff-e189-4614-bdf9-3f3a95ab5386 · inbound

The Superalignment of Superhuman Intelligence with Large Language Models cites this paper.

The Superalignment of Superhuman Intelligence with Large Language Models LLM Critics Help Catch LLM Bugs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T15:18:17.820149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:18:17.820149Z digest=sha256:4bcc52699ac3e66fff05a50f691cf77f0a1c4617ab4aa152f68393d88ca25883

Observation 4c7c6204-94bc-45af-85ed-6d0161a80de5 · inbound

Private Yet Social: How LLM Chatbots Support and Challenge Eating Disorder Recovery cites this paper.

Private Yet Social: How LLM Chatbots Support and Challenge Eating Disorder Recovery LLM Critics Help Catch LLM Bugs

Reference 104

Resolution
unresolved
no resolver link, observed 2026-08-11T14:45:10.817020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:45:10.817020Z digest=sha256:b1e5ffc245dbbef8e1b365f1dc86ead521f43cda41e29d237ea16b69d1093eb6

Observation 3960071c-58ae-468d-b6a3-da9ae9e73077 · inbound

Algebraic Evaluation Theorems cites this paper.

Algebraic Evaluation Theorems LLM Critics Help Catch LLM Bugs

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T11:58:27.866007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:58:27.866007Z digest=sha256:d070dc0f36e391dd0d4fcbc477f724de542c0698ca0bf39efa6e8258cc840748

Observation 5c75c41e-3a5a-4e49-b466-6577666deebf · inbound

Teaching LLMs to Refine with Tools cites this paper.

Teaching LLMs to Refine with Tools LLM Critics Help Catch LLM Bugs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T06:05:31.176033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T06:05:31.176033Z digest=sha256:c8d7af0a7a2e0d6cfc7d8298e4274e235d8f2db9c25c5010e3dc1cc42d0445df

Observation 7776990f-41bc-480f-abc3-1f265ee59a59 · inbound

Online Preference-based Reinforcement Learning with Self-augmented Feedback from Large Language Model cites this paper.

Online Preference-based Reinforcement Learning with Self-augmented Feedback from Large Language Model LLM Critics Help Catch LLM Bugs

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T06:05:40.211743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T06:05:40.211743Z digest=sha256:f9481bbdfd066224f6d5af730db4b2c77d1fa11d1da5c10bd8b08b39ad5195bb

Observation fa4b0705-096b-4dae-951b-3826112411cc · inbound

Distilling Desired Comments for Enhanced Code Review with Large Language Models cites this paper.

Distilling Desired Comments for Enhanced Code Review with Large Language Models LLM Critics Help Catch LLM Bugs

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T23:28:51.156470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:28:51.156470Z digest=sha256:be360b5940cd5fe0437c8bb5bcee8c868f88120aa8d2460df4dd494c6b5c5149

Observation 103d1241-3dc4-4e3c-89d8-bec9e63a76df · inbound

Iterative Label Refinement Matters More than Preference Optimization under Weak Supervision cites this paper.

Iterative Label Refinement Matters More than Preference Optimization under Weak Supervision LLM Critics Help Catch LLM Bugs

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:38.200106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:38.200106Z digest=sha256:59b1ce86c21d10c2f9cf819516b90e4a4da1fac7f8ade1b4d7d2201a63590474

Observation bfc1b0ac-95ab-4255-8260-d903ff18887b · inbound

PairJudge RM: Perform Best-of-N Sampling with Knockout Tournament cites this paper.

PairJudge RM: Perform Best-of-N Sampling with Knockout Tournament LLM Critics Help Catch LLM Bugs

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T16:39:31.221082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:39:31.221082Z digest=sha256:4f7709f635f24e49890774a9979d93c1b68a5f0e6d8026caa8cc4c440b3e9140

Observation d77895bb-706e-4f44-a70e-a052df001dbc · inbound

Debate Helps Weak-to-Strong Generalization cites this paper.

Debate Helps Weak-to-Strong Generalization LLM Critics Help Catch LLM Bugs

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T17:50:56.439642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:50:56.439642Z digest=sha256:4ea1eb230aa6d03480146bf33c41fbfab4632a9b25eef84ccbb959f1237d22e5

Observation 5cdd2b78-d9b6-4281-8027-413f78e3733b · inbound

BitsAI-CR: Automated Code Review via LLM in Practice cites this paper.

BitsAI-CR: Automated Code Review via LLM in Practice LLM Critics Help Catch LLM Bugs

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T14:40:07.190847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:40:07.190847Z digest=sha256:62764de7984ca4adfdf6f7983e81e173e0f88e36de6e389719f9792fb3297db6

Observation 2d120828-c66a-41d1-9a16-1800fc0e996a · inbound

LongDPO: Unlock Better Long-form Generation Abilities for LLMs via Critique-augmented Stepwise Information cites this paper.

LongDPO: Unlock Better Long-form Generation Abilities for LLMs via Critique-augmented Stepwise Information LLM Critics Help Catch LLM Bugs

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T13:25:52.035971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:25:52.035971Z digest=sha256:3b2f1cc5cc1f0a9ddf4ad1f5454e1b9909d083a320187b688d24b5fb252c021f

Observation 9bbd0236-5e0d-4e77-9466-a2c223a80c5e · inbound

Automated Capability Discovery via Foundation Model Self-Exploration cites this paper.

Automated Capability Discovery via Foundation Model Self-Exploration LLM Critics Help Catch LLM Bugs

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T12:17:55.260097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:17:55.260097Z digest=sha256:009bf63b4618c797bccc60c55e54c83f60c76b6dbd1f3752b2b4afec726d87aa

Observation ee956037-fb1a-4883-9db4-5f2edfa415a8 · inbound

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models cites this paper.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models LLM Critics Help Catch LLM Bugs

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:41:08.064965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:b3a39888c2e6fd67eade37023223c2f5c1d054866ecac265d76d95dec4bf6e14

Observation 16947f13-c97c-49be-be09-eb5680b7a1da · inbound

DeepCritic: Deliberate Critique with Large Language Models cites this paper.

DeepCritic: Deliberate Critique with Large Language Models LLM Critics Help Catch LLM Bugs

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T04:42:07.455016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:42:07.455016Z digest=sha256:e39a69fa187eb718ca4ac6b5ebdc2e9560194e08a3b304968d8d32e91bee1063

Observation ee70976a-dda7-4ce8-83d6-b3643658689c · inbound

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge cites this paper.

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge LLM Critics Help Catch LLM Bugs

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T20:50:10.606695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:50:10.606695Z digest=sha256:21caecd5a728c47d75b4250089d5a15ef52d33566d3408a566c6b151c83579ce

Observation 71262754-d894-4f4a-a2d8-3c63aa8ffb1a · inbound

Optimizing LLM-Based Multi-Agent System with Textual Feedback: A Case Study on Software Development cites this paper.

Optimizing LLM-Based Multi-Agent System with Textual Feedback: A Case Study on Software Development LLM Critics Help Catch LLM Bugs

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:12:36.041322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:12:36.041322Z digest=sha256:2312f4d7329dcedde8bc9137b192d5e39245a6a54588070116b1e04f2e39db80

Observation d664ad6a-ee01-489f-9aac-0f635ddb0cc3 · inbound

Breakpoint: Scalable evaluation of system-level reasoning in LLM code agents cites this paper.

Breakpoint: Scalable evaluation of system-level reasoning in LLM code agents LLM Critics Help Catch LLM Bugs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:17:44.997683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:17:44.997683Z digest=sha256:0a9dbf14b8d3501a5647030512a9957e6610e38bd1c9b1e9c74d5892ffaf0ce9

Observation b36489cf-d06b-4897-b3c9-75ceb9958269 · inbound

CRScore++: Reinforcement Learning with Verifiable Tool and AI Feedback for Code Review cites this paper.

CRScore++: Reinforcement Learning with Verifiable Tool and AI Feedback for Code Review LLM Critics Help Catch LLM Bugs

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:29.658971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:29.658971Z digest=sha256:d8e66ada6c9c569370176f1ae1854670f0b3a7f351ce3f3f102984be68adf883

Observation b97f700a-87c7-4fa0-8cb7-7840c538bd65 · inbound

Exchange of Perspective Prompting Enhances Reasoning in Large Language Models cites this paper.

Exchange of Perspective Prompting Enhances Reasoning in Large Language Models LLM Critics Help Catch LLM Bugs

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:00.096836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:00.096836Z digest=sha256:e54e3a2ad429715453cc0f69534c8758e0cabfe57043342116a907de7f0d8eb1

Observation 12932292-e91b-4517-9c01-14b43c65cf96 · inbound

A Survey of Automatic Evaluation Methods on Text, Visual and Speech Generations cites this paper.

A Survey of Automatic Evaluation Methods on Text, Visual and Speech Generations LLM Critics Help Catch LLM Bugs

Reference 261

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:47.148271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:47.148271Z digest=sha256:55266be998aabcbd33b0f9868fbfb874f65d6fe28923c883c9760fbce22fe811

Observation 02bc1dfe-46d6-417a-bbef-fef24ef561b9 · inbound

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios cites this paper.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios LLM Critics Help Catch LLM Bugs

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T04:44:03.053326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:44:03.053326Z digest=sha256:b8f40b862093fb1d7fe0a123e580dd5b8d84f9b8c3651dc2b341ac2b25caf5ef

Observation 7fae0959-5729-4815-b117-045361ba19bc · inbound

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset cites this paper.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset LLM Critics Help Catch LLM Bugs

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.764833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.764833Z digest=sha256:68b2b3a67bc9b83c500ea7f58dc20a4954c679cac84934269279549dedee9f64

Observation 6868fb9f-f436-4f7d-959a-d7d92e1ba5c1 · inbound

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks cites this paper.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks LLM Critics Help Catch LLM Bugs

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T17:32:32.445876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:32:32.445876Z digest=sha256:8da9b85222c85cbe056ad314b724045525886a4cd5552afd9e57bcfb5e14a6ab

Observation d82d28eb-19e1-4866-b18f-23d8360e1bd6 · inbound

Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety cites this paper.

Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety LLM Critics Help Catch LLM Bugs

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T14:19:44.746613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-20T14:19:44.695462Z digest=sha256:76bfe9d9028fc5af2e8a2337c8ceaaa049df7e103143ac03e32749a2f54bc1c9

Observation 395ff2a7-3a94-42ad-afc1-0e52bc8d9e71 · inbound

RefCritic: Training Long Chain-of-Thought Critic Models with Refinement Feedback cites this paper.

RefCritic: Training Long Chain-of-Thought Critic Models with Refinement Feedback LLM Critics Help Catch LLM Bugs

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T15:47:32.141323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:47:32.141323Z digest=sha256:0df1aad0774418f393cb8a4f1a8ffb74ff1720b2b0ea270d16a765badf192b3e

Observation 9adddc11-31b6-4e55-a087-ba9a64a6eb06 · inbound

CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning cites this paper.

CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning LLM Critics Help Catch LLM Bugs

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-21T23:25:45.268233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-21T23:24:43.556606Z digest=sha256:5bb45f5c7b52463604a561f56b2b86b2ba2939cb78874c67e38383ab808600b7

Observation da4951f9-e021-4c94-b175-b9086c68171d · inbound

ViseGPT: Towards Better Alignment of LLM-generated Data Wrangling Scripts and User Prompts cites this paper.

ViseGPT: Towards Better Alignment of LLM-generated Data Wrangling Scripts and User Prompts LLM Critics Help Catch LLM Bugs

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T05:46:52.666443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:46:52.666443Z digest=sha256:d641fc9876dc5af52f05439046a56a5f02d4310670a19f7fafd93aefb17a6c0e

Observation b0539d43-abd0-461c-8d1c-864fe2e38c9e · inbound

Are Today's LLMs Ready to Explain Well-Being Concepts? cites this paper.

Are Today's LLMs Ready to Explain Well-Being Concepts? LLM Critics Help Catch LLM Bugs

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T01:02:51.111485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T01:02:51.111485Z digest=sha256:b5eb6c1fe404f080b42735a0070c5b8a851661a766afc479a95f5902ca54e464

Observation 48faffc0-1935-46f9-a8d2-a314fba9378f · inbound

Let's Revise Step-by-Step: A Unified Local Search Framework for Code Generation with LLMs cites this paper.

Let's Revise Step-by-Step: A Unified Local Search Framework for Code Generation with LLMs LLM Critics Help Catch LLM Bugs

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-05T22:09:46.553605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:09:46.553605Z digest=sha256:de37ff62cd4f655d2d363897cd6fc762fa9b60b08b3e7cb4e1cc1ef27d495032

Observation dd190c0b-b0b9-4c20-bef5-218520ea3ce5 · inbound

Mind the Generation Process: Fine-Grained Confidence Estimation During LLM Generation cites this paper.

Mind the Generation Process: Fine-Grained Confidence Estimation During LLM Generation LLM Critics Help Catch LLM Bugs

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T17:30:45.807580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:30:45.807580Z digest=sha256:f2386de1a4bee2606cc2d5a13e34de00b5cc031ef17cba877bd74734a4360828

Observation e1347977-d5ef-4634-a2c0-d3068ce5e254 · inbound

Reinforcement Learning with Rubric Anchors cites this paper.

Reinforcement Learning with Rubric Anchors LLM Critics Help Catch LLM Bugs

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T17:22:55.449553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:22:55.449553Z digest=sha256:21e705ff978ab8f432d60e83f34d746e518a86086f47222c34a8fe6036a0765c

Observation 8a10b93e-dbd6-43be-9e52-76ceea82646f · inbound

OnGoal: Tracking and Visualizing Conversational Goals in Multi-Turn Dialogue with Large Language Models cites this paper.

OnGoal: Tracking and Visualizing Conversational Goals in Multi-Turn Dialogue with Large Language Models LLM Critics Help Catch LLM Bugs

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T14:38:06.861926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:38:06.861926Z digest=sha256:c12fe7d092db5cd4fd17a28d1726ede354d8767b46df52c5b74c204b55eccf47

Observation 007db2f9-132e-4e26-8ae4-18318fdfbb5d · inbound

Dream-Coder 7B: An Open Diffusion Language Model for Code cites this paper.

Dream-Coder 7B: An Open Diffusion Language Model for Code LLM Critics Help Catch LLM Bugs

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T12:56:35.136328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:56:35.136328Z digest=sha256:31d43761cbc255e915d30537c5047e1df43912fb70caf35e6da7c41ab573cbac

Observation 983d0e01-9175-409f-93ce-b594b6db0d16 · inbound

Human-AI Complementarity: A Goal for Amplified Oversight cites this paper.

Human-AI Complementarity: A Goal for Amplified Oversight LLM Critics Help Catch LLM Bugs

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:22.511988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:22.511988Z digest=sha256:db0c92c7768b6dde61f4b3cf7c694dc36e8154fbc6c65e8b84c1410ea2e59242

Observation b1730ac7-4802-48f8-9c5f-ae186195875e · inbound

No More Stale Feedback: Co-Evolving Critics for Open-World Agent Learning cites this paper.

No More Stale Feedback: Co-Evolving Critics for Open-World Agent Learning LLM Critics Help Catch LLM Bugs

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T16:03:04.279092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-16T16:01:48.789986Z digest=sha256:9927a8f8307035afad8757de206c5fc0aeac086c79ebc7fe505d95053e9ef7b2

Observation 28fe0286-729e-4c34-97ad-3e29aef28991 · inbound

ReCodeAgent: A Multi-agent Workflow for Language-Agnostic Translation and Validation of Large-Scale Repositories cites this paper.

ReCodeAgent: A Multi-agent Workflow for Language-Agnostic Translation and Validation of Large-Scale Repositories LLM Critics Help Catch LLM Bugs

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:41:36.059248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T17:29:37.642356Z digest=sha256:1fbb1288db20ef4dbedb1ab195fe646c45fd14c5ef90e7799594415b92a6bd80

Observation a5806370-9914-4158-9a1a-0b5fc8847c67 · inbound

Building a Precise Video Language with Human-AI Oversight cites this paper.

Building a Precise Video Language with Human-AI Oversight LLM Critics Help Catch LLM Bugs

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:04.557005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:b38d0bf1c79f30dce2307fcd547478b0d105cd5da77a89c1ed42ce1475d0da77

Observation 954bd355-01ed-4240-a08c-369ca313ce8e · inbound

BenchGuard: Who Guards the Benchmarks? Automated Auditing of LLM Agent Benchmarks cites this paper.

BenchGuard: Who Guards the Benchmarks? Automated Auditing of LLM Agent Benchmarks LLM Critics Help Catch LLM Bugs

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T00:24:28.369167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T03:30:55.419225Z digest=sha256:2d6337aada2df3dc3646218b7bd4859c371e942f772375933e2f7e82b5c53270

Observation 0435cd55-1b72-478f-b536-62a92a510796 · inbound

AI Alignment via Incentives and Correction cites this paper.

AI Alignment via Incentives and Correction LLM Critics Help Catch LLM Bugs

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:01:10.325818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-09T14:07:44.717260Z digest=sha256:86ed326aa70d787fa50c899ce65e344d7894038e14504639f98709f906cf100b

Observation 647e6ab5-a8ed-487c-80aa-51a9564a395b · inbound

AI Alignment via Incentives and Correction cites this paper.

AI Alignment via Incentives and Correction LLM Critics Help Catch LLM Bugs

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:51:26.693683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-12T04:51:31.357544Z digest=sha256:d037f34aa8e35dbe2f1ec93b10745a97b3e438a71fb30f46738d11117f76ae55

Observation f47c57aa-73df-4737-95fe-b9752c8cb113 · inbound

LLM Wardens: Mitigating Adversarial Persuasion with Third-Party Conversational Oversight cites this paper.

LLM Wardens: Mitigating Adversarial Persuasion with Third-Party Conversational Oversight LLM Critics Help Catch LLM Bugs

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:41:24.528520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-12T00:51:17.434889Z digest=sha256:1ba8b4e01a58a0436fc1bd533744f6e3779a4091d5b756a4a4836cd97be852bb

Observation 21af213a-b98c-4ed4-820b-f4daf76d597f · inbound

Beyond Binary: Reframing GUI Critique as Continuous Semantic Alignment cites this paper.

Beyond Binary: Reframing GUI Critique as Continuous Semantic Alignment LLM Critics Help Catch LLM Bugs

Reference 99

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:58:28.867558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-15T01:58:19.295247Z digest=sha256:852f1269bcc069279ec139b19c01f773e6f19bf342c20c97138722d66a803583

Observation fd0fc45f-6ddd-4935-a697-a8f60dc628b8 · inbound

Beyond Binary: Reframing GUI Critique as Continuous Semantic Alignment cites this paper.

Beyond Binary: Reframing GUI Critique as Continuous Semantic Alignment LLM Critics Help Catch LLM Bugs

Reference 99

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:47:40.465434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-19T16:45:16.963802Z digest=sha256:ac81d6120d7b8db1deb39ee5ce838cc43d12226a17cbb18c1867480acb15a6c5

Observation 580e838f-3c39-49ec-ae90-046e36b15931 · inbound

Philosophical Dispositions as Behavioral Constraints for AI-Assisted Code Review: An Empirical Study cites this paper.

Philosophical Dispositions as Behavioral Constraints for AI-Assisted Code Review: An Empirical Study LLM Critics Help Catch LLM Bugs

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:10:22.403624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-25T05:07:12.832683Z digest=sha256:8aa7fe53839e0f9d96ad9c2687cc82ebfdd5bed0cb673a291dc549c94853ea32

Observation e8979a3c-728b-483d-bbb9-c65035c546bb · inbound

Weak Critics Make Strong Learners: On-Policy Critique Distillation for Scalable Oversight cites this paper.

Weak Critics Make Strong Learners: On-Policy Critique Distillation for Scalable Oversight LLM Critics Help Catch LLM Bugs

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-01T19:56:10.983153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-28T21:55:58.645062Z digest=sha256:1d9aa6c7207eec7483f2917b2f1c5023abbb1aedaab23b01a8738d08a86b503d

Observation 75ea0ca6-6ce5-4326-9f02-775fcb3bd834 · inbound

Proof-or-Stop: Don't Trust the Agent, Trust the Evidence -- Loop Engineering for Verifiable Evidence-Gated Lifecycle Control cites this paper.

Proof-or-Stop: Don't Trust the Agent, Trust the Evidence -- Loop Engineering for Verifiable Evidence-Gated Lifecycle Control LLM Critics Help Catch LLM Bugs

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T00:50:39.714101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:50:39.714101Z digest=sha256:cd3be4be96c6261ede6b12b4a0a176020fb3d702f64f41163814eec6f3dbe892

Observation ee3d1936-9e66-4f5c-933f-b42f661aa8ea · inbound

Fantastic Adaptive Taxonomies and How to Use Them cites this paper.

Fantastic Adaptive Taxonomies and How to Use Them LLM Critics Help Catch LLM Bugs

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T21:14:42.470394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:14:42.470394Z digest=sha256:79c2c3f64e8a79debd890f30c9e4f73d1ac192328f67fdea1916dcf96ead3fca

Observation 5a5900ed-38ad-4d3e-b70c-10be22d5ff0d · inbound

Code Monitor Red Teaming for Public-Test-Passing Code cites this paper.

Code Monitor Red Teaming for Public-Test-Passing Code LLM Critics Help Catch LLM Bugs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T09:12:25.169891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:12:25.169891Z digest=sha256:21b6ba314c12122637e67f3a19a2d0cd07d2ff781c21bdf9fc9a66dc4e8f59bd

Observation b5b8ce2a-122f-47d0-b665-1c83909c0789 · inbound

A dataset of rated conceptual arguments cites this paper.

A dataset of rated conceptual arguments LLM Critics Help Catch LLM Bugs

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-31T01:28:21.684871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T01:28:21.684871Z digest=sha256:33e91597b3df392dc46a7d10a3cfafc239959f7feeb78694e551a28325e32634

Observation c37075e3-5fbb-44c3-91c3-1f907f29e722 · inbound

Beyond a Single Judge: The Evidence-Grounded, Social-Weighted Persona Panel for Generative UI Evaluation cites this paper.

Beyond a Single Judge: The Evidence-Grounded, Social-Weighted Persona Panel for Generative UI Evaluation LLM Critics Help Catch LLM Bugs

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-31T07:20:10.899680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T07:20:10.899680Z digest=sha256:c02fda585496e5e106fc67b538a34b7f8625257f7fab678d9ca91b81f2fc512d

Observation 088a8d31-2fec-43ca-9875-79c8f24ad036 · inbound

Beyond a Single Judge: The Evidence-Grounded, Social-Weighted Persona Panel for Generative UI Evaluation cites this paper.

Beyond a Single Judge: The Evidence-Grounded, Social-Weighted Persona Panel for Generative UI Evaluation LLM Critics Help Catch LLM Bugs

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T01:25:30.883047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T01:25:30.883047Z digest=sha256:6248fa2519df3d40a9b469f91e5adff3c9e59ab6962f4e86da49c0a1975a408c

Observation 12947eb9-15c3-41cd-ac6d-73160fddd018 · inbound

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets cites this paper.

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets LLM Critics Help Catch LLM Bugs

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T00:43:00.039128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:43:00.039128Z digest=sha256:2675e6b08da71293030e296e64cc87258cf6510315969335a64df5fe633c5928

Observation 1b97cf54-a4ec-4b8a-a023-7d41cfc8b182 · inbound

Quo Vadis, World Modeling? cites this paper.

Quo Vadis, World Modeling? LLM Critics Help Catch LLM Bugs

Reference 112

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:07.111956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:07.111956Z digest=sha256:bfe80c63f5ddf1f5ce16935ce4d36d0253742cdf1272765f4ef824a86118661b

Observation eae49f4d-c649-49a3-ad1a-6ff3ca62d586 · inbound

Apodex Discovery: Reality Benchmarks and Environments for Evaluating and Building Discoverative Artificial Intelligence cites this paper.

Apodex Discovery: Reality Benchmarks and Environments for Evaluating and Building Discoverative Artificial Intelligence LLM Critics Help Catch LLM Bugs

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-15T14:17:47.051909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:17:47.051909Z digest=sha256:666b135c3b7a2f4c68f5432963a60a17b43fd51962736167aca496e20046b9de

Observation 04cf6c93-592c-4d2a-afaa-6fbdbd114091 · inbound

Benchmarking LLM Judges for Mobile Agent Evaluation cites this paper.

Benchmarking LLM Judges for Mobile Agent Evaluation LLM Critics Help Catch LLM Bugs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:31.951140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:31.951140Z digest=sha256:2da951d6030df53cb0a1fa8653a702c59ee65ae998e3b6144a67277ed1bd921e