Pith. sign in

Paper Citation Record · LEDGER

ProbeLLM: Automating Principled Diagnosis of LLM Failures

As of 6 August 2026, this Paper Citation Record lists 12 of 12 outbound references and 6 inbound Pith citation observations for arXiv:2602.12966.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2602.12966 v2

Coverage vector

measured 12 of 12 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T23:43:07.803719Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-14T06:09:46.933171Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-03T14:28:31.386601Z

Reference resolution

12 of 12 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved11
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d7db0cbe-6faa-4159-83b1-d5822ad74e8d · outbound

This paper cites URL https: //aclanthology.org/2025.acl-long.17/.

ProbeLLM: Automating Principled Diagnosis of LLM Failures URL https: //aclanthology.org/2025.acl-long.17/

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T23:43:06.888055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:43:06.888055Z digest=sha256:a1c3a18dfac49f9fa08365a64640d2815141105e9caf8065f629fc4f5884c19a

Observation 5cf49950-f877-4f14-98ac-021460384fb4 · outbound

This paper cites Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models.

ProbeLLM: Automating Principled Diagnosis of LLM Failures Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T23:43:07.037472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:43:07.037472Z digest=sha256:ebe012da5af9fa0c05a73a7e6a84dc5da2c88856d0709a60eb7e4f2709b62108

Observation 02beb234-b5d3-47f7-bdcd-86b5ceee21f2 · outbound

This paper cites gpt-oss-120b & gpt-oss-20b Model Card.

ProbeLLM: Automating Principled Diagnosis of LLM Failures gpt-oss-120b & gpt-oss-20b Model Card

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T23:43:07.177211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:43:07.177211Z digest=sha256:590df5ff82032145ab535c9840cf63651ef5cceaa6499d2ded1533ec69d35bc0

Observation 62c7fc2c-a713-4d3c-9f1f-d8ae70d65ebf · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

ProbeLLM: Automating Principled Diagnosis of LLM Failures DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T23:43:07.294716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:43:07.294716Z digest=sha256:09f43aa042029c98f6a09886ebe512ce54d279e4816e8c8e816322d2bf1967a3

Observation 9b38479a-5fe6-48ff-b489-5aa0bc94c903 · outbound

This paper cites Discovering Knowledge Deficiencies of Language Models on Massive Knowledge Base.

ProbeLLM: Automating Principled Diagnosis of LLM Failures Discovering Knowledge Deficiencies of Language Models on Massive Knowledge Base

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T23:43:07.464377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:43:07.464377Z digest=sha256:4b5fdbf9f4946dfec44339ad57fd303f5fb3b0871c0bcc5e5938424249e5cc05

Observation 7e203e4a-8339-481e-8f50-b0f2ab11aa5e · outbound

This paper cites Accessing GPT-4 level Mathematical Olympiad Solutions via Monte Carlo Tree Self-refine with LLaMa-3 8B.

ProbeLLM: Automating Principled Diagnosis of LLM Failures Accessing GPT-4 level Mathematical Olympiad Solutions via Monte Carlo Tree Self-refine with LLaMa-3 8B

Reference 12

Resolution
malformed identifier
no resolver link, observed 2026-08-02T23:43:07.803719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:43:07.803719Z digest=sha256:f4e7205cb172c0ab4bd46b3a525d0b9c558fa366479f1c2dbdceb62395a271dc

Observation a2578719-f5dc-470c-9026-44b8627cdff5 · outbound

This paper cites org/CorpusID:15184765.

ProbeLLM: Automating Principled Diagnosis of LLM Failures org/CorpusID:15184765

Reference 2006

Resolution
unresolved
no resolver link, observed 2026-08-02T23:43:06.773853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:43:06.773853Z digest=sha256:bfb5e34e25b67d1cb9d39b24cbb92f2ff44a3a1454cef62493ea32a964263ca6

Observation 746a367d-6d75-4b97-850f-e419c3b57eb1 · outbound

This paper cites A Comprehensive Survey in LLM(-Agent) Full Stack Safety: Data, Training and Deployment.

ProbeLLM: Automating Principled Diagnosis of LLM Failures A Comprehensive Survey in LLM(-Agent) Full Stack Safety: Data, Training and Deployment

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-02T23:43:07.687190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:43:07.687190Z digest=sha256:e2fdb92ae4c5e29ccf8623b7931f738d9c644472a87be4b03504f7a1036bf127

Observation 35f6cc96-e0d7-4b2e-b2c7-488aac151725 · outbound

This paper cites Toward Automated Robustness Evaluation of Mathematical Reasoning.

ProbeLLM: Automating Principled Diagnosis of LLM Failures Toward Automated Robustness Evaluation of Mathematical Reasoning

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-02T23:43:06.664084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:43:06.664084Z digest=sha256:d9b4956e0b4d8eabf0300c92518c5612e8fa9a7d30d182a012af21505ca46b60

Observation d7328d18-8700-4e10-9f7c-525745f77d99 · outbound

This paper cites Efficient Process Reward Model Training via Active Learning.

ProbeLLM: Automating Principled Diagnosis of LLM Failures Efficient Process Reward Model Training via Active Learning

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-02T23:43:06.529225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:43:06.529225Z digest=sha256:0a8e45f83a9b31f10293ff5b17cf4c80dbcf88c74e93303d962f4bffc4f55127

Observation af7a3fd0-9c46-43f7-bb56-7917466a3894 · outbound

This paper cites Program Synthesis with Large Language Models.

ProbeLLM: Automating Principled Diagnosis of LLM Failures Program Synthesis with Large Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-02T23:43:06.429401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:43:06.429401Z digest=sha256:9d87baaf4ab1eec70a6d6c880d154b1c38354881bfebc5bc3b60d621629637d4

Observation cfca5e10-3479-4c8a-821a-b05cb32972c2 · outbound

This paper cites org/CorpusID:277452302.

ProbeLLM: Automating Principled Diagnosis of LLM Failures org/CorpusID:277452302

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-02T23:43:07.587407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:43:07.587407Z digest=sha256:7f339ce3c9fccf1433e7beeeb4240ae8834b2efc3cc1c5afd877d64ad4b0490c

Pith citing papers

Observation 72ccd56b-2c5d-483f-84e1-aabe6c4cb8b2 · inbound

Guardian-as-an-Advisor: Advancing Next-Generation Guardian Models for Trustworthy LLMs cites this paper.

Guardian-as-an-Advisor: Advancing Next-Generation Guardian Models for Trustworthy LLMs ProbeLLM: Automating Principled Diagnosis of LLM Failures

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-06-10T02:11:49.648006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T17:27:13.339411Z digest=sha256:e35120c135fb0c5a2fbb84275d93c0500c8354dffd1a2f66afa307f60f3e5c4c

Observation 1e1b4f2a-747c-4ddf-9282-4b6491c86229 · inbound

Chain of Risk: Safety Failures in Large Reasoning Models and Mitigation via Adaptive Multi-Principle Steering cites this paper.

Chain of Risk: Safety Failures in Large Reasoning Models and Mitigation via Adaptive Multi-Principle Steering ProbeLLM: Automating Principled Diagnosis of LLM Failures

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-06-10T02:11:49.648006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T11:49:47.994456Z digest=sha256:da96a6ad6790e3864bf4cba355a3fed1edfbac8b6805a1254f795caafb2d2df7

Observation 835d23ac-818a-45cb-b47f-0dbb2c9695c1 · inbound

FailureScope: Cross-Regime Behavioral Diagnosis of Language Model Weaknesses cites this paper.

FailureScope: Cross-Regime Behavioral Diagnosis of Language Model Weaknesses ProbeLLM: Automating Principled Diagnosis of LLM Failures

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-07-02T06:56:44.731661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T07:16:30.725358Z digest=sha256:5cdfab44ba290a234065cc5d8bf0aab836fa1af62579e175883ec682f9d073b0

Observation 1b007848-5601-4196-97d2-e454efe05a34 · inbound

Getting Better at Working With You: Compiling User Corrections into Runtime Enforcement for Coding Agents cites this paper.

Getting Better at Working With You: Compiling User Corrections into Runtime Enforcement for Coding Agents ProbeLLM: Automating Principled Diagnosis of LLM Failures

Reference 30

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T14:28:31.387876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-27T07:04:51.970049Z digest=sha256:c5ed2ba3135c004817005dbd7fd6762f399389871375f7a476506182dfd47b03

Observation d69426c1-14cb-4eb9-bfc8-27e268cffb13 · inbound

DeepBias: Adaptive In-depth Probing of Social Biases in LVLMs cites this paper.

DeepBias: Adaptive In-depth Probing of Social Biases in LVLMs ProbeLLM: Automating Principled Diagnosis of LLM Failures

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-14T06:09:46.933171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T06:09:46.933171Z digest=sha256:490f4fd0118e21c8f9bcb55ab4374334197b0261792e63d1187543c43f4b1731

Observation 5d4d1f21-7196-433c-a1f7-36cb705dc80b · inbound

Inside the Unfair Judge: A Mechanistic Interpretability Account of LLM-as-Judge Bias cites this paper.

Inside the Unfair Judge: A Mechanistic Interpretability Account of LLM-as-Judge Bias ProbeLLM: Automating Principled Diagnosis of LLM Failures

Reference 141

Resolution
unresolved
no resolver link, observed 2026-07-14T02:33:34.084111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T02:33:34.084111Z digest=sha256:f14e6a09093bbf25c1780fe9595effbd5e686eb0350e982e7111f63d63e34bdb