Pith. sign in

Paper Citation Record · LEDGER

Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2505.03814.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.03814 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:45:13.720220Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T20:41:10.421116Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 4c05919c-aa1e-4ef6-8b37-46ce0ca263c2 · inbound

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models cites this paper.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:13.720220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:13.720220Z digest=sha256:7dc6170e4e9a648ef6062d316eace75035203bc95842b347c10f32d2c01362e5

Observation 8a0a392f-3bd4-4fc9-ab66-29e8e58cca04 · inbound

ReasonBENCH: Benchmarking the (In)Stability of LLM Reasoning cites this paper.

ReasonBENCH: Benchmarking the (In)Stability of LLM Reasoning Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T17:53:56.098545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:53:56.098545Z digest=sha256:9e07c57b5aeba78aa6d1b554ff2e23414f38909d7d24aad7c89c76abb2888fef

Observation b3c40ee8-e93d-4b1a-a8fd-2c4f649f8ba9 · inbound

ProEval: Proactive Failure Discovery and Efficient Performance Estimation for Generative AI Evaluation cites this paper.

ProEval: Proactive Failure Discovery and Efficient Performance Estimation for Generative AI Evaluation Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:41:10.426607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-08T08:21:55.648930Z digest=sha256:2be0a266e45801286ebfcbea6c05ff3681a72013472879cdb728ee845f877501

Observation bb21be98-65cc-4441-808f-cc85b0c17896 · inbound

How Many Iterations to Jailbreak? Dynamic Budget Allocation for Multi-Turn LLM Evaluation cites this paper.

How Many Iterations to Jailbreak? Dynamic Budget Allocation for Multi-Turn LLM Evaluation Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:21:07.933340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T12:13:16.896486Z digest=sha256:878a8b1fa3dc6692bfc021cc0e4fd16022db1e73b62a22254f5d4b1ebe00b800

Observation c271690c-8241-43d3-93c2-3d91eaeedb8a · inbound

How Many Iterations to Jailbreak? Dynamic Budget Allocation for Multi-Turn LLM Evaluation cites this paper.

How Many Iterations to Jailbreak? Dynamic Budget Allocation for Multi-Turn LLM Evaluation Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-02T14:49:15.527713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:49:15.527713Z digest=sha256:07814503d8f4bdc3e86920d5885f4555c792912ecc57fe26fb3219b42875320b

Observation 05c98c1b-116c-44c4-800f-9e4655137da9 · inbound

BayesAME: Bayesian Active Model Evaluation cites this paper.

BayesAME: Bayesian Active Model Evaluation Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-07-30T13:33:13.002509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T13:33:13.002509Z digest=sha256:7481f062c0d6f8823e2d9cf7255358f7328b8002420a6b1baafaa62a6621f900

Observation 97598b0f-a068-43aa-84ad-40580d288ceb · inbound

ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision cites this paper.

ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T07:19:03.239193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:19:03.239193Z digest=sha256:cf8b75b7d25de53f7af87380ceac6f3e2b2811556e3286c42563dbce88481bfa