Pith. sign in

Paper Citation Record · LEDGER

AI Evaluation Should Measure Verification Cost, Not Correctness Alone

As of 21 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 0 inbound Pith citation observations for arXiv:2608.08709.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.08709 v1

Coverage vector

measured 14 of 14 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:33:34.992330Z

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

14 of 14 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b1605f40-1a17-4c9a-8c0f-0af89339ed7b · outbound

This paper cites Concrete Problems in AI Safety.

AI Evaluation Should Measure Verification Cost, Not Correctness Alone Concrete Problems in AI Safety

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-14T04:33:34.605882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:33:34.605882Z digest=sha256:ad610d9a1fd924863e81f2310728461f0148b5cc00d9cc6111774cc08f4255df

Observation 75964956-e509-47b8-a7ea-4f16a366c623 · outbound

This paper cites Sparks of Artificial General Intelligence: Early experiments with GPT-4.

AI Evaluation Should Measure Verification Cost, Not Correctness Alone Sparks of Artificial General Intelligence: Early experiments with GPT-4

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-14T04:33:34.689041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:33:34.689041Z digest=sha256:b68005ae5da0b0bde136e600c18d1bd55bd9db23b8c745bdbff77fc21fe631b2

Observation 052835b1-af32-4850-8f51-64849255884e · outbound

This paper cites Susanne Förster and Yarden Skop.

AI Evaluation Should Measure Verification Cost, Not Correctness Alone Susanne Förster and Yarden Skop

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-14T04:33:34.714599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:33:34.714599Z digest=sha256:944c90b6594d0433a163d3f48301453b7a9fe907926f05880067176fbe7e6493

Observation 81834d87-def8-4435-b782-347c01781d89 · outbound

This paper cites Luyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon, Pengfei Liu, Yiming Yang, Jamie Callan, and Graham Neubig.

AI Evaluation Should Measure Verification Cost, Not Correctness Alone Luyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon, Pengfei Liu, Yiming Yang, Jamie Callan, and Graham Neubig

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-14T04:33:34.728945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:33:34.728945Z digest=sha256:07ddc3aa3de0d81a8727c9cd48d33d0eace6fd74a6325dd53ab9b9283645e9f1

Observation 60f7ec44-764b-4fca-9202-b842061708ed · outbound

This paper cites Daniel Kahneman.Thinking, Fast and Slow.

AI Evaluation Should Measure Verification Cost, Not Correctness Alone Daniel Kahneman.Thinking, Fast and Slow

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-14T04:33:34.744259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:33:34.744259Z digest=sha256:98f05a46aba67bf14a4575c639015b35ad55c2c305897b8b88d4436feb92a08c

Observation 1b6494b6-994d-4945-b546-7ea07615f41d · outbound

This paper cites Potsawee Manakul, Adian Liusie, and Mark J.

AI Evaluation Should Measure Verification Cost, Not Correctness Alone Potsawee Manakul, Adian Liusie, and Mark J

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-14T04:33:34.899748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:33:34.899748Z digest=sha256:51f778711ad76d14d118178b647568c656bdb5cab87ffc3a038e3a434a06cef8

Observation 4d888af8-4f10-483a-8fc1-045a40d3b549 · outbound

This paper cites Sewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis, Wen-tau Yih, Pang Wei Koh, Mohit Iyyer, Luke Zettlemoyer, and Hannaneh Hajishirzi.

AI Evaluation Should Measure Verification Cost, Not Correctness Alone Sewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis, Wen-tau Yih, Pang Wei Koh, Mohit Iyyer, Luke Zettlemoyer, and Hannaneh Hajishirzi

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-14T04:33:34.970806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:33:34.970806Z digest=sha256:b9d005708fa0f229d51fec5b15e8edd6e0f0892aaaeeedb45aa9bc466c872cff

Observation 667a3ac3-32d1-4f8d-b107-a82685e74365 · outbound

This paper cites Towards a Science of AI Agent Reliability.

AI Evaluation Should Measure Verification Cost, Not Correctness Alone Towards a Science of AI Agent Reliability

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-14T04:33:34.985430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:33:34.985430Z digest=sha256:8049eb31d7fe83bdcf2eef08f3e7507dbc4280550fd671593b050811c338924a

Observation 37589ff0-6b28-4539-b08e-461b640db4e9 · outbound

This paper cites Bernhard Schölkopf, Francesco Locatello, Stefan Bauer, Nan Rosemary Ke, Nal Kalchbrenner, Anirudh Goyal, and Yoshua Bengio.

AI Evaluation Should Measure Verification Cost, Not Correctness Alone Bernhard Schölkopf, Francesco Locatello, Stefan Bauer, Nan Rosemary Ke, Nal Kalchbrenner, Anirudh Goyal, and Yoshua Bengio

Reference 16

Resolution
metadata mismatch
raw_fallback, observed 2026-08-14T04:33:35.480668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T04:33:34.992330Z digest=sha256:110d9fa621b3f02571ca2579816d52678357f120aea016a2802104cbaf7d6dc0

Observation 747765ba-92eb-43ce-a1a4-a1b894d3504e · outbound

This paper cites Scaling Laws for Neural Language Models.

AI Evaluation Should Measure Verification Cost, Not Correctness Alone Scaling Laws for Neural Language Models

Reference 2011

Resolution
unresolved
no resolver link, observed 2026-08-14T04:33:34.799903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:33:34.799903Z digest=sha256:05d0bdae5a70f99b37e775089fa03b9315afe21bd9c08ca22416235c5194058a

Observation 108c6eec-cab0-4005-852c-fcca86b584d7 · outbound

This paper cites URL https://aclanthology.org/2023.ijcnlp-main.20/.

AI Evaluation Should Measure Verification Cost, Not Correctness Alone URL https://aclanthology.org/2023.ijcnlp-main.20/

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-14T04:33:34.859455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:33:34.859455Z digest=sha256:151766fc7c0b64f9da5f956ad745574bbe2b28c6d9077676490b0ca30c7e1b4e

Observation 0abbdd35-a239-4609-bfe2-bff815ceb173 · outbound

This paper cites Chunyuan Deng, Yilun Zhao, Xiangru Tang, Mark Gerstein, and Arman Cohan.

AI Evaluation Should Measure Verification Cost, Not Correctness Alone Chunyuan Deng, Yilun Zhao, Xiangru Tang, Mark Gerstein, and Arman Cohan

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-14T04:33:34.702807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:33:34.702807Z digest=sha256:51121bdf4fd8736393152f1b7a9b1fef2b080b2387cc6560d92a4ef0fd831ff8

Observation 30e3fe52-d0f5-4a0c-b6e2-d98b7ce69d92 · outbound

This paper cites Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity.

AI Evaluation Should Measure Verification Cost, Not Correctness Alone Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-14T04:33:34.673519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:33:34.673519Z digest=sha256:12ab6f52c43acd36a673ae7f283108615690a3ef86d69b4d671ff43893a9672f

Observation 1831bd58-92e3-4c0f-9c53-ac4e30fcc30c · outbound

This paper cites URLhttps://www.worldscientific.com/doi/10.1142/S1793351X26410011.

AI Evaluation Should Measure Verification Cost, Not Correctness Alone URLhttps://www.worldscientific.com/doi/10.1142/S1793351X26410011

Reference 2026

Resolution
verified exact
doi, observed 2026-08-14T04:33:35.198209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T04:33:34.738199Z digest=sha256:4c3329277550fcec150b21114d24d8f3bb452a428bb1227851373ebba71d94ee

Pith citing papers

No inbound Pith citation observations are available.