Pith. sign in

Paper Citation Record · LEDGER

Accuracy Hides How Language Models Fail: Measuring Failure States Under Matched Output Budgets

As of 18 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 0 inbound Pith citation observations for arXiv:2607.24268.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.24268 v2

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:31:50.457893Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

29 of 29 outbound references displayed

  • verified exact1
  • verified fuzzy22
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d75de6bf-d70a-4e23-a4d4-afe513f9e0ae · outbound

This paper cites G-eval: NLG evaluation using GPT-4 with better human alignment,.

Accuracy Hides How Language Models Fail: Measuring Failure States Under Matched Output Budgets G-eval: NLG evaluation using GPT-4 with better human alignment,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:31:50.836135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T15:31:50.353991Z digest=sha256:0bf147d58c4530382bebb9c476d37a127f8fd4b0ff12f1dc4fc8c1aad28ded0e

Observation 5c66f487-d9b1-41b9-84e0-777ea94cef8f · outbound

This paper cites Judging LLM-as-a-judge with MT-Bench and chatbot arena,.

Accuracy Hides How Language Models Fail: Measuring Failure States Under Matched Output Budgets Judging LLM-as-a-judge with MT-Bench and chatbot arena,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:31:50.825044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T15:31:50.357946Z digest=sha256:e2baeda13274ed623894a6af7620e782f201865569c39a460ebf76ba36f61d08

Observation da6bdfa2-b864-45a8-9519-79986e161a0a · outbound

This paper cites Large language models are not fair evaluators,.

Accuracy Hides How Language Models Fail: Measuring Failure States Under Matched Output Budgets Large language models are not fair evaluators,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:31:50.813236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T15:31:50.361614Z digest=sha256:e8398422e3ee8cdd29fed124441a8ab68148ccf5a402193bb9bec3a852ea12f5

Observation fabbaf1c-095b-43ac-ae11-712faaf88e04 · outbound

This paper cites Length- controlled AlpacaEval: A simple way to debias automatic evaluators,.

Accuracy Hides How Language Models Fail: Measuring Failure States Under Matched Output Budgets Length- controlled AlpacaEval: A simple way to debias automatic evaluators,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:31:50.789095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T15:31:50.368490Z digest=sha256:92f9764cb252edcf0fd30bdf7bdd05addd129d07173f4621605ab493a7c89b65

Observation 8a2af24b-4d52-49a7-ad76-ac48d2d73ca4 · outbound

This paper cites Finding blind spots in evaluator LLMs with interpretable checklists,.

Accuracy Hides How Language Models Fail: Measuring Failure States Under Matched Output Budgets Finding blind spots in evaluator LLMs with interpretable checklists,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:31:50.777598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T15:31:50.372053Z digest=sha256:d42de172d20ce877530949ce6bb7ec4be213813fec93e857428c432abc5f14ac

Observation 40ec0a0a-e7b1-4460-838a-cfe15bb4894a · outbound

This paper cites A Survey on LLM-as-a-Judge.

Accuracy Hides How Language Models Fail: Measuring Failure States Under Matched Output Budgets A Survey on LLM-as-a-Judge

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T15:31:50.375706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:31:50.375706Z digest=sha256:4559d1b7aeabb3ba153f91002113b8a8f652f6c7c78eaa85a0b505f26bf1bdb0

Observation 921af750-2776-453b-92cc-0d22ee0372b3 · outbound

This paper cites When the Judge Changes, So Does the Measurement: Auditing LLM-as-Judge Reliability.

Accuracy Hides How Language Models Fail: Measuring Failure States Under Matched Output Budgets When the Judge Changes, So Does the Measurement: Auditing LLM-as-Judge Reliability

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-15T15:31:50.549766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T15:31:50.379275Z digest=sha256:aa6d132b73f3add39ee551f40bcb917f54039704f3122919244765c9fcc86fff

Observation dd879589-7aaf-4a69-bf39-ee792da652bb · outbound

This paper cites Self- refine: Iterative refinement with self-feedback,.

Accuracy Hides How Language Models Fail: Measuring Failure States Under Matched Output Budgets Self- refine: Iterative refinement with self-feedback,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:31:50.766115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T15:31:50.383087Z digest=sha256:b52ae6d99a7d7bc52130dcf95648429f6f2cd77898cbc360f9f4ff4574c7ec0d

Observation b28cc601-9f98-4768-9380-afaa294fa5f1 · outbound

This paper cites Reflexion: Language agents with verbal reinforcement learning,.

Accuracy Hides How Language Models Fail: Measuring Failure States Under Matched Output Budgets Reflexion: Language agents with verbal reinforcement learning,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:31:50.754828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T15:31:50.387492Z digest=sha256:a3ae81b5e8cc8931c9fc3bc0512e1d9b344c14a08fe42043f85cd018b1ff2a02

Observation ca42c752-0fd2-4167-9895-e64840db8dd4 · outbound

This paper cites Large language models cannot self- correct reasoning yet,.

Accuracy Hides How Language Models Fail: Measuring Failure States Under Matched Output Budgets Large language models cannot self- correct reasoning yet,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:31:50.743071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T15:31:50.391256Z digest=sha256:045d8c951655268b8a32460bf55bd8c7c999e409563cf21892559cb8272fb3a9

Observation 07c5eeca-5556-4117-8535-0fc07ae76568 · outbound

This paper cites Large language models can self-correct with key condition verification,.

Accuracy Hides How Language Models Fail: Measuring Failure States Under Matched Output Budgets Large language models can self-correct with key condition verification,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:31:50.730242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T15:31:50.394873Z digest=sha256:68ce3a70b3a73ad4ce5b8a7841f4c1c2f442ad1d7e7064392132bbdeaa4be6eb

Observation 3642869e-3a41-4ea7-9962-526b811e5a2f · outbound

This paper cites When can LLMs actually correct their own mistakes? a critical survey of self-correction of LLMs,.

Accuracy Hides How Language Models Fail: Measuring Failure States Under Matched Output Budgets When can LLMs actually correct their own mistakes? a critical survey of self-correction of LLMs,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:31:50.718264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T15:31:50.398445Z digest=sha256:4215e2afee5123d3a30d1772626fa40cd461e19eea591d2b529176101ad40ff2

Observation 3080dec1-efe7-451b-90ad-42d5e4b8a65c · outbound

This paper cites Automatically correcting large language models: Surveying the landscape of diverse automated correction strategies,.

Accuracy Hides How Language Models Fail: Measuring Failure States Under Matched Output Budgets Automatically correcting large language models: Surveying the landscape of diverse automated correction strategies,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:31:50.704086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T15:31:50.402226Z digest=sha256:6d52d8442cb661ec64a5a43f060dfb2eb5e9d63911fced191e21e65aaf07270e

Observation 25d25750-2f72-4bb9-b2c9-04581b52e961 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models,.

Accuracy Hides How Language Models Fail: Measuring Failure States Under Matched Output Budgets Chain-of-thought prompting elicits reasoning in large language models,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:31:50.680104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T15:31:50.409957Z digest=sha256:0b9afdea8467f8ebe3eff9c82ea2d688c535743bd62828cd8c18761f398123bf

Observation ade4312d-30e1-414d-b895-5a1a88b5dd4a · outbound

This paper cites Self-consistency improves chain of thought reasoning in language models,.

Accuracy Hides How Language Models Fail: Measuring Failure States Under Matched Output Budgets Self-consistency improves chain of thought reasoning in language models,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T15:31:50.413411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:31:50.413411Z digest=sha256:ac1f4765a5c059a2a4b6339a385b8a11bc0baa9e1e219afa79915bd38fc79a8b

Observation 50114e22-3636-4e88-885f-dd3bd2c83f94 · outbound

This paper cites Scaling LLM test-time compute optimally can be more effective than scaling parameters for reasoning,.

Accuracy Hides How Language Models Fail: Measuring Failure States Under Matched Output Budgets Scaling LLM test-time compute optimally can be more effective than scaling parameters for reasoning,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:31:50.661638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T15:31:50.416852Z digest=sha256:1921c1ca9814a09f0f90be1edfe3de830a1f11af6d7cceaa0fca62ca10921bb0

Observation 7314de4d-bd1f-4887-a915-23c7ca31f8c5 · outbound

This paper cites Adaptive Inference-Time Compute: LLMs Can Predict if They Can Do Better, Even Mid-Generation.

Accuracy Hides How Language Models Fail: Measuring Failure States Under Matched Output Budgets Adaptive Inference-Time Compute: LLMs Can Predict if They Can Do Better, Even Mid-Generation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T15:31:50.420577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:31:50.420577Z digest=sha256:0dc7f141b73cd4a99095377bdaf28423ec57c28aaee6077c4b0117bda167439d

Observation ecb34dfa-cb8b-4ab3-b587-21017041c133 · outbound

This paper cites s1: Simple test-time scaling.

Accuracy Hides How Language Models Fail: Measuring Failure States Under Matched Output Budgets s1: Simple test-time scaling

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T15:31:50.424448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:31:50.424448Z digest=sha256:f53969db42ddde80c1cbdb61d90fb9dc5c1f4b905a5cc833dc0adbea9b634c35

Observation 6ca74e2b-4fb4-46fa-a8d7-80406f20dc76 · outbound

This paper cites Beyond accuracy: Behavioral testing of NLP models with CheckList,.

Accuracy Hides How Language Models Fail: Measuring Failure States Under Matched Output Budgets Beyond accuracy: Behavioral testing of NLP models with CheckList,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:31:50.650193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T15:31:50.428160Z digest=sha256:e509fd97d573d261c52516a307ca1c96e336c705dc0bf191e824211284dc45ac

Observation bcc1da1e-bebe-43fd-a055-9858bb3c3b5d · outbound

This paper cites Evaluating evaluation metrics: A framework for analyzing NLG evaluation metrics using measurement theory,.

Accuracy Hides How Language Models Fail: Measuring Failure States Under Matched Output Budgets Evaluating evaluation metrics: A framework for analyzing NLG evaluation metrics using measurement theory,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:31:50.635562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T15:31:50.432593Z digest=sha256:fc72e5fec88df36318ba5724ccd6cdeebb7d3b6e8a0a49f9ab8c9727e079c44e

Observation 0a34877f-8e1b-4069-bf9f-148d9abeaedb · outbound

This paper cites Capturing failures of large language models via human cognitive biases,.

Accuracy Hides How Language Models Fail: Measuring Failure States Under Matched Output Budgets Capturing failures of large language models via human cognitive biases,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:31:50.621651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T15:31:50.436117Z digest=sha256:c04f8416c3e51bd003fcce433b3b839a01ed2a6776d12701d9d2d3624470972a

Observation 83e14cfb-c70a-4b60-a2a3-aca1262bb7c7 · outbound

This paper cites What will it take to fix benchmarking in natural language understanding?.

Accuracy Hides How Language Models Fail: Measuring Failure States Under Matched Output Budgets What will it take to fix benchmarking in natural language understanding?

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:31:50.607959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T15:31:50.439905Z digest=sha256:f4654d5541108c61133d9b10517afce537ee4954608c24c605acf5d80543f264

Observation 28797bff-1bd2-442a-b0e6-6e089837288d · outbound

This paper cites A systematic classification of knowledge, reasoning, and context within the ARC dataset,.

Accuracy Hides How Language Models Fail: Measuring Failure States Under Matched Output Budgets A systematic classification of knowledge, reasoning, and context within the ARC dataset,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:31:50.595915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T15:31:50.443972Z digest=sha256:00d17adad1c99ee2977f93eed7ac95fcfa0bdee579e41acc59f3b8625dd5ebdb

Observation 48cfcb7d-00fa-4ef9-9057-4cc377d49652 · outbound

This paper cites Towards a consensus taxonomy for annotating errors in automatically generated text,.

Accuracy Hides How Language Models Fail: Measuring Failure States Under Matched Output Budgets Towards a consensus taxonomy for annotating errors in automatically generated text,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:31:50.584621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T15:31:50.447539Z digest=sha256:f155b9f5103d446aa53a2d20ed89c5fdeb6c0d9e04e962a283a59a29dbeb986c

Observation fd2e61cb-a871-42f3-8500-ee89c3652bc7 · outbound

This paper cites Measuring mathematical problem solving with the MATH dataset,.

Accuracy Hides How Language Models Fail: Measuring Failure States Under Matched Output Budgets Measuring mathematical problem solving with the MATH dataset,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:31:50.573378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T15:31:50.450911Z digest=sha256:55bf171a0497024189ea818f0a1c771505e2bf976eb580ecfe4ab0f8622b26c6

Observation ff5c7b8b-98bb-484a-98fd-010581f035ab · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Accuracy Hides How Language Models Fail: Measuring Failure States Under Matched Output Budgets Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T15:31:50.454276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:31:50.454276Z digest=sha256:182e4ba29627c38691a1c7ab90fd44082a06e7fc04658dde9a84abbea01a51a9

Observation 41c5571a-5ef5-4610-a38d-46573890dddd · outbound

This paper cites Qwen3 Technical Report.

Accuracy Hides How Language Models Fail: Measuring Failure States Under Matched Output Budgets Qwen3 Technical Report

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T15:31:50.457893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:31:50.457893Z digest=sha256:6bc65714e13cdbf8e4a8e33d7145b57cbc815343e6bed612c45ccb896fe1eded

Observation efda8ccd-5b59-43e6-937c-45ba78e7daeb · outbound

This paper cites Available: https://aclanthology.org/2024.tacl-1.27/.

Accuracy Hides How Language Models Fail: Measuring Failure States Under Matched Output Budgets Available: https://aclanthology.org/2024.tacl-1.27/

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:31:50.692569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T15:31:50.406226Z digest=sha256:e25525d92c22994958f4e40488b73f59f2f823df1b3da0f6f1fadd9ed9ca2075

Observation 0f2b6586-f2b0-486d-84b9-4091edc43772 · outbound

This paper cites Available: https://aclanthology.org/2024.acl-long.511/.

Accuracy Hides How Language Models Fail: Measuring Failure States Under Matched Output Budgets Available: https://aclanthology.org/2024.acl-long.511/

Reference 9450

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:31:50.800533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T15:31:50.365254Z digest=sha256:7afb783f1ae5df459b7430474d731a034338aa3032caf7b53516f259e4bb2b5b

Pith citing papers

No inbound Pith citation observations are available.