Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T15:31:50.457893Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 0 inbound Pith citation observations for arXiv:2607.24268.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T15:31:50.457893Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
29 of 29 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation d75de6bf-d70a-4e23-a4d4-afe513f9e0ae · outbound
Accuracy Hides How Language Models Fail: Measuring Failure States Under Matched Output Budgets G-eval: NLG evaluation using GPT-4 with better human alignment,
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5c66f487-d9b1-41b9-84e0-777ea94cef8f · outbound
Accuracy Hides How Language Models Fail: Measuring Failure States Under Matched Output Budgets Judging LLM-as-a-judge with MT-Bench and chatbot arena,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation da6bdfa2-b864-45a8-9519-79986e161a0a · outbound
Accuracy Hides How Language Models Fail: Measuring Failure States Under Matched Output Budgets Large language models are not fair evaluators,
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation fabbaf1c-095b-43ac-ae11-712faaf88e04 · outbound
Accuracy Hides How Language Models Fail: Measuring Failure States Under Matched Output Budgets Length- controlled AlpacaEval: A simple way to debias automatic evaluators,
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8a2af24b-4d52-49a7-ad76-ac48d2d73ca4 · outbound
Accuracy Hides How Language Models Fail: Measuring Failure States Under Matched Output Budgets Finding blind spots in evaluator LLMs with interpretable checklists,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 40ec0a0a-e7b1-4460-838a-cfe15bb4894a · outbound
Accuracy Hides How Language Models Fail: Measuring Failure States Under Matched Output Budgets A Survey on LLM-as-a-Judge
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 921af750-2776-453b-92cc-0d22ee0372b3 · outbound
Accuracy Hides How Language Models Fail: Measuring Failure States Under Matched Output Budgets When the Judge Changes, So Does the Measurement: Auditing LLM-as-Judge Reliability
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation dd879589-7aaf-4a69-bf39-ee792da652bb · outbound
Accuracy Hides How Language Models Fail: Measuring Failure States Under Matched Output Budgets Self- refine: Iterative refinement with self-feedback,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b28cc601-9f98-4768-9380-afaa294fa5f1 · outbound
Accuracy Hides How Language Models Fail: Measuring Failure States Under Matched Output Budgets Reflexion: Language agents with verbal reinforcement learning,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ca42c752-0fd2-4167-9895-e64840db8dd4 · outbound
Accuracy Hides How Language Models Fail: Measuring Failure States Under Matched Output Budgets Large language models cannot self- correct reasoning yet,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 07c5eeca-5556-4117-8535-0fc07ae76568 · outbound
Accuracy Hides How Language Models Fail: Measuring Failure States Under Matched Output Budgets Large language models can self-correct with key condition verification,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 3642869e-3a41-4ea7-9962-526b811e5a2f · outbound
Accuracy Hides How Language Models Fail: Measuring Failure States Under Matched Output Budgets When can LLMs actually correct their own mistakes? a critical survey of self-correction of LLMs,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 3080dec1-efe7-451b-90ad-42d5e4b8a65c · outbound
Accuracy Hides How Language Models Fail: Measuring Failure States Under Matched Output Budgets Automatically correcting large language models: Surveying the landscape of diverse automated correction strategies,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 25d25750-2f72-4bb9-b2c9-04581b52e961 · outbound
Accuracy Hides How Language Models Fail: Measuring Failure States Under Matched Output Budgets Chain-of-thought prompting elicits reasoning in large language models,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ade4312d-30e1-414d-b895-5a1a88b5dd4a · outbound
Accuracy Hides How Language Models Fail: Measuring Failure States Under Matched Output Budgets Self-consistency improves chain of thought reasoning in language models,
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50114e22-3636-4e88-885f-dd3bd2c83f94 · outbound
Accuracy Hides How Language Models Fail: Measuring Failure States Under Matched Output Budgets Scaling LLM test-time compute optimally can be more effective than scaling parameters for reasoning,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 7314de4d-bd1f-4887-a915-23c7ca31f8c5 · outbound
Accuracy Hides How Language Models Fail: Measuring Failure States Under Matched Output Budgets Adaptive Inference-Time Compute: LLMs Can Predict if They Can Do Better, Even Mid-Generation
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ecb34dfa-cb8b-4ab3-b587-21017041c133 · outbound
Accuracy Hides How Language Models Fail: Measuring Failure States Under Matched Output Budgets s1: Simple test-time scaling
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ca74e2b-4fb4-46fa-a8d7-80406f20dc76 · outbound
Accuracy Hides How Language Models Fail: Measuring Failure States Under Matched Output Budgets Beyond accuracy: Behavioral testing of NLP models with CheckList,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation bcc1da1e-bebe-43fd-a055-9858bb3c3b5d · outbound
Accuracy Hides How Language Models Fail: Measuring Failure States Under Matched Output Budgets Evaluating evaluation metrics: A framework for analyzing NLG evaluation metrics using measurement theory,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 0a34877f-8e1b-4069-bf9f-148d9abeaedb · outbound
Accuracy Hides How Language Models Fail: Measuring Failure States Under Matched Output Budgets Capturing failures of large language models via human cognitive biases,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 83e14cfb-c70a-4b60-a2a3-aca1262bb7c7 · outbound
Accuracy Hides How Language Models Fail: Measuring Failure States Under Matched Output Budgets What will it take to fix benchmarking in natural language understanding?
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 28797bff-1bd2-442a-b0e6-6e089837288d · outbound
Accuracy Hides How Language Models Fail: Measuring Failure States Under Matched Output Budgets A systematic classification of knowledge, reasoning, and context within the ARC dataset,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 48cfcb7d-00fa-4ef9-9057-4cc377d49652 · outbound
Accuracy Hides How Language Models Fail: Measuring Failure States Under Matched Output Budgets Towards a consensus taxonomy for annotating errors in automatically generated text,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation fd2e61cb-a871-42f3-8500-ee89c3652bc7 · outbound
Accuracy Hides How Language Models Fail: Measuring Failure States Under Matched Output Budgets Measuring mathematical problem solving with the MATH dataset,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ff5c7b8b-98bb-484a-98fd-010581f035ab · outbound
Accuracy Hides How Language Models Fail: Measuring Failure States Under Matched Output Budgets Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41c5571a-5ef5-4610-a38d-46573890dddd · outbound
Accuracy Hides How Language Models Fail: Measuring Failure States Under Matched Output Budgets Qwen3 Technical Report
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation efda8ccd-5b59-43e6-937c-45ba78e7daeb · outbound
Accuracy Hides How Language Models Fail: Measuring Failure States Under Matched Output Budgets Available: https://aclanthology.org/2024.tacl-1.27/
Reference 2024
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 0f2b6586-f2b0-486d-84b9-4091edc43772 · outbound
Accuracy Hides How Language Models Fail: Measuring Failure States Under Matched Output Budgets Available: https://aclanthology.org/2024.acl-long.511/
Reference 9450
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
No inbound Pith citation observations are available.