Pith. sign in

Paper Citation Record · LEDGER

Growing Pains: Extensible and Efficient LLM Benchmarking Via Fixed Parameter Calibration

As of 5 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 2 inbound Pith citation observations for arXiv:2604.12843.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.12843 v2

Coverage vector

measured 17 of 17 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T16:37:12.962958Z

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T04:45:20.445703Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-03T16:58:43.279807Z

Reference resolution

17 of 17 outbound references displayed

  • verified exact4
  • verified fuzzy3
  • unresolved1
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch8

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0b598f78-1ffb-4e76-adb7-2b07ddcf6e06 · outbound

This paper cites Revealing the structure of language model capabilities.

Growing Pains: Extensible and Efficient LLM Benchmarking Via Fixed Parameter Calibration Revealing the structure of language model capabilities

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T08:30:56.713555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T16:37:12.962958Z digest=sha256:f22a4e01027ac7b1d800371364eb9e869ed9b109b9a6a080bbc182f41bec3b75

Observation e77699b9-3d1b-493a-b72a-2bf8eb368f5f · outbound

This paper cites Measuring massive multitask language understanding.

Growing Pains: Extensible and Efficient LLM Benchmarking Via Fixed Parameter Calibration Measuring massive multitask language understanding

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:29:46.093857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T16:37:12.962958Z digest=sha256:8d4f0c867ef8a16261fda405d051d89887978f10b8a725f321f5600a9cff7e87

Observation a47e77c2-6c2e-4fb1-9ce6-4bb9f5440cd7 · outbound

This paper cites A rosetta stone for ai benchmarks.

Growing Pains: Extensible and Efficient LLM Benchmarking Via Fixed Parameter Calibration A rosetta stone for ai benchmarks

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T08:30:56.706135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T16:37:12.962958Z digest=sha256:a990fe0c638f9361cbc3f5473f23dd99e0142d34c6645070cddf137a7d964ba5

Observation 5ea3ca8f-8041-41f7-bb62-e0f3cf9200c3 · outbound

This paper cites Fluid Language Model Benchmarking.

Growing Pains: Extensible and Efficient LLM Benchmarking Via Fixed Parameter Calibration Fluid Language Model Benchmarking

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T08:30:56.698614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T16:37:12.962958Z digest=sha256:a310bc562f7a9c48afc0b9c104112eff4b810f2e556d8c138b4e00c66b6ff7e3

Observation 520082a7-d528-42ce-b08f-fb2e696bcec3 · outbound

This paper cites 12 Liisa Järvilehto, Yongjie Sun, Nami Aiba, Shumpei Haginoya, Hasse Hallström, Julia Korkman, and Pekka Santtila.

Growing Pains: Extensible and Efficient LLM Benchmarking Via Fixed Parameter Calibration 12 Liisa Järvilehto, Yongjie Sun, Nami Aiba, Shumpei Haginoya, Hasse Hallström, Julia Korkman, and Pekka Santtila

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T21:25:50.076377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T16:37:12.962958Z digest=sha256:9c4b1dcd5fc0a9b0f5df672b7afb090cde8e2ab65cbbcccd8024d200fc340552

Observation b0379a1e-4cce-4041-b22a-2f2c6b485316 · outbound

This paper cites Dynabench: Rethinking benchmarking in nlp.

Growing Pains: Extensible and Efficient LLM Benchmarking Via Fixed Parameter Calibration Dynabench: Rethinking benchmarking in nlp

Reference 6

Resolution
metadata mismatch
doi, observed 2026-05-10T21:25:50.079372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T16:37:12.962958Z digest=sha256:063f6cc368e9004bddf91469eff6d31428c8c3e586ee507ed5dea0e3e0f93b37

Observation 471cbdb5-6644-4b77-8c5b-0595ff536f0d · outbound

This paper cites metabench -- A Sparse Benchmark of Reasoning and Knowledge in Large Language Models.

Growing Pains: Extensible and Efficient LLM Benchmarking Via Fixed Parameter Calibration metabench -- A Sparse Benchmark of Reasoning and Knowledge in Large Language Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:30:56.694144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T16:37:12.962958Z digest=sha256:75a25b4aed149c8d5ba0f967a13fec16338e05610e17b2ffe81175542ce9cc5d

Observation a82166b1-a2ce-4875-a12b-acba89703f2b · outbound

This paper cites an unresolved cited work.

Growing Pains: Extensible and Efficient LLM Benchmarking Via Fixed Parameter Calibration Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-05-17T14:29:46.097265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T16:37:12.962958Z digest=sha256:eaff2b2ca9a1420a0fd5da5d7490dc62e86c8ee4ee9a489a52296c3d410b6335

Observation 249dd8f0-5c01-41ed-b27b-4589a87640f4 · outbound

This paper cites and Wu, Hao and Yu, Hong.

Growing Pains: Extensible and Efficient LLM Benchmarking Via Fixed Parameter Calibration and Wu, Hao and Yu, Hong

Reference 9

Resolution
metadata mismatch
doi, observed 2026-05-10T21:25:50.023398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T16:37:12.962958Z digest=sha256:8763d6b8230e64f97a95e97c395667b513e956a004b0cf464d27950fd44d5306

Observation 05ebc3d2-978a-46f9-8be2-eb156ab75157 · outbound

This paper cites Adaptive Testing for LLM Evaluation: A Psychometric Alternative to Static Benchmarks.

Growing Pains: Extensible and Efficient LLM Benchmarking Via Fixed Parameter Calibration Adaptive Testing for LLM Evaluation: A Psychometric Alternative to Static Benchmarks

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-15T01:20:49.385427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T16:37:12.962958Z digest=sha256:9554a462249ec1ff9550537105d27a1425266ffec496654d71cc726cf401ebdf

Observation b887eb82-d5c0-411e-9f9b-c2a441b0f422 · outbound

This paper cites Sasha Luccioni, Yacine Jernite, and Emma Strubell.

Growing Pains: Extensible and Efficient LLM Benchmarking Via Fixed Parameter Calibration Sasha Luccioni, Yacine Jernite, and Emma Strubell

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:29:46.100599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T16:37:12.962958Z digest=sha256:58e7a74761ae4dcaf0edbf63fa070ba309e12533c2c0df5aa46d0d807ca04a31

Observation 90f7d9e2-aaf6-434f-9291-b101cdcd3c25 · outbound

This paper cites Beyond Individual Accountability: (Re-)Asserting Democratic Control of AI.

Growing Pains: Extensible and Efficient LLM Benchmarking Via Fixed Parameter Calibration Beyond Individual Accountability: (Re-)Asserting Democratic Control of AI

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T21:25:50.067120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T16:37:12.962958Z digest=sha256:c93c4317b17b4c2d362e7d451a40d1f437a63754b08fb84e5e73f11681bce66b

Observation 545cf548-90c2-4386-afa0-2b04eaef10d1 · outbound

This paper cites From Benchmarks to Skills: Low-Rank Factors for LLM Evaluation.

Growing Pains: Extensible and Efficient LLM Benchmarking Via Fixed Parameter Calibration From Benchmarks to Skills: Low-Rank Factors for LLM Evaluation

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-06-12T02:09:07.519296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T16:37:12.962958Z digest=sha256:bb953b913d94bf2c9ee7658a4255a52df7e7b97a4ebb7255aaea3655ce687c57

Observation 31b38521-984a-43fd-95eb-dc0200fa0863 · outbound

This paper cites URLhttps://openreview.net/forum?id=qAml3FpfhG.

Growing Pains: Extensible and Efficient LLM Benchmarking Via Fixed Parameter Calibration URLhttps://openreview.net/forum?id=qAml3FpfhG

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:29:46.103509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T16:37:12.962958Z digest=sha256:fdf7eefd588e035a37119fe65cef6d8097929cadf0ad21c2b1f22c153bf54c93

Observation 36b24013-0a66-4c4c-b120-27387ceb91b3 · outbound

This paper cites and Jia, Robin and Boyd-Graber, Jordan.

Growing Pains: Extensible and Efficient LLM Benchmarking Via Fixed Parameter Calibration and Jia, Robin and Boyd-Graber, Jordan

Reference 15

Resolution
metadata mismatch
doi, observed 2026-05-10T21:25:50.070265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T16:37:12.962958Z digest=sha256:2c2b0c1a475f6e8b2eaaf06542c54673580dddd0d0393714b8a9a051208734bf

Observation b6b74848-b548-4bdb-890d-7a7bbfeedb80 · outbound

This paper cites Chothia and A.M.

Growing Pains: Extensible and Efficient LLM Benchmarking Via Fixed Parameter Calibration Chothia and A.M

Reference 16

Resolution
malformed identifier
doi_truncated, observed 2026-05-10T21:25:50.083508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T16:37:12.962958Z digest=sha256:c190b6d7a4a9ebaaf9c1475f033c6e44c1aabeb245dec84637e6e1d9e20db883

Observation ff199c6f-5d59-4821-9ec8-c55a6d0806bb · outbound

This paper cites Reliable and Efficient Amortized Model-based Evaluation.

Growing Pains: Extensible and Efficient LLM Benchmarking Via Fixed Parameter Calibration Reliable and Efficient Amortized Model-based Evaluation

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:30:56.688972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T16:37:12.962958Z digest=sha256:1127e481e9c390b26073dc43830daa66748841f81a75a26dc4a7cd069f91b53d

Pith citing papers

Observation 1b951344-6798-455c-add9-b9fc88e9595e · inbound

Beyond Fixed Benchmarks and Worst-Case Attacks: Dynamic Boundary Evaluation for Language Models cites this paper.

Beyond Fixed Benchmarks and Worst-Case Attacks: Dynamic Boundary Evaluation for Language Models Growing Pains: Extensible and Efficient LLM Benchmarking Via Fixed Parameter Calibration

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-11T20:06:13.583631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T10:13:35.777910Z digest=sha256:1b4fbd931b62681c15dc1f6c940e942fe0a0866159c24aa103e0a7a1cb05eebf

Observation 9c1d734e-88ce-4bb5-a346-8329f24e8f85 · inbound

Every Eval Ever: A Unifying Schema and Community Repository for AI Evaluation Results cites this paper.

Every Eval Ever: A Unifying Schema and Community Repository for AI Evaluation Results Growing Pains: Extensible and Efficient LLM Benchmarking Via Fixed Parameter Calibration

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-07-03T16:58:43.280944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T04:45:20.445703Z digest=sha256:c94e05ce373bdbb508717a1708df02fd2468f991fc1f0b52c9cc9ec6d71c095f