Pith. sign in

Paper Citation Record · LEDGER

metabench -- A Sparse Benchmark of Reasoning and Knowledge in Large Language Models

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2407.12844.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.12844 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T04:23:53.806181Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T21:28:58.087704Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e90a0b89-3763-4ded-b0ac-f717c25c9ddd · inbound

A Causality-aware Paradigm for Evaluating Creativity of Multimodal Large Language Models cites this paper.

A Causality-aware Paradigm for Evaluating Creativity of Multimodal Large Language Models metabench -- A Sparse Benchmark of Reasoning and Knowledge in Large Language Models

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-10T14:38:25.862946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:38:25.862946Z digest=sha256:0dced21983cb6ed1590bbdc0a96a7977d14d5058df04834820753cbee09b985d

Observation 5d40505a-3946-4f22-8c5e-6efaccecf898 · inbound

Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs cites this paper.

Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs metabench -- A Sparse Benchmark of Reasoning and Knowledge in Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T04:23:53.806181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:23:53.806181Z digest=sha256:1c80e7db40c972966e7ca20a73f88e24caca15a5ff812cc1034293180f07983c

Observation bd35b406-d089-423a-ade4-f0b92e37ae66 · inbound

MiniLongBench: The Low-cost Long Context Understanding Benchmark for Large Language Models cites this paper.

MiniLongBench: The Low-cost Long Context Understanding Benchmark for Large Language Models metabench -- A Sparse Benchmark of Reasoning and Knowledge in Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T14:08:09.475795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:08:09.475795Z digest=sha256:65d7fcd44abc8731d609f4a4f1618d39f0243906f2893ece8e70fe8cf74a0480

Observation 3d6133d1-c2e2-4143-8d7f-6fd5889bdb46 · inbound

Capability Salience Vector: Fine-grained Alignment of Loss and Capabilities for Downstream Task Scaling Law cites this paper.

Capability Salience Vector: Fine-grained Alignment of Loss and Capabilities for Downstream Task Scaling Law metabench -- A Sparse Benchmark of Reasoning and Knowledge in Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T00:42:00.775227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:42:00.775227Z digest=sha256:9f5fac1b1f978401c2a5b360462bcd663564a6f343d0d6338b6d04eccde5edf2

Observation 5e1005c1-f974-401e-b4b1-1b1040bab0af · inbound

Position: Stop Evaluating AI with Human Tests, Develop Principled, AI-specific Tests instead cites this paper.

Position: Stop Evaluating AI with Human Tests, Develop Principled, AI-specific Tests instead metabench -- A Sparse Benchmark of Reasoning and Knowledge in Large Language Models

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-19T02:12:55.759627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-19T02:12:48.586913Z digest=sha256:943664d78fffa5e36d2c5609729ecc4c3a4d47c2da02aabd6708c6690c823613

Observation 1e313c25-af52-415a-8245-306e89ebdefc · inbound

Select Smarter, Not More: Prompt-Aware Evaluation Scheduling with Submodular Guarantees cites this paper.

Select Smarter, Not More: Prompt-Aware Evaluation Scheduling with Submodular Guarantees metabench -- A Sparse Benchmark of Reasoning and Knowledge in Large Language Models

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T09:50:59.571592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T15:47:45.677717Z digest=sha256:5c818c97dc03ab8290f457cb40cdbd9b5a0b4585cfc7101ef7cab1289fe5c0b7

Observation 471cbdb5-6644-4b77-8c5b-0595ff536f0d · inbound

Growing Pains: Extensible and Efficient LLM Benchmarking Via Fixed Parameter Calibration cites this paper.

Growing Pains: Extensible and Efficient LLM Benchmarking Via Fixed Parameter Calibration metabench -- A Sparse Benchmark of Reasoning and Knowledge in Large Language Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:30:56.694144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T16:37:12.962958Z digest=sha256:39723b3b97627f56fa58b497324ed68bf9b8ed5344cdbd38b0ed00e2f2fdb45a

Observation d6229044-6503-4080-99b2-6ee5ecc8fcca · inbound

Beyond Fixed Benchmarks and Worst-Case Attacks: Dynamic Boundary Evaluation for Language Models cites this paper.

Beyond Fixed Benchmarks and Worst-Case Attacks: Dynamic Boundary Evaluation for Language Models metabench -- A Sparse Benchmark of Reasoning and Knowledge in Large Language Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:06:13.574332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T10:13:35.777910Z digest=sha256:e767ee6d5a3ea116f6a3f8baf56c678eba0ebc06c496e34ccd912c8d285f935a

Observation 51e162ca-9c1a-4458-8237-0c2b8b11f6a3 · inbound

Item Response Scaling Laws: A Measurement Theory Approach for Efficient and Generalizable Neural Scaling Estimation cites this paper.

Item Response Scaling Laws: A Measurement Theory Approach for Efficient and Generalizable Neural Scaling Estimation metabench -- A Sparse Benchmark of Reasoning and Knowledge in Large Language Models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-06-28T22:52:45.629897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T22:48:20.193699Z digest=sha256:25167d50d7fd935cf38ff1fc2a3b7dc329fe64ae3812584b8e6745900e5f1414

Observation 39807472-4816-4985-b804-9c3523e29efa · inbound

Every Eval Ever: A Unifying Schema and Community Repository for AI Evaluation Results cites this paper.

Every Eval Ever: A Unifying Schema and Community Repository for AI Evaluation Results metabench -- A Sparse Benchmark of Reasoning and Knowledge in Large Language Models

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-07-03T16:58:43.293342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T04:45:20.445703Z digest=sha256:1f0013b116a9882c33cff8547a5c974b13206f91c278c5d419c5dff0d03b60d6

Observation 9f99d01b-c72f-486b-ac57-0e4c82e19be4 · inbound

AGC-Bench: Measuring Artificial General Creativity cites this paper.

AGC-Bench: Measuring Artificial General Creativity metabench -- A Sparse Benchmark of Reasoning and Knowledge in Large Language Models

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T12:36:56.069253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-07-02T12:33:53.029578Z digest=sha256:64559cc1d63e06b71b31cb44cec350146f06c85c137c5e9ae39b6fe7aca2b05f

Observation 23705eed-4c7b-4b28-81e6-563dc571beeb · inbound

AGC-Bench: Measuring Artificial General Creativity cites this paper.

AGC-Bench: Measuring Artificial General Creativity metabench -- A Sparse Benchmark of Reasoning and Knowledge in Large Language Models

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T21:28:58.089705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-07-03T21:25:14.030920Z digest=sha256:5c67776d7ae6e49c46d6a58bfddfb185769193bdbe71942db4a3e763d8fbc01f

Observation aa4f8dd3-6f04-4252-9483-c0a11dc34989 · inbound

Safety, or Just Capability? A Validity Audit of Agent-Safety Benchmarks cites this paper.

Safety, or Just Capability? A Validity Audit of Agent-Safety Benchmarks metabench -- A Sparse Benchmark of Reasoning and Knowledge in Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T00:45:58.290920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T00:45:58.290920Z digest=sha256:031b8846e4512fc66fcadb0f2187688d1be702dae107a28cf39f26a088b24243

Observation 9241b1ca-3a78-41bb-9ad6-6c35e8772641 · inbound

Item Response Theory for AI Safety cites this paper.

Item Response Theory for AI Safety metabench -- A Sparse Benchmark of Reasoning and Knowledge in Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T05:39:49.447997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:39:49.447997Z digest=sha256:2de2039b21f0d4016486512faa5c139464d09c24397d54f8847557f0d9586e15