Pith. sign in

Paper Citation Record · LEDGER

Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights

As of 13 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2507.06893.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.06893 v1

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:55:34.331074Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

13 of 13 outbound references displayed

  • verified exact0
  • verified fuzzy7
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 78cbdc7a-d2f5-4e63-b11c-d016b53f4b04 · outbound

This paper cites Pairwise analysis of model performance.

Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights Pairwise analysis of model performance

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:55:34.600472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-06T18:55:34.247692Z digest=sha256:74ae0a2151a42cfb3e15753c03756856c17606a662194a18854828b3e856a189

Observation 77bc16fb-21d6-4a91-8814-da81f917fa3e · outbound

This paper cites Calculating optimal resampling for model evaluation.

Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights Calculating optimal resampling for model evaluation

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:55:34.579888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-06T18:55:34.259478Z digest=sha256:90220eaedf2e848f9011f0a246f9b53d5e763918a4ed2ac9ba2848cf7fd2f45f

Observation 5d2cc3ba-3566-4d4c-a820-19a7c16aafe4 · outbound

This paper cites The AI evaluation substack, 2024.

Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights The AI evaluation substack, 2024

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:55:34.553817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-06T18:55:34.264917Z digest=sha256:a41f3b28ac85a5ae19f3184c2c6278bfbe1bc47ea4dc041d62eea2363cb1a367

Observation e71c644c-e5ec-43b0-88c7-04992f764326 · outbound

This paper cites AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents.

Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T18:55:34.271238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:55:34.271238Z digest=sha256:129795fe1425ab7ea8c72138e4136f7b26c4dfb5c3ff76112151869fa1d22b01

Observation 12197ec2-6a20-49da-8bea-0eed1617a721 · outbound

This paper cites Autonomous systems evaluation standard.

Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights Autonomous systems evaluation standard

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:55:34.536321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-06T18:55:34.280724Z digest=sha256:252722bf0e70754daea22190552afdca13aaba1cdf5a6b06b5c668ad3d054157

Observation 368d4930-bed7-4405-a097-373a5921c3c2 · outbound

This paper cites GAIA: a benchmark for General AI Assistants.

Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights GAIA: a benchmark for General AI Assistants

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T18:55:34.286229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:55:34.286229Z digest=sha256:ba7b3a02c714d6392952c1da34917f0d68558e316c0d79b7cd73cde033c9caa7

Observation b3cfc9c6-33ac-45ac-924a-853631b4403a · outbound

This paper cites Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations.

Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T18:55:34.292785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:55:34.292785Z digest=sha256:174a6435f02f85514a29eb4e5588c223cb37ededa780334e5deebe4d2bf2b3e1

Observation 1feabfbd-b7ff-4243-b054-eb37a744cf80 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:55:34.300173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:55:34.300173Z digest=sha256:881294400bb1c5995e0a7df29651bea2a5db81d8f3f758ef542a6c780acdb4fa

Observation aeb4f00d-108e-4a45-bbae-dee22bea3d1c · outbound

This paper cites inspect\_ai: A framework for large language model evaluations, 2024.

Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights inspect\_ai: A framework for large language model evaluations, 2024

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:55:34.519830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-06T18:55:34.308264Z digest=sha256:f2295cb53ea852456d20dc0eb04ee455441c1a832113d71ac41706a0e177d38f

Observation 02c2801e-22b1-4a59-bb3b-0df2e8d85e00 · outbound

This paper cites Research agenda.

Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights Research agenda

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:55:34.503280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-06T18:55:34.313755Z digest=sha256:37dea11b69a9920a6dedc8e3fc377b6682ce0a60e9490a2cae812f6c164d90fa

Observation 82e3a5a0-0d7a-4421-a559-d4231eeb1fe1 · outbound

This paper cites inspect\_evals: Collection of evals for Inspect AI , 2024.

Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights inspect\_evals: Collection of evals for Inspect AI , 2024

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:55:34.487244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-06T18:55:34.320409Z digest=sha256:b6ae03dae1993d587a8dac2e46008fa7473545e273b74e26d34ab60e04cf54b6

Observation fc9251f0-e653-4768-ba17-b7bf3c19b967 · outbound

This paper cites Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models.

Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T18:55:34.325643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:55:34.325643Z digest=sha256:9706b49bd23cdfe0077a54a3830b6302933bd1c0cefadac751c9822404256a78

Observation f14fcdfe-824c-47f4-a1c7-a7fd4d137ddd · outbound

This paper cites General Scales Unlock AI Evaluation with Explanatory and Predictive Power.

Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights General Scales Unlock AI Evaluation with Explanatory and Predictive Power

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T18:55:34.331074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:55:34.331074Z digest=sha256:f0773848dcb32d851ecaf02ecaabb6a7dfa5f2e45dcc89d1bb879869b1b5ed26

Pith citing papers

No inbound Pith citation observations are available.