Pith. sign in

Paper Citation Record · LEDGER

Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights

As of 9 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2507.06893.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.06893 v1

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:55:34.331074Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

13 of 13 outbound references displayed

  • verified exact0
  • verified fuzzy7
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 78cbdc7a-d2f5-4e63-b11c-d016b53f4b04 · outbound

This paper cites Pairwise analysis of model performance.

Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights Pairwise analysis of model performance

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:55:34.600472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:55:34.247692Z digest=sha256:3248c0a8958ffe6c903240ff5dc9b770193a34be71eee9a69efe7a5895b45605

Observation 77bc16fb-21d6-4a91-8814-da81f917fa3e · outbound

This paper cites Calculating optimal resampling for model evaluation.

Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights Calculating optimal resampling for model evaluation

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:55:34.579888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:55:34.259478Z digest=sha256:9ca8919c84a1f7e82c896c00b207d8b58b7284f69fc0d4a4d2856946cd93d0a0

Observation 5d2cc3ba-3566-4d4c-a820-19a7c16aafe4 · outbound

This paper cites The AI evaluation substack, 2024.

Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights The AI evaluation substack, 2024

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:55:34.553817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:55:34.264917Z digest=sha256:dff1e99cb3e9afb995a1d2ddc4b6a7b3ccbfbb9c5041af8003bec7947f1eda21

Observation e71c644c-e5ec-43b0-88c7-04992f764326 · outbound

This paper cites AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents.

Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T18:55:34.271238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:55:34.271238Z digest=sha256:a76444bc8272490f97e549aef20b3b45716484c07d55f7e3d9fe2f21ab832c58

Observation 12197ec2-6a20-49da-8bea-0eed1617a721 · outbound

This paper cites Autonomous systems evaluation standard.

Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights Autonomous systems evaluation standard

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:55:34.536321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:55:34.280724Z digest=sha256:6648c904a390b8c9df0376f339d00c58fcc7b296f239fd4a589dfb67d3bd0601

Observation 368d4930-bed7-4405-a097-373a5921c3c2 · outbound

This paper cites GAIA: a benchmark for General AI Assistants.

Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights GAIA: a benchmark for General AI Assistants

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T18:55:34.286229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:55:34.286229Z digest=sha256:62d7691c04c3722b520610be3de4b4e83e0fe6c338170ff402b120682d2fa0ce

Observation b3cfc9c6-33ac-45ac-924a-853631b4403a · outbound

This paper cites Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations.

Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T18:55:34.292785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:55:34.292785Z digest=sha256:3a606ff1e0ac2da0c9b4a87a5d36af73964479cf159bd94ae29f7a5ab18a2a9e

Observation 1feabfbd-b7ff-4243-b054-eb37a744cf80 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:55:34.300173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:55:34.300173Z digest=sha256:35f0a3230dce1a47d6408b5644fe63a55078053aa46c546ba92dcb51dfc43cea

Observation aeb4f00d-108e-4a45-bbae-dee22bea3d1c · outbound

This paper cites inspect\_ai: A framework for large language model evaluations, 2024.

Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights inspect\_ai: A framework for large language model evaluations, 2024

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:55:34.519830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:55:34.308264Z digest=sha256:820f40a96cd36a35e07e40f3fcf6c5f587b97f52b8eddcd6764773bd39068166

Observation 02c2801e-22b1-4a59-bb3b-0df2e8d85e00 · outbound

This paper cites Research agenda.

Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights Research agenda

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:55:34.503280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:55:34.313755Z digest=sha256:0c6ad68ccc06acae8398ec1272615cfe0eeb5f53df75c5353f6355bc04269e67

Observation 82e3a5a0-0d7a-4421-a559-d4231eeb1fe1 · outbound

This paper cites inspect\_evals: Collection of evals for Inspect AI , 2024.

Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights inspect\_evals: Collection of evals for Inspect AI , 2024

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:55:34.487244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:55:34.320409Z digest=sha256:d07651606e6437ad442c2b375d7e40cff551ca1fcffcec66d4903f2159eb10f4

Observation fc9251f0-e653-4768-ba17-b7bf3c19b967 · outbound

This paper cites Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models.

Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T18:55:34.325643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:55:34.325643Z digest=sha256:684ed9d78f7948f816aa0289361b13333d954a9dbf66a2d42669a31eab5cabf8

Observation f14fcdfe-824c-47f4-a1c7-a7fd4d137ddd · outbound

This paper cites General Scales Unlock AI Evaluation with Explanatory and Predictive Power.

Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights General Scales Unlock AI Evaluation with Explanatory and Predictive Power

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T18:55:34.331074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:55:34.331074Z digest=sha256:877abd1f6dd4d87373bd2510b7293adbb33c27cebefafcc76c8a18df1d3b1811

Pith citing papers

No inbound Pith citation observations are available.