Pith. sign in

Paper Citation Record · LEDGER

Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights

As of 9 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2507.06893.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.06893 v1

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:55:34.331074Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

13 of 13 outbound references displayed

  • verified exact0
  • verified fuzzy7
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 78cbdc7a-d2f5-4e63-b11c-d016b53f4b04 · outbound

This paper cites Pairwise analysis of model performance.

Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights Pairwise analysis of model performance

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:55:34.600472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T18:55:34.247692Z digest=sha256:9bb7dcd531e0874dcf5cb44afc8419e1e19bf20bd8238d98edb59dc838eddfe4

Observation 77bc16fb-21d6-4a91-8814-da81f917fa3e · outbound

This paper cites Calculating optimal resampling for model evaluation.

Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights Calculating optimal resampling for model evaluation

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:55:34.579888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T18:55:34.259478Z digest=sha256:f3bb355b60c151de4b81405d52e9bb45b4c711362dcf575d9e807a8dbb2134b1

Observation 5d2cc3ba-3566-4d4c-a820-19a7c16aafe4 · outbound

This paper cites The AI evaluation substack, 2024.

Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights The AI evaluation substack, 2024

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:55:34.553817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T18:55:34.264917Z digest=sha256:aa85e72d52e2d3b4a777b73ff2dcb61e780c86b1b22e01532522b0f5f5b1e358

Observation e71c644c-e5ec-43b0-88c7-04992f764326 · outbound

This paper cites AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents.

Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T18:55:34.271238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:55:34.271238Z digest=sha256:a76444bc8272490f97e549aef20b3b45716484c07d55f7e3d9fe2f21ab832c58

Observation 12197ec2-6a20-49da-8bea-0eed1617a721 · outbound

This paper cites Autonomous systems evaluation standard.

Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights Autonomous systems evaluation standard

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:55:34.536321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T18:55:34.280724Z digest=sha256:205b03d5bcc8412617c5a0f735179b036c3a26a411bfbbc93f6916d181a38ad0

Observation 368d4930-bed7-4405-a097-373a5921c3c2 · outbound

This paper cites GAIA: a benchmark for General AI Assistants.

Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights GAIA: a benchmark for General AI Assistants

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T18:55:34.286229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:55:34.286229Z digest=sha256:62d7691c04c3722b520610be3de4b4e83e0fe6c338170ff402b120682d2fa0ce

Observation b3cfc9c6-33ac-45ac-924a-853631b4403a · outbound

This paper cites Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations.

Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T18:55:34.292785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:55:34.292785Z digest=sha256:3a606ff1e0ac2da0c9b4a87a5d36af73964479cf159bd94ae29f7a5ab18a2a9e

Observation 1feabfbd-b7ff-4243-b054-eb37a744cf80 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:55:34.300173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:55:34.300173Z digest=sha256:35f0a3230dce1a47d6408b5644fe63a55078053aa46c546ba92dcb51dfc43cea

Observation aeb4f00d-108e-4a45-bbae-dee22bea3d1c · outbound

This paper cites inspect\_ai: A framework for large language model evaluations, 2024.

Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights inspect\_ai: A framework for large language model evaluations, 2024

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:55:34.519830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T18:55:34.308264Z digest=sha256:5345d9f00129b44dcfc8539cba7aab3ee91bd4879d3e02364cb82a4ee6dba335

Observation 02c2801e-22b1-4a59-bb3b-0df2e8d85e00 · outbound

This paper cites Research agenda.

Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights Research agenda

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:55:34.503280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T18:55:34.313755Z digest=sha256:252e61146037aec76c4b4ccfe093d3391269677d090d6a4e41453bedd7986bd0

Observation 82e3a5a0-0d7a-4421-a559-d4231eeb1fe1 · outbound

This paper cites inspect\_evals: Collection of evals for Inspect AI , 2024.

Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights inspect\_evals: Collection of evals for Inspect AI , 2024

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:55:34.487244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T18:55:34.320409Z digest=sha256:9d21949f6be2e87bafe4b51cfd354eee9cb1adf2c31b63ff93cdb6c7c51e9f2f

Observation fc9251f0-e653-4768-ba17-b7bf3c19b967 · outbound

This paper cites Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models.

Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T18:55:34.325643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:55:34.325643Z digest=sha256:684ed9d78f7948f816aa0289361b13333d954a9dbf66a2d42669a31eab5cabf8

Observation f14fcdfe-824c-47f4-a1c7-a7fd4d137ddd · outbound

This paper cites General Scales Unlock AI Evaluation with Explanatory and Predictive Power.

Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights General Scales Unlock AI Evaluation with Explanatory and Predictive Power

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T18:55:34.331074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:55:34.331074Z digest=sha256:877abd1f6dd4d87373bd2510b7293adbb33c27cebefafcc76c8a18df1d3b1811

Pith citing papers

No inbound Pith citation observations are available.