Pith. sign in

Paper Citation Record · LEDGER

AutoEvoEval: An Automated Framework for Evolving Close-Ended LLM Evaluation Data

As of 7 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2506.23735.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.23735 v1

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:45:02.174033Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

13 of 13 outbound references displayed

  • verified exact1
  • verified fuzzy2
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b59611fa-8696-4bb1-9e6f-0df2cd2a98a9 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

AutoEvoEval: An Automated Framework for Evolving Close-Ended LLM Evaluation Data Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T21:45:01.136247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:45:01.136247Z digest=sha256:8ad6e48fff66ea1f03521cb433e8100d6282630eb2e429d7b2a491dea5ebcabb

Observation 025d6689-70c1-4393-902d-71418d1ccd7c · outbound

This paper cites Deepseek-v3: Scaling open-source language models with mixture of experts.

AutoEvoEval: An Automated Framework for Evolving Close-Ended LLM Evaluation Data Deepseek-v3: Scaling open-source language models with mixture of experts

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:45:02.796646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:45:01.188367Z digest=sha256:0d635edde050409103a41b12791d68e04ba094766621aabc6a085562f7dda130

Observation 2a67eaf4-1d63-4ccc-9c49-ca13d65b5496 · outbound

This paper cites Black-box generation of adversarial text sequences to evade deep learning classifiers.

AutoEvoEval: An Automated Framework for Evolving Close-Ended LLM Evaluation Data Black-box generation of adversarial text sequences to evade deep learning classifiers

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T21:45:01.255575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:45:01.255575Z digest=sha256:bd2f6d2f543c5c2e18efbaf9659c196300823fc14a4f75f86fe94fb98fd87381

Observation e6813975-75c2-4d2f-ad10-1df82cbb37cd · outbound

This paper cites Measuring Massive Multitask Language Understanding.

AutoEvoEval: An Automated Framework for Evolving Close-Ended LLM Evaluation Data Measuring Massive Multitask Language Understanding

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T21:45:01.381819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:45:01.381819Z digest=sha256:6fc1be03db506bb90a9b7653173a2b44cf8e9f5fb7448731261475e4fbc71243

Observation e14cfab9-b745-4349-ba58-6caa471c5fef · outbound

This paper cites C-Eval: A Multi-Level Multi-Discipline Chinese Evaluation Suite for Foundation Models.

AutoEvoEval: An Automated Framework for Evolving Close-Ended LLM Evaluation Data C-Eval: A Multi-Level Multi-Discipline Chinese Evaluation Suite for Foundation Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T21:45:01.465673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:45:01.465673Z digest=sha256:ed9ffa80ea3c139fba83f87c36973cf8fcef9dc0adeef3a50e29001f0fd4275b

Observation ea7df36b-156b-492d-ab37-f536b1963162 · outbound

This paper cites Adversarial text generation by search and learning.

AutoEvoEval: An Automated Framework for Evolving Close-Ended LLM Evaluation Data Adversarial text generation by search and learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T21:45:01.552009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:45:01.552009Z digest=sha256:6f3edbaf014fa2c293be92c44521237155b128f3aae2468440bef814db69aa67

Observation baaf731e-2d06-4501-9f47-67da1aceedf5 · outbound

This paper cites Beyond Static Datasets: A Deep Interaction Approach to LLM Evaluation.

AutoEvoEval: An Automated Framework for Evolving Close-Ended LLM Evaluation Data Beyond Static Datasets: A Deep Interaction Approach to LLM Evaluation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T21:45:01.648196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:45:01.648196Z digest=sha256:dd94407f957b9edc5fae76365376bd2089613c8b7515100b8c3d82edd3d534a4

Observation 53df4759-7c2d-4fb4-9e65-93c92e196499 · outbound

This paper cites PertEval: Unveiling Real Knowledge Capacity of LLMs with Knowledge-Invariant Perturbations.

AutoEvoEval: An Automated Framework for Evolving Close-Ended LLM Evaluation Data PertEval: Unveiling Real Knowledge Capacity of LLMs with Knowledge-Invariant Perturbations

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T21:45:01.719577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:45:01.719577Z digest=sha256:e211156cd94c0d88b109c8c976f46c984ba9c01dec00cb8cea2a70f7a63a7152

Observation ca5d8135-ce8e-42d7-bc25-ef1220d1d270 · outbound

This paper cites Using Adversarial Attacks to Reveal the Statistical Bias in Machine Reading Comprehension Models.

AutoEvoEval: An Automated Framework for Evolving Close-Ended LLM Evaluation Data Using Adversarial Attacks to Reveal the Statistical Bias in Machine Reading Comprehension Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T21:45:01.788221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:45:01.788221Z digest=sha256:eb15cab44ff2949ccaf4a515b48680e704ba3a079a8c28b875aee77f524f3d2b

Observation 0b0ab9bf-24d6-45b8-b827-73b7a419d91d · outbound

This paper cites Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering.

AutoEvoEval: An Automated Framework for Evolving Close-Ended LLM Evaluation Data Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T21:45:01.892149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:45:01.892149Z digest=sha256:c46877646e8d5861fcf88cd71968503ed2307b357da44ce543afb38b855c8b31

Observation 6fab41cb-ecc3-4953-aca6-697cfc9ef00b · outbound

This paper cites MedMCQA : A Large-scale Multi-Subject Multi-Choice Dataset for Medical domain Question Answering.

AutoEvoEval: An Automated Framework for Evolving Close-Ended LLM Evaluation Data MedMCQA : A Large-scale Multi-Subject Multi-Choice Dataset for Medical domain Question Answering

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T21:45:01.978036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:45:01.978036Z digest=sha256:78b13448ade918c9ff793f969e8b8cc36a945856545cc5f43089b5fed26cb1e9

Observation 6c0eae8f-4d19-48d9-ac76-bfaf9f4e6033 · outbound

This paper cites Alcuna: Large language models meet new knowledge.

AutoEvoEval: An Automated Framework for Evolving Close-Ended LLM Evaluation Data Alcuna: Large language models meet new knowledge

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:45:02.656878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:45:02.073822Z digest=sha256:be9917836d8e21a8df290d26849a5206ddb07b0ec70582ab5121f286e15cc14c

Observation bfd065d6-b13c-4793-8dbe-5967b7222084 · outbound

This paper cites ALCUNA: Large Language Models Meet New Knowledge.

AutoEvoEval: An Automated Framework for Evolving Close-Ended LLM Evaluation Data ALCUNA: Large Language Models Meet New Knowledge

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:45:02.341402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:45:02.174033Z digest=sha256:de25d0d19cfe3f7c852e837c6e29419ba52ea7863fb869c2f1cc58a6fae55252

Pith citing papers

No inbound Pith citation observations are available.