Pith. sign in

Paper Citation Record · LEDGER

HERO: Improving the Reliability and Sensitivity of Generative Model Evaluation Using Historical Data

As of 20 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 0 inbound Pith citation observations for arXiv:2606.29784.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.29784 v1

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-30T05:41:56.435040Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

30 of 30 outbound references displayed

  • verified exact3
  • verified fuzzy0
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8930321a-5022-41e7-a325-c30f7235164f · outbound

This paper cites Nuanced metrics for measuring unintended bias with real data for text classification,.

HERO: Improving the Reliability and Sensitivity of Generative Model Evaluation Using Historical Data Nuanced metrics for measuring unintended bias with real data for text classification,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-30T05:41:56.435040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T05:41:56.435040Z digest=sha256:a4d4d3c7842c451efe5cdcee04e8842d19c6098f8ebed4d3c4835904115a9fa6

Observation 8a62762d-3055-4f84-963e-b4a593dbe85a · outbound

This paper cites Mllm-as-a-judge: Assessing multimodal llm-as-a-judge with vision-language benchmark,.

HERO: Improving the Reliability and Sensitivity of Generative Model Evaluation Using Historical Data Mllm-as-a-judge: Assessing multimodal llm-as-a-judge with vision-language benchmark,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-30T05:41:56.435040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T05:41:56.435040Z digest=sha256:69b2c207e46fd516932239b96f3ee9dda56ded4a13bf89b20231f197be2f97fd

Observation 667234eb-960d-40bd-80e8-99a2519905b3 · outbound

This paper cites Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference.

HERO: Improving the Reliability and Sensitivity of Generative Model Evaluation Using Historical Data Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-06-30T13:54:44.499146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T05:41:56.435040Z digest=sha256:5a6a4eaea398780a074c0cc42f27f401fdb8fc7564e496161ffd87a17a1d0b4a

Observation 339d2873-2cce-4696-83b0-c60c997ec6bb · outbound

This paper cites A Coefficient of Agreement for Nominal Scales,.

HERO: Improving the Reliability and Sensitivity of Generative Model Evaluation Using Historical Data A Coefficient of Agreement for Nominal Scales,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-30T05:41:56.435040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T05:41:56.435040Z digest=sha256:e7a316e1753cd980f9871f21cdd60d7049f7c48d70b1b1564c881a98289bf5a9

Observation 6e413cf8-56e5-46ce-b2ba-b98c8473f04a · outbound

This paper cites Maximum Likelihood Estimation of Observer Error-Rates Using the EM Algorithm,.

HERO: Improving the Reliability and Sensitivity of Generative Model Evaluation Using Historical Data Maximum Likelihood Estimation of Observer Error-Rates Using the EM Algorithm,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-30T05:41:56.435040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T05:41:56.435040Z digest=sha256:cde7a643012a4218adee1dc27e37ee8206c5c38d9a072c06a563a1426bb5d931

Observation b2a1f518-c3fd-40f8-b1f4-28381f220679 · outbound

This paper cites Maximum likelihood from incomplete data via the EM algorithm,.

HERO: Improving the Reliability and Sensitivity of Generative Model Evaluation Using Historical Data Maximum likelihood from incomplete data via the EM algorithm,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-30T05:41:56.435040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T05:41:56.435040Z digest=sha256:f5edee496321c8cdcd4e918c356182aaccf9bd08b7656a7009e32623a145b64e

Observation b2c8a07d-f209-432e-ba50-2f0ce4578ad6 · outbound

This paper cites Improving the sensitivity of online controlled ex- periments by utilizing pre-experiment data,.

HERO: Improving the Reliability and Sensitivity of Generative Model Evaluation Using Historical Data Improving the sensitivity of online controlled ex- periments by utilizing pre-experiment data,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-30T05:41:56.435040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T05:41:56.435040Z digest=sha256:9458276047696f0b840d7570d184eee7e04629d5145ed3a829590f6b3ae747fd

Observation aedef17a-002e-4a22-a2af-f894a9bc6ec1 · outbound

This paper cites Measuring Nominal Scale Agreement Among Many Raters,.

HERO: Improving the Reliability and Sensitivity of Generative Model Evaluation Using Historical Data Measuring Nominal Scale Agreement Among Many Raters,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-30T05:41:56.435040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T05:41:56.435040Z digest=sha256:7f8e9e1f17e68bd38d87cbcf2208c5e0bb24617aedbf46db7415ed018fd4c5c7

Observation 3ffe8712-b068-48be-b53f-1c937d2f0e05 · outbound

This paper cites Classification in the presence of label noise: a survey,.

HERO: Improving the Reliability and Sensitivity of Generative Model Evaluation Using Historical Data Classification in the presence of label noise: a survey,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-30T05:41:56.435040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T05:41:56.435040Z digest=sha256:0e9f8f68f41d29036da2c5ed09c8e4906f2ee1e6b6947fea960b244a829935d0

Observation cc39b2e8-1b9d-48b8-b798-e8046b194951 · outbound

This paper cites Realtoxicityprompts: Eval- uating neural toxic degeneration in language models,.

HERO: Improving the Reliability and Sensitivity of Generative Model Evaluation Using Historical Data Realtoxicityprompts: Eval- uating neural toxic degeneration in language models,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-30T05:41:56.435040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T05:41:56.435040Z digest=sha256:01ef3f16872bce9a84d50deb82ac6b4a0bfec5875403abc0a92688d1ed550523

Observation 5fc3011a-b3a6-431f-8171-82d740b74be8 · outbound

This paper cites (2004).Monte Carlo methods in financial engineering,53: Springer.

HERO: Improving the Reliability and Sensitivity of Generative Model Evaluation Using Historical Data (2004).Monte Carlo methods in financial engineering,53: Springer

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-30T05:41:56.435040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T05:41:56.435040Z digest=sha256:479da1963bb263b22bc4ec2c44378b9b7e7028a7154504b0e110e4c1f0487b44

Observation 4b5c7a33-17a8-443c-a34b-f5a29e86fcee · outbound

This paper cites A survey on llm-as-a-judge,.

HERO: Improving the Reliability and Sensitivity of Generative Model Evaluation Using Historical Data A survey on llm-as-a-judge,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-30T05:41:56.435040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T05:41:56.435040Z digest=sha256:f0856b187734d93c0d91083ef7ff73760864e03fd82f431da36b7b77ff946056

Observation e076f9c0-b773-4c49-9a80-f12973768470 · outbound

This paper cites Holistic Evaluation of Language Models.

HERO: Improving the Reliability and Sensitivity of Generative Model Evaluation Using Historical Data Holistic Evaluation of Language Models

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-06-30T13:54:44.496914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T05:41:56.435040Z digest=sha256:9efa4c8cda8e260a373552680e0e73a286e14cd45c4ced630716ee3ba57e9026

Observation 4d2f0ddf-7731-46b7-af6a-4c5ababd72d7 · outbound

This paper cites Agnostic notes on regression adjustments to experimental data: Reexamining Freed- man’s critique,.

HERO: Improving the Reliability and Sensitivity of Generative Model Evaluation Using Historical Data Agnostic notes on regression adjustments to experimental data: Reexamining Freed- man’s critique,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-30T05:41:56.435040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T05:41:56.435040Z digest=sha256:b09778232d07d2e034e7a3ddb0e9cb5b60e701a4cae3643dd99f9bd97808a255

Observation 5f1d5b3f-3790-4d53-addd-29738465aa68 · outbound

This paper cites Confident learning: Estimating uncertainty in dataset labels,.

HERO: Improving the Reliability and Sensitivity of Generative Model Evaluation Using Historical Data Confident learning: Estimating uncertainty in dataset labels,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-30T05:41:56.435040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T05:41:56.435040Z digest=sha256:f35ed26c99b94156f731df1951d9644bcae9b4c7623684437d0b332bb0ff0849

Observation b262923b-dbf8-4c97-a316-6bf156b8ba27 · outbound

This paper cites an unresolved cited work.

HERO: Improving the Reliability and Sensitivity of Generative Model Evaluation Using Historical Data Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-30T05:41:56.435040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T05:41:56.435040Z digest=sha256:e93e846c14919bb7612b160dc1f0534b273914ac43aacc820b9311a91891489d

Observation 3bd3ba48-5f6f-47b1-b61d-22cc1e7950b0 · outbound

This paper cites Learning from crowds.,.

HERO: Improving the Reliability and Sensitivity of Generative Model Evaluation Using Historical Data Learning from crowds.,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-30T05:41:56.435040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T05:41:56.435040Z digest=sha256:a455c890b02af1c5cd94d885b892812fa413f38c72975b263c5b9c3e4ed818f8

Observation f8ad6d80-83c5-4fda-9d6c-559f1d53613f · outbound

This paper cites Learning From Crowds,.

HERO: Improving the Reliability and Sensitivity of Generative Model Evaluation Using Historical Data Learning From Crowds,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-30T05:41:56.435040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T05:41:56.435040Z digest=sha256:849ff8779771c07ae7c8f9cf3df1b71c15617b8bcda352f9f4ebd6e355c7cbfe

Observation 6af76126-a76e-442a-8375-90db9996c5a1 · outbound

This paper cites Get another label? improving data quality and data mining using multiple, noisy labelers,.

HERO: Improving the Reliability and Sensitivity of Generative Model Evaluation Using Historical Data Get another label? improving data quality and data mining using multiple, noisy labelers,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-30T05:41:56.435040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T05:41:56.435040Z digest=sha256:5a4236faa07e83d2a40df9c79735f1113a0a42741ecb3a85af2c4ce952c48ed7

Observation 362ddcc2-bc21-4bae-8535-cb93b51a882d · outbound

This paper cites Cheap and Fast—But is it Good? Evaluating Non-Expert Annotations for Natural Language Tasks,.

HERO: Improving the Reliability and Sensitivity of Generative Model Evaluation Using Historical Data Cheap and Fast—But is it Good? Evaluating Non-Expert Annotations for Natural Language Tasks,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-30T05:41:56.435040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T05:41:56.435040Z digest=sha256:ebb0606086be0ddd3348312d033b6f4925b539dde24808c9622dc2a8beaffb02

Observation f9e28ac7-8026-4007-b63c-4f1a3763a9ac · outbound

This paper cites Cheap and fast–but is it good? eval- uating non-expert annotations for natural language tasks,.

HERO: Improving the Reliability and Sensitivity of Generative Model Evaluation Using Historical Data Cheap and fast–but is it good? eval- uating non-expert annotations for natural language tasks,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-06-30T05:41:56.435040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T05:41:56.435040Z digest=sha256:a700c8b9c19422a7179881f56c534d948497d339a753e62f8955a092bc6840e7

Observation bc6b0f54-30b6-46c8-8725-e0d68b5db499 · outbound

This paper cites Exposing flaws of generative model evaluation metrics and their unfair treatment of diffusion models,.

HERO: Improving the Reliability and Sensitivity of Generative Model Evaluation Using Historical Data Exposing flaws of generative model evaluation metrics and their unfair treatment of diffusion models,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-30T05:41:56.435040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T05:41:56.435040Z digest=sha256:005956888e33c9b5a2e680b7f260de6e3acd9af24482ca962b912bafaae6bb53

Observation 00acb5c9-0c8b-4a25-afe5-aee987b96c11 · outbound

This paper cites an unresolved cited work.

HERO: Improving the Reliability and Sensitivity of Generative Model Evaluation Using Historical Data Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-06-30T05:41:56.435040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T05:41:56.435040Z digest=sha256:7080248e7bc98b1c3e56b3393ba95fd4974da60d64565677b923765e76c42a3e

Observation d977f61b-0acd-4377-94fa-d7616d432a21 · outbound

This paper cites An Algorithm for the Validation of Image Segmentation,.

HERO: Improving the Reliability and Sensitivity of Generative Model Evaluation Using Historical Data An Algorithm for the Validation of Image Segmentation,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-30T05:41:56.435040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T05:41:56.435040Z digest=sha256:8cfe48950038e20d3f1f640a0dd29981e85d3fc078cedf4832b35b854e3beba5

Observation 6f756615-b9db-4248-ba42-eb1faca750f8 · outbound

This paper cites Toward an Evaluation Science for Generative AI Systems.

HERO: Improving the Reliability and Sensitivity of Generative Model Evaluation Using Historical Data Toward an Evaluation Science for Generative AI Systems

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:54:44.494494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T05:41:56.435040Z digest=sha256:63537299c9bdacf4ac2e26c63ae02e55a826ec7e8228f27e68fe0acd35d0643d

Observation 5e2a9976-fd1c-406e-bd4c-ddf93737fdc5 · outbound

This paper cites The multidimensional wisdom of crowds,.

HERO: Improving the Reliability and Sensitivity of Generative Model Evaluation Using Historical Data The multidimensional wisdom of crowds,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-06-30T05:41:56.435040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T05:41:56.435040Z digest=sha256:7572616bf17fab847c1144f65eeea01205e9ab2e97eb4697bc2b80d8ba58b3ac

Observation 15afc851-9250-4a15-aae7-6f34323a3b7e · outbound

This paper cites Whose Vote Should Count More: Optimal Integration of Labels from Labelers of Unknown Expertise,.

HERO: Improving the Reliability and Sensitivity of Generative Model Evaluation Using Historical Data Whose Vote Should Count More: Optimal Integration of Labels from Labelers of Unknown Expertise,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-06-30T05:41:56.435040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T05:41:56.435040Z digest=sha256:431a3b3c378b70018d7638270a547d7ec5568a3c68404b87f1c8e52667e75b95

Observation 4c049112-1f97-4c72-87c5-edcbfaa50f9e · outbound

This paper cites Whose vote should count more: Optimal integration of labels from labelers of unknown expertise,.

HERO: Improving the Reliability and Sensitivity of Generative Model Evaluation Using Historical Data Whose vote should count more: Optimal integration of labels from labelers of unknown expertise,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-30T05:41:56.435040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T05:41:56.435040Z digest=sha256:7014d62e8b11ecd538968b8726e6cf20176854343db71e6f37dc19e43be87383

Observation 2b9dd796-8d29-4643-af25-f4e809665a0c · outbound

This paper cites On the convergence properties of the EM algorithm,.

HERO: Improving the Reliability and Sensitivity of Generative Model Evaluation Using Historical Data On the convergence properties of the EM algorithm,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-06-30T05:41:56.435040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T05:41:56.435040Z digest=sha256:3aa08a02bc096d6bba3dd946d5070cd69e349df3cb7a657c5d637733881140d7

Observation e76243c8-4001-472b-b575-1527a56cfeb7 · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena,.

HERO: Improving the Reliability and Sensitivity of Generative Model Evaluation Using Historical Data Judging llm-as-a-judge with mt-bench and chatbot arena,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-06-30T05:41:56.435040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T05:41:56.435040Z digest=sha256:0941ac1c221d158d7d78407e14d471dc20cefd12df9b0a28343c29d3b6430eef

Pith citing papers

No inbound Pith citation observations are available.